July 26, 2026 · The Decoder
Claude Opus 5 nearly quadruples prior record on ARC-AGI-3 benchmark
Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, up from the previous best of 7.8% set by GPT-5.6 Sol. The benchmark's developers said Opus 5 independently formulated reflection equations during testing, a problem-solving behavior they hadn't seen from any other model, and attribute the jump to stronger logical reasoning.
Why it matters: ARC-AGI benchmarks are built specifically to resist memorization and test novel reasoning, so a near-fourfold jump is a meaningful signal about genuine reasoning gains rather than just more training data. Paired with Opus 5's other recent results, it points to a broader capability step-up in this model generation rather than a single cherry-picked win.
Related updates
- Survey: 68% of computer science educators changed exams because of AIJul 26
- OpenAI downgraded GPT-5's risk rating after flagging bio-hazard dangerJul 26
- Induction Labs' Photon-1 learns world simulation without action labelsJul 26
- Kuaishou's KAT-Coder-V2.5 trains coding agents on 100,000 verified environmentsJul 26