parallelquant
July 26, 2026 · The Decoder

Claude Opus 5 nearly quadruples prior record on ARC-AGI-3 benchmark

Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, up from the previous best of 7.8% set by GPT-5.6 Sol. The benchmark's developers said Opus 5 independently formulated reflection equations during testing, a problem-solving behavior they hadn't seen from any other model, and attribute the jump to stronger logical reasoning.

Why it matters: ARC-AGI benchmarks are built specifically to resist memorization and test novel reasoning, so a near-fourfold jump is a meaningful signal about genuine reasoning gains rather than just more training data. Paired with Opus 5's other recent results, it points to a broader capability step-up in this model generation rather than a single cherry-picked win.

Related updates