---
title: "Claude Opus 5 nearly quadruples prior record on ARC-AGI-3 benchmark"
url: https://www.parallelquant.com/posts/claude-opus-5-nearly-quadruples-prior-record-on-arc-agi-3-benchmark-91c8f3
source_name: "The Decoder"
source_url: https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/
published: 2026-07-26T09:43:02.000Z
topics: ["llms", "research"]
publisher: "Parallel Quant"
---

# Claude Opus 5 nearly quadruples prior record on ARC-AGI-3 benchmark

*2026-07-26 · Source: [The Decoder](https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/)*

Anthropic's Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, up from the previous best of 7.8% set by GPT-5.6 Sol. The benchmark's developers said Opus 5 independently formulated reflection equations during testing, a problem-solving behavior they hadn't seen from any other model, and attribute the jump to stronger logical reasoning.

**Why it matters:** ARC-AGI benchmarks are built specifically to resist memorization and test novel reasoning, so a near-fourfold jump is a meaningful signal about genuine reasoning gains rather than just more training data. Paired with Opus 5's other recent results, it points to a broader capability step-up in this model generation rather than a single cherry-picked win.

**Topics:** llms, research

---
Read the original: https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/
Canonical: https://www.parallelquant.com/posts/claude-opus-5-nearly-quadruples-prior-record-on-arc-agi-3-benchmark-91c8f3
Published by Parallel Quant — https://www.parallelquant.com
