parallelquant
September 5, 2026 · The Decoder

Benchmark site revises index after GPT-6 Astra score doubts

Artificial Analysis released version 4.2 of its Intelligence Index after criticism that earlier benchmarks understated GPT-6 Astra's real-world progress. Under the new scoring, Astra rates four points above its predecessor but still trails Anthropic's Claude Fable 5.1.

Why it matters: Benchmark credibility is central to how the industry judges model quality, so a public revision under pressure signals that evaluation methods are struggling to keep pace with newer model capabilities. That Astra still trails Claude Fable 5.1 even after recalibration reinforces Anthropic's competitive standing at the frontier.

Related updates