GPT-6 Astra benchmarks disagree, but ARC-AGI-3 result stands out
Benchmark results for OpenAI's GPT-6 Astra are inconsistent: Epoch AI ranks it in the lead with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. On ARC-AGI-3, though, Astra is more efficient than the average human for the first time. ARC Prize's Francois Chollet says progress there is running "twice as fast" as he expected and is moving up his AGI forecast.
Why it matters: The split verdicts show how much model evaluation still depends on methodology choices rather than raw capability alone, useful context whenever a single benchmark claim circulates. Chollet's forecast revision carries extra weight because he has been one of the more skeptical voices on near-term AGI claims, so a shift from him signals more than typical hype-driven predictions.