July 30, 2026 · The Decoder
OpenAI's ARC-AGI-3 benchmark win relies on a custom test harness
OpenAI says its GPT-5.6 Sol model scored 38.3% on the ARC-AGI-3 benchmark, beating Anthropic's Opus 5, but only when run through OpenAI's own API with retained reasoning and context compaction enabled. In the official, standardized test environment, GPT-5.6 Sol scored just 7.8%, below Opus 5's 30.2%.
Why it matters: This is a direct rebuttal to the standardized ARC-AGI-3 result covered earlier, and the roughly 5x gap between OpenAI's custom-harness score and its official-environment score is a useful reminder that benchmark claims depend heavily on evaluation conditions. It's a live example of the benchmark-gaming dynamics that make it increasingly hard to compare frontier model capability claims at face value.