July 29, 2026 · OpenAI
Two API settings tripled OpenAI's score on ARC-AGI-3
OpenAI found that enabling two existing API settings, retaining reasoning between turns and enabling context compaction, tripled GPT-5.6's score on the ARC-AGI-3 benchmark. The change also improved token efficiency, according to OpenAI's write-up.
Why it matters: This is a useful reminder that a model's raw capability and its benchmark score are two different things: context and reasoning-retention configuration can matter as much as the underlying model. For developers building agents, it's a concrete, actionable tip rather than a vague 'better prompting' claim, and it shows ARC-AGI-3 remains sensitive to engineering choices, not just model scale.