---
title: "Two API settings tripled OpenAI's score on ARC-AGI-3"
url: https://www.parallelquant.com/posts/two-api-settings-tripled-openai-s-score-on-arc-agi-3-ec381a
source_name: "OpenAI"
source_url: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores
published: 2026-07-29T15:00:00.000Z
topics: ["llms", "research"]
publisher: "Parallel Quant"
---

# Two API settings tripled OpenAI's score on ARC-AGI-3

*2026-07-29 · Source: [OpenAI](https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores)*

OpenAI found that enabling two existing API settings, retaining reasoning between turns and enabling context compaction, tripled GPT-5.6's score on the ARC-AGI-3 benchmark. The change also improved token efficiency, according to OpenAI's write-up.

**Why it matters:** This is a useful reminder that a model's raw capability and its benchmark score are two different things: context and reasoning-retention configuration can matter as much as the underlying model. For developers building agents, it's a concrete, actionable tip rather than a vague 'better prompting' claim, and it shows ARC-AGI-3 remains sensitive to engineering choices, not just model scale.

**Topics:** llms, research

---
Read the original: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores
Canonical: https://www.parallelquant.com/posts/two-api-settings-tripled-openai-s-score-on-arc-agi-3-ec381a
Published by Parallel Quant — https://www.parallelquant.com
