Saturday, September 12, 2026Verified technology journalism

OpenAI says enabling two settings tripled its scores on the ARC-AGI-3 benchmark

OpenAI reports that enabling two settings tripled its scores on the ARC-AGI-3 benchmark, the latest iteration of the abstraction and reasoning test designed to measure progress toward artificial general intelligence. The dramatic improvement from adjusting two parameters underscores how much performance may hinge on configuration choices in frontier AI systems rather than raw model scale alone.

OpenAI says enabling two settings tripled its scores on the ARC-AGI-3 benchmark

Two Settings, Three Times the Score: What ARC-AGI-3 Actually Measures

OpenAI tripled its score on ARC-AGI-3 without changing the model. The jump came from turning on two API settings.

GPT-5.6 Sol, OpenAI's reasoning system, scored 13.3% on the ARC-AGI-3 public benchmark using the official test harness 1. When OpenAI enabled retained reasoning and compaction, two features the company already deploys in ChatGPT and Codex, that score climbed to 38.3% 1. For context, the average human tester scores around 48%, according to OpenAI's estimate based on official ARC Prize gameplay logs 1. GPT-5.5, the previous generation, scored just 0.4% 1.

A threefold swing from configuration changes. Not from retraining. Not from a bigger model.

Here is what those two settings do. The first, retained reasoning, lets the model carry its private thinking across turns. The official ARC-AGI-3 harness discards a model's internal reasoning after every action, forcing it to re-derive the rules of each game from scratch on each move 1. The second, compaction, replaces rolling truncation, a system where older actions silently disappear from context as the conversation grows. With compaction, the model summarizes earlier history instead of deleting it 1.

Together, these changes let GPT-5.6 Sol remember what it had figured out about each game and build on that knowledge. OpenAI says the result was roughly 3x the score with 6x fewer output tokens 1.

The tension sits in why ARC-AGI-3 uses a bare harness in the first place.

ARC-AGI-3, created by the ARC Prize Foundation, is the first interactive reasoning benchmark designed to measure human-like intelligence in AI agents 2. Instead of static puzzles, agents explore 2D game environments where they must infer the rules without instructions, plan across multiple steps, and adapt as they go 2. The benchmark scores performance using Relative Human Action Efficiency, which compares how many actions an AI takes to solve each level against a human baseline 3. The ARC Prize Foundation's rationale for a simple, generic harness is that it makes model comparisons fair: every system faces identical constraints, and model shortcomings become more visible 1.

OpenAI's counterargument is that this fairness imposes a cost. If a benchmark strips away the features a model was trained to use, it measures the model in a degraded state. OpenAI's models are trained with retained reasoning and compaction as core parts of how they operate. Running them without those features is like testing a race car on the wrong tires and calling the lap time a measure of the engine.

But OpenAI's proposed fix has a commercial dimension that undercuts the neutrality of its recommendation. The two settings that tripled the score are specific to OpenAI's Responses API 1. The company recommends that developers comparing models use the same settings OpenAI deploys in its own products 1. In practice, that means each lab would need its own optimized harness to post its best score. A leaderboard where every entry runs a different configuration stops being a comparison of models and starts being a showcase for whichever company has the most refined deployment wrapper.

OpenAI's blog post is notably direct about the core problem. The company writes that benchmarks rarely measure models in isolation and also capture less visible choices about API settings, harness design, and prompting 1. That observation is correct and worth taking seriously. It also points, conveniently, toward OpenAI's product.

For builders and investors who track AI progress, the ARC-AGI-3 episode crystallizes a problem that has been building since reasoning models arrived. The ARC Prize Foundation frames the gap between AI and human learning as the defining measure for AGI 2. But when that gap can narrow by 25 percentage points through API configuration alone, what exactly is being measured? The score reflects a bundle of model capability, harness design, and deployment infrastructure, and nobody has established the proportions.

The benchmark worth watching next is not the one posting the highest numbers. It is the one whose methodology forces every model to run under identical conditions, with no lab-specific harness optimizations. Until that standard exists, treat every leaderboard position as provisional: the model is in there somewhere, but so is the plumbing.

References

1.OpenAIopenai.com
2.ARC Prize Foundationarcprize.org
3.ARC Prize Foundationdocs.arcprize.org

Cite this story

ProvenBrief (2026). "OpenAI says enabling two settings tripled its scores on the ARC-AGI-3 benchmark." ProvenBrief. https://provenbrief.com/story/openai-says-enabling-two-settings-tripled-its-scores-on-the-arc-agi-3-benchmark

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Verified19 factual claims in this story were independently checked against primary sources before publication. Read our editorial standards.

Get the next brief in your inbox

One weekly email. Every claim verified against primary sources before we hit send.

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.