Home / Blog / How enabling two settings tripled our scores on the…
Tech News

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

**OpenAI** reported that enabling two API settings **tripled** scores for **GPT-5.6** on the **ARC-AGI-3** benchmark. The gains came from **retaining…

By Dillip Chowdary • Aug 04, 2026 • Source: OpenAI News

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

**OpenAI** reported that enabling two API settings **tripled** scores for **GPT-5.6** on the **ARC-AGI-3** benchmark. The gains came from **retaining reasoning** and **enabling compaction**, which together raised measured performance while also improving efficiency on that evaluation.

On the product side, the change is configuration, not a new model release. **Retaining reasoning** keeps intermediate reasoning available across the run instead of discarding it. **Enabling compaction** compresses or condenses that retained state so longer multi-step work stays within usable context. On **ARC-AGI-3**, which stresses abstract, multi-step problem solving, those two levers changed how the same **GPT-5.6** endpoint uses tokens and carries state through a task.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For engineers building agentic or multi-step systems, this is a concrete lever: benchmark and product behavior can shift hard from API flags alone. If you evaluate models without matching **reasoning retention** and **compaction**, you may understate what **GPT-5.6** can do on **ARC-AGI-3**-style workloads and mis-size latency, cost, and failure modes in production agents.

In market terms, the write-up frames progress on **ARC-AGI-3** as partly an inference- and session-configuration problem, not only a training problem. Competitors and evaluators who publish scores without documenting equivalent settings risk non-comparable leaderboards; builders comparing vendors need the same two controls in the test harness if they want a fair read against OpenAI’s reported **3x** result.

Practical takeaway: when you next run **GPT-5.6** on hard multi-step or abstraction tasks, explicitly test **reasoning retention** on and off and **compaction** on and off, log score, token use, and wall time, and pin the winning pair in your client config. Watch for whether OpenAI documents defaults, limits, and cost side effects of those two settings, and whether later evals require them to reproduce the **ARC-AGI-3** triple.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →