CONNECT WITH US
AI & Deeptech

AI & Deeptech

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI logo

Published on

Add as a preferred source on Google
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

A sped-up video of GPT‑5.6 Sol attempting to solve puzzles in the ARC-AGI-3 benchmark, with the official harness (left) and our Responses API harness (right), which retains reasoning and enables compaction. On the leaderboard for this game(opens in a new window), no frontier model solves any level beyond the first. With our harness, GPT‑5.6 Sol solves all six.

When we first saw GPT‑5.6 Sol’s low scores on the ARC-AGI-3(opens in a new window) benchmark, we were puzzled.

GPT‑5.6 Sol has solved longstanding open problems in mathematics like the cycle double cover conjecture(opens in a new window) and beaten games like Pokémon FireRed. But on ARC-AGI-3, a benchmark of 2D puzzle games, GPT‑5.6 Sol scored just 7.8%, and GPT‑5.5 could barely play the games at all, scoring a paltry 0.4%.

Were 2D puzzle games unusually difficult for our models? Or was something else going on?

Benchmarks rarely measure AI models in isolation. They also measure less visible choices about API settings, harness design, and prompting. In the case of ARC-AGI-3, we discovered that turning on two API settings we use in ChatGPT and Codex—retained reasoning and compaction—tripled scores and cut output tokens by 6x on the public task set.

3% on the ARC-AGI-3 public set. 3%. Scores measure Relative Human Action Efficiency ( RHAE ⁠ (opens in a new window) ) — a metric comparing model performance to a human baseline. Based on official gameplay logs ⁠ (opens in a new window) , we estimate the average human tester scored 48%.


Source link

Disclaimer

We strive to uphold the highest ethical standards in all of our reporting and coverage. We TheMorningPulse.fyi want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It's possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.