There was no mistaking the divide in March with the release of ARC-AGI-3. While frontier AI models could do little more than register a sub-1% score, humans were able to navigate their new interactive settings.
OpenAI reports a different story for GPT-6 Astra six months on: 98.6%. Put that against the GPT-5.6 Sol it has superseded, which OpenAI puts at 7.8%, and the improvement is hard to miss.
Then again, one has to consider what ARC-AGI is designed for. The whole point is to put models in uncharted interactive territory where they cannot simply rely on training data to find an answer but must work out the mechanics of the environment themselves. Given what ARC-AGI-3 was built to test, a jump to 98.6% is enormous. But the number comes with an important caveat.
Given what ARC-AGI-3 was built to test, a jump to 98.6% is enormous. But the number comes with an important caveat.
Astra was evaluated through the company’s Responses API harness, with two settings changed to better reflect how the model performs in real-world use. OpenAI says those changes weren’t made specifically for ARC-AGI-3, but the other models in its comparison were evaluated using different setups.
ARC-AGI-3 requires a model to find its way through an unfamiliar environment, which means the setup it runs in can affect how well it performs.
The gains aren’t limited to ARC-AGI-3. 2% on SRE-Bench with four attempts. 6%.
Source link







