A coding agent running for hours can make dozens of decisions as it edits files, runs tests, and works through errors. One wrong turn can carry through the rest of the task unless the agent catches it. SpaceXAI appears to be training Grok for exactly that problem.
The company released Grok 4.7 on Sunday, and its training approach is uniquely different. SpaceXAI used a longer reinforcement learning run deliberately weighted toward harder tasks, including problems that take “many hours” to complete. The company says that training also made Grok better at verifying its own work and managing longer context.
Every failed approach from an agent adds more history for the model to keep straight, and one bad assumption can follow it through the rest of the task. SpaceXAI is trying to address that with better context management and self-verification, so Grok can catch a wrong turn before it builds on it.
SpaceXAI used a longer reinforcement learning run deliberately weighted toward harder tasks, including problems that take “many hours” to complete.
Grok 4.7 scored 38.0% on Terminal-Bench 4.0, up from 20.3% for Grok 4.6. It also improved from 40.4% to 46.3% on CursorBench 4.0, which tests longer-running coding workflows inside the editor, and from 1,546 to 1,657 on AA Briefcase v1.1, an evaluation of multi-hour professional work.
1 here. 6. SpaceXAI says the improvements came from pairing the larger base model with an extended reinforcement learning run deliberately shifted toward harder, multi-hour problems, and that the model specifically improved at two capabilities critical to long-horizon execution: self-verification and long-context management.
Source link







