When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score.
In this benchmark, a model gets a terminal and a real scientific research problem to solve independently. Fable 5.1 scores 52.6%, and Fable 5 scores 24.7%. By Anthropic’s scoring, the new model more than doubles the old one.
Anthropic’s published score was produced under conditions most users don’t have access to. The benchmark allows each model up to eight hours per task, and Anthropic has not said what harness or budget it used to get its numbers. When the benchmark’s leaderboard tested Fable 5, it ran the model through Claude Code at maximum effort and spent $14,180 across 210 attempts, about $67 each.
Most people don’t use Fable in a lab, so I wanted to know what an average user would get on these tasks. I recently tested Fable 5 and Fable 5.1 on everyday work and found them far closer than the benchmark suggests. That made me want to run the benchmark’s own tasks the way a home user would and see where the models actually differ.
The benchmark’s 70 tasks are public, so I pulled five of them, one from each science field, and ran both models myself.
Terminal-Bench-Science has five categories, each with multiple tests. I chose one test per category that could run in a Python environment. Here’s what I picked:
Each run got a plain terminal, and I set a $12 limit and 60 turns for each test.
Source link







