An independent test of Anthropic's Claude Fable 5.1 under everyday user constraints found the model solved only one of five scientific benchmark tasks, scoring 20%, far below the company's published 52.6% result. The analysis, published by The New Stack this month, ran both Fable 5.1 and its predecessor Fable 5 through selected problems from Terminal-Bench-Science using a $12 spending cap and 60-turn limit per task, conditions closer to what home users face than the eight-hour, unlimited-budget environment Anthropic used for its official score.

Across the five tasks—spanning mathematics, Earth sciences, engineering, life sciences, and physical sciences—Fable 5.1 completed only the symbolic regression problem, finishing in 11.8 minutes after 27 turns and costing $1.96. Fable 5 failed all five tests, hitting the $12 cost ceiling four times and exhausting all 60 turns twice. The most expensive pair of attempts came on the Lorenz-96 atmospheric reconstruction task, where Fable 5 spent $12.63 over 97.7 minutes and Fable 5.1 used $10.70 across 126 minutes, with both models failing. The reactor safety control challenge generated the highest token output at 157,710 from Fable 5.1, which ran for 40.9 minutes and cost $11.53 before hitting the turn limit without passing. On the foraging cognitive model task, Fable 5.1 declared itself finished after 43 turns and 53.5 minutes, spending $5.65, but the official grader rejected its submission. Fable 5 never completed any task, running for as long as 139.3 minutes on the foraging problem before exhausting its budget at $12.13.

The tester notes that while the 0% and 20% results fall well below Anthropic's published figures of 24.7% for Fable 5 and 52.6% for Fable 5.1, the five-task sample is too small to confirm or contradict the company's doubling claim. The direction matched company assertions, since the newer model performed better. The one task Fable 5.1 solved came from mathematics, the same field where the independent leaderboard shows Fable 5 achieving its strongest performance.

The gap between lab benchmarks and real-world use highlights how testing infrastructure shapes results. When Anthropic announced Fable 5.1, it centered the launch around Terminal-Bench-Science, where models receive a terminal and eight hours to solve research problems independently. The company hasn't disclosed what harness or budget it used to reach its published scores. When the benchmark's own leaderboard evaluated Fable 5, it ran the model through Claude Code at maximum effort, spending $14,180 across 210 attempts—roughly $67 per task. That's more than five times the $12 limit the independent test applied. The tester observed that Fable 5.1 consistently failed faster and cheaper than its predecessor, never hitting the cost ceiling while Fable 5 exceeded it in four of five runs.

The report concludes that typical users likely won't experience much difference between the two models. For work that resembles these benchmark challenges, the harness and budget matter as much as the model itself, according to the analysis. With a purpose-built harness, hours available per task, and substantially larger spending limits, users may approach results closer to Anthropic's published figures. The clearest distinction between the models showed up on cost, where the newer version burned through attempts more quickly but stopped short of budget limits that repeatedly trapped its predecessor. Organizations weighing an upgrade should consider whether their deployment mirrors benchmark conditions or real constraints, since performance diverges sharply between the two environments. If most enterprise applications operate under tighter time and cost boundaries than research labs allow, the practical advantage of the newer model may narrow to incremental speed gains rather than the capability leap the headline benchmark suggests.