1. A Big Jump in Scientific Work
On Terminal-Bench-Science, Fable 5.1 scored 52.6%.
Fable 5 scored 24.7%.
That is more than double.
These are not simple science Q&A tasks. The benchmark involves work that requires tools and multi-step problem solving.
Anthropic’s examples include protein design, GPU kernel optimization, and generating a high-resolution map of Venus from older NASA data.
That Venus example is especially eye-catching: the resulting map reportedly had close to 10 times the previous resolution.
The hard part here is not knowing scientific facts. It is using tools, checking the result, spotting what went wrong, and continuing from there.
Getting one step right is easy. Staying on track for dozens of steps is much harder.
2. Better at Coding and Automation
Fable 5.1 also improved on Terminal-Bench 4.0 and AutomationBench compared with the previous generation.
One detail matters more than the raw scores: reasoning settings.
According to Anthropic, Fable 5.1 at low or medium reasoning levels can approach the performance that previously required higher reasoning settings on Fable 5.
That matters because stronger reasoning usually means more tokens and higher cost.
If a task that once needed the highest setting can now be handled at a lower one, that can save a meaningful amount of money.
A stronger model is nice. A stronger model that does not burn through tokens as quickly is much more useful.
3. It Can Stay on a Task for a Long Time
Ramp ran Fable 5.1 continuously for 38 hours in one test.
Who normally needs an AI working for 38 straight hours?
Most people do not.
But teams building agents, automation systems, or data workflows absolutely can.
During that run, the model was not just waiting for a human to feed it the next instruction. It found problems, corrected data, launched experiments, and eventually organized the results into a report.
That is where long-running agents tend to fall apart.
Doing three steps correctly is one thing. Reaching step thirty and still remembering why you started is another.
4. Less of the “AI Formatting” Habit
This is a smaller change, but it may be easier to notice in daily use.
Some researchers have reported that Fable 5.1 uses fewer unnecessary headings, bold sections, and bullet lists, and does a better job following requested writing styles.
Anyone who uses AI for long-form writing knows the pattern.
A simple point gets turned into three subheadings, five bullets, and a final “summary.”
It starts to feel like the model is formatting a school assignment.
Fable 5.1 appears to tone that down.
It does not suddenly make the model a great writer, but it does reduce some of the obvious “AI formatting” habits.
What Changed From Fable 5?
Fable 5 felt more like a very strong problem-solving model.
Fable 5.1 pushes further toward something that can actually carry a job from start to finish.
Not just produce a good answer, but investigate, use tools, check its work, keep going, and recover when something breaks.
Combined with the much cheaper cache pricing, Anthropic’s direction is pretty clear: Fable 5.1 is built for people who are treating Claude less like a chatbot and more like a digital worker.
Comments (0)