FINALLY OFFLINE

ANTHROPIC'S BEST CODING MODEL IS THE ONE YOU CANNOT BUY

By FINALLY OFFLINE | Approved by Will Nichols, Editor in Chief | 9/1/2026

In September 2026 Anthropic released Claude Fable 5.1, model ID claude-fable-5-1, and Claude Mythos 5.1, which share the same underlying model but carry different safeguards. Fable 5.1 is generally available through the API and on Amazon Web Services, Google Cloud and Microsoft Azure, as well as Claude.ai, Claude Code, Claude Enterprise and Claude Platform. Mythos 5.1 is restricted to vetted cyberdefenders and life scientists at United States organizations through the Cyber Verification Program and the Life Sciences Verification Program, the latter developed with a US government partner. On Terminal Bench 4.0, Mythos 5.1 scored 60.9 percent, Fable 5.1 scored 55.8 percent, Opus 5 scored 52.3 percent, Fable 5 scored 42.0 percent and GPT 5.6 Sol scored 37.3 percent. Fable 5.1 scored 52.6 percent on Terminal Bench Science 0.1 versus 24.7 percent for Fable 5, 31.4 percent on AutomationBench versus 17.1 percent, 1853 on GDPval AA v2, 73.4 percent on CursorBench 3.2.0, 77.9 percent on OSWorld 2.0 partial and 41.7 percent strict, and 60.9 percent on Humanity's Last Exam without tools and 65.0 percent with tools. Pricing is $10 per million input tokens and $50 per million output tokens, with cache reads reduced 75 percent to $0.25 per million, which Anthropic says amounts to roughly 25 percent lower cost for typical workloads and up to 45 percent for highly agentic tasks. Reported science results include protein binding affinities ten times higher than previous competition winners with about a 50 percent hit rate across 12 targets versus a typical 10 to 15 percent, Venus elevation maps improved from 10 to 20 kilometer resolution to 2 to 3 kilometers with 25 percent greater height accuracy, and GPU kernel optimizations accelerating seven deep learning models by up to 2.5 times with identical outputs. Safety changes include a 60 percent reduction in false positive interventions, 85 percent fewer false positives on benign medical and elementary biology queries, Enterprise Frontier Safeguards rolling out in phases starting fall 2026, invisible watermarking for EU AI Act compliance, and anti distillation measures preventing manual context editing for new API accounts.

Key Points

Two models shipped in September, built on the same underlying system, separated only by what they are allowed to do. Claude Fable 5.1 goes to everyone, on every cloud, in every Anthropic product. Claude Mythos 5.1 goes to vetted cyberdefenders and life scientists at US organizations, and nobody else. On Terminal Bench 4.0, the agentic coding test both were measured against, Fable 5.1 scores 55.8 percent. Mythos 5.1 scores 60.9.

The thesis: Anthropic has stopped pretending its frontier is a product. The best version is now an access program, and the version you can actually buy got cheaper to make you feel fine about that.

60.9 Percent, and You Cannot Have It

Start with the gap, because it is the whole story. On Terminal Bench 4.0, Mythos 5.1 posts 60.9 percent against Fable 5.1 at 55.8, Opus 5 at 52.3, last generation's Fable 5 at 42.0, and GPT 5.6 Sol at 37.3. Five points of separation between the public model and the restricted one, on the benchmark that most closely resembles what an engineer actually does all day.

The rest of the card is genuinely strong for the model you can buy. Terminal Bench Science 0.1 more than doubles generation over generation, 52.6 percent against Fable 5's 24.7. AutomationBench nearly doubles, 31.4 against 17.1. GDPval AA v2 lands at 1853, ahead of Opus 5 at 1824 and well ahead of GPT 5.6 Sol at 1711. Humanity's Last Exam reaches 60.9 percent with no tools at all. OSWorld 2.0 improves on both the partial and strict scoring, 77.9 and 41.7.

That is a real generational jump, and it arrives cheaper. Input runs $10 per million tokens, output $50, and cache reads dropped to $0.25 per million, a 75 percent cut. Anthropic puts the practical saving around 25 percent for typical work and up to 45 percent for heavily agentic loads, framing it as "important steps towards addressing feedback we've received from customers on price, data retention, and safeguards." We covered the last round of this when Opus 5 undercut Fable 5. The direction has not changed.

Mythos Is a Verification Program, Not a Product

You cannot buy Mythos 5.1 at any price. Access runs through two doors: the Cyber Verification Program for defensive security work, and the Life Sciences Verification Program, which Anthropic built with a US government partner. Both are US only.

This is a familiar shape for anyone who has followed the line. Anthropic ended an earlier Mythos lockout back in June, and the reasoning for gating it in the first place was never mysterious, given that a Mythos preview turned up thousands of zero days. Capability that finds vulnerabilities at scale is capability that finds them for anyone holding the keys.

So the split is defensible. It is also a market structure. There is now a tier of model capability that is allocated by institutional vetting and nationality rather than by willingness to pay, and the gap between tiers is measurable in public benchmark points. Fable 5.1 can identify software vulnerabilities but cannot write exploits or run penetration tests. That boundary is the product.

The Protein Numbers Are the Actual Headline

Bury the benchmarks for a second, because the science results are the part that should change how you think about this release. In protein design, Anthropic reports binding affinities ten times higher than previous competition winners, with roughly a 50 percent hit rate across 12 targets against a typical range of 10 to 15 percent. On Venus elevation mapping, resolution improved from 10 to 20 kilometers down to 2 to 3, with 25 percent better height accuracy. Seven deep learning models got GPU kernels up to 2.5 times faster with identical outputs.

Those are not chatbot numbers. A 50 percent hit rate against a 10 to 15 percent baseline is the difference between a screening tool and a design tool, and it is exactly why the life sciences door has a verification program bolted to it.

Read the Anti Distillation Clause Before You Build

Here is the quiet feature nobody will demo. New API accounts can no longer manually edit context, a change made specifically to close distillation techniques. If your product pattern depends on rewriting conversation history server side, check whether you are grandfathered, because the door closed behind you.

There is more in the fine print worth knowing. Enterprise Frontier Safeguards begin phasing in this fall, running detection inside customer controlled cloud infrastructure. Outputs now carry invisible watermarking for EU AI Act compliance, with the detection API in private preview. On the friendlier side, false positive interventions dropped 60 percent versus Fable 5, and benign medical and elementary biology questions see 85 percent fewer false refusals, which is the fix everyone doing real work has been asking for.

The fair counter is that vendor benchmarks are vendor benchmarks. Every number above comes from the company selling the model, none of it is independently replicated yet, and Terminal Bench in particular rewards exactly the agentic behavior Anthropic optimized for. Treat the protein and Venus results as claims until somebody outside the building repeats them.

Verdict: upgrade, and read the API terms before you architect around them. The prediction is less comfortable. The public and restricted tiers were four points apart at Opus 5 and are five apart now, and if that spread keeps widening, the interesting question stops being which model is best and becomes who is cleared to use it.

Frequently Asked Questions

What is the difference between Claude Fable 5.1 and Mythos 5.1?

They are the same underlying model with different safeguards. Fable 5.1 is generally available to everyone. Mythos 5.1 has fewer restrictions on security and biology work and is limited to vetted cyberdefenders and life scientists at US organizations.

How do you get access to Claude Mythos 5.1?

Only through two programs: the Cyber Verification Program for defensive security work, and the Life Sciences Verification Program, developed with a US government partner. Both are currently limited to US organizations. It cannot be bought on the open market.

How much does Claude Fable 5.1 cost?

$10 per million input tokens and $50 per million output tokens, with cache reads at $0.25 per million, a 75 percent reduction. Anthropic estimates roughly 25 percent lower cost for typical workloads and up to 45 percent for highly agentic tasks.

How does Claude Fable 5.1 perform on benchmarks?

55.8 percent on Terminal Bench 4.0, 52.6 percent on Terminal Bench Science 0.1, 1853 on GDPval AA v2, 73.4 percent on CursorBench 3.2.0, 77.9 percent on OSWorld 2.0 partial, and 60.9 percent on Humanity's Last Exam without tools, rising to 65.0 percent with tools.

Where can you use Claude Fable 5.1?

Through the API as claude-fable-5-1, on Amazon Web Services, Google Cloud and Microsoft Azure, and in Claude.ai, Claude Code, Claude Enterprise and Claude Platform.

What is the anti distillation change in Claude Fable 5.1?

New API accounts can no longer manually edit context. Anthropic made the change to close distillation techniques, and it can affect products that rewrite conversation history on the server side.

Topics: ai-safety, llm, benchmarks, developer-tools, terminal-bench, claude-fable-5-1, claude, anthropic, claude-mythos, ai-models

More in tech