Anthropic · Intelligence Briefing · September 2026
Claude Fable 5.1: one frontier model, two carefully different fables
Anthropic's newest flagship pairs record reasoning benchmarks with cache reads 75% cheaper than its predecessor — while a tightly restricted twin, Claude Mythos 5.1, quietly carries the capabilities Fable was built to hold back.
66
Intelligence Index, Max effort
↓75%
Cache-read cost vs. Fable 5
1M
Token context window
5
Configurable reasoning levels
I.
The split foundation
Same underlying weights, two names. Fable 5.1 ships broadly, wrapped in extra safeguards; Mythos 5.1 stays behind an approval wall reserved for vetted cybersecurity and life-sciences organizations who need the version without the training wheels.
How a request gets routed
1
Request submitted to Claude Fable 5.1
Developer or consumer sends a prompt at a chosen reasoning effort level.
2
Automated safety classifiers evaluate it
Rebuilt safeguards cut false-positive interventions by 85% on benign biology requests, while still supporting source-code security analysis without enabling exploit generation.
3
A trigger reroutes the request automatically
Configured through the Fallback API; routing can stay "sticky" for the rest of a multi-turn conversation.
Cybersecurity trigger → Claude Opus 4.8
Biology trigger → Claude Opus 5
4
Response returns with full transparency
Consumer apps show a model-switch notice; the API returns a fallback content block naming which model actually answered, its reasoning effort, and tokens used.
Enterprise data governance
Standard deployment
30-day data retention for safety monitoring.
Enterprise Frontier Safeguards
Customer-controlled infrastructure for sensitive data.
Zero-data-retention
Available to qualifying organizations on request.
II.
The benchmarks
At maximum reasoning effort, Fable 5.1 posts the highest Artificial Analysis Intelligence Index score Anthropic has published for the family — and roughly doubles its predecessor on the hardest agentic and scientific tests.
Artificial Analysis Intelligence Index (0–100, Max effort)
Claude Fable 5.1New flagship
66
Claude Opus 5Prior top model
63
Claude Fable 5Previous generation
62
GPT-5.6 SolCompeting model
61
Scale 0–100 · full axis shown, not truncated
ARC-AGI, at Max effort
97.5%
ARC-AGI-1
Near-total clearance of the original benchmark.
90.0%
ARC-AGI-2
The harder successor benchmark, still cleared at 90%.
Knowledge work (AA-Briefcase, 1,694 Elo — tied with Opus 5's 1,685)
Fable 5.1 attempts 93.4% of knowledge questions, versus Opus 5's 87.8% — and gets 67.2% of what it attempts right.
Analytical quality
2,025 Elo
vs. Opus 5 at 1,980 — Fable 5.1 ahead
Presentation quality
1,495 Elo
vs. Opus 5 at 1,572 — Opus 5 ahead
III.
The ledger
Reasoning effort is now a dial, not a fixed setting — and Anthropic cut the cost of remembering by three-quarters.
Five reasoning effort levels
Low$0.77 / task
58
Medium$1.00 / task
60
High$1.43 / task · default
62
XHigh$2.72 / task
65
MaxHighest cost
66
Intelligence Index by effort level, scale 0–100 · darker = higher effort
Fable 5.1 pricing, per million tokens
Service
Rate
Standard uncached input
$10.00
Generated output
$50.00
5-minute cache write
$12.50
1-hour cache write
$20.00
Cache read −75%
$0.25
Batch input (−50%)
$5.00
Batch output (−50%)
$25.00
Cache-read cost across the Claude line-up
Fable 5.1
$0.25
Fable 5
$1.00
Opus 5
$0.50
Sonnet 5
$0.20
What that means in practice
~25%
Estimated cost reduction on typical cached workloads
~45%
Savings ceiling on highly agentic, cache-heavy workloads
Where it runs
Claude APIAmazon BedrockGoogle CloudMicrosoft FoundryClaude Platform on AWSOpenRouter
IV.
The discoveries
Outside the benchmark suite, early-access partners pointed Fable 5.1 — and its restricted sibling Mythos 5.1 — at three genuinely hard scientific problems.
Planetary science
Venus, remapped from 30-year-old radar
Fable 5.1 reprocessed NASA's Magellan radar imagery — gathered more than three decades ago — into a new high-resolution elevation map covering roughly a third of the planet, released under a Creative Commons license.
2–3 km
New observable detail scale, vs. 10–20 km before
+25%
Improvement in height accuracy
CC
Released under Creative Commons license
Molecular design · Claude Mythos 5.1
Protein binders at 3–5× the usual hit rate
Tested across 12 protein targets and validated through external laboratory testing, with binding affinity up to ~10× stronger than baseline on three of them.
Typical industry rate
10–15%
Mythos-designed
~50%
Computational biology
GPU optimization in days, not weeks
Seven open-source deep-learning models re-optimized for genome-scale analysis, at a fraction of the usual manual engineering time.
Up to 2.5×
GPU performance improvement
30–60%
Estimated GPU cost savings
V.
The field
Ten early-access organizations, one pattern: Fable 5.1 stays on task longer, and keeps finding the thing nobody was looking for.
Millennium
Investment management
Surfaced a vendor library bug that had gone unresolved for years.
MongoDB
Database software
Built a functional prototype in roughly 3 days.
Ramp
Financial technology
Completed a 38-hour unattended ML investigation.
Jane Street
Quantitative finance
Sharper coding and trading-intuition performance.
IMC
Trading research
Surfaced novel analytical approaches.
Rakuten
Clinical research
Identified a research gap earlier review had missed.
Red Hat
Enterprise software
Diagnosed root causes across multiple builds.
Datadog
Cloud monitoring
Diagnosed complex production incidents.
Plaid
Financial technology
Traced workflows across multiple codebases.
Shopify
E-commerce
Maintained task continuity over an extended session.
Field dispatch
Inside Ramp's 38-hour unattended run
1
Opens with the assigned hypothesis
2
Reassesses it against the data — and doubts it
3
Finds a labeling artifact skewing the results
4
Corrects the analysis, launches 6 parallel experiments overnight
5
Returns with findings and a recommended next step
Unattended, start to finish.
VI.
The verdict
Faster benchmarks and cheaper caching come with fine print worth reading before you scale a workload on it.
1
Subscription math differs from API math
API developers capture most of the cache-read savings. Max and premium seats can run Fable 5.1 for up to 50% of their weekly limit; Pro and standard seats fall back to pay-as-you-go usage credits — and adaptive thinking runs at High effort by default, with thinking tokens counted against that usage.
2
Fallback routing is designed to be visible
A model switch shows as a notice in consumer apps; the API returns a dedicated fallback content block; usage metadata logs every model actually queried, and routing can stay sticky across a conversation.
The asterisk on every score above
Published benchmarks ran with production safeguards switched on. Some OSWorld tasks scored zero when a safeguard intervened mid-task, and cybersecurity or biology triggers routed those specific tasks to Opus 4.8 or Opus 5 instead. That means Fable 5.1's real ceiling is likely higher than the figures reported in Chapter II.
"
Fable and Mythos share one foundation model split across two safeguard profiles — not two different levels of intelligence. It's how Anthropic ships frontier capability broadly without lowering the floor for everyday use, while keeping the sharpest edge of it behind a door only a few vetted teams can open.