Anthropic
Claude Fable 5.1
Anthropic's current flagship with adaptive thinking, production safeguards, and the strongest Terminal-Bench 4.0 result published at launch.
API pricing
USD per million tokens · Claude API Standard- Input
- $10
- Output
- $50
- Cached input
- $0.25
- Context window
- 1M
Try the visual runs
Open the generated scene to test it, or compare that prompt with another model.

City Scroll Journey
Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Ember Glider
Sunset gliding journey testing flight energy management, checkpoint flow, and atmospheric scene craft.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low-Poly Tower Defense
Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Neon Drift
Synthwave time-trial racing testing drift physics, lap timing, ghost replay, and unlock progression.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Starfall Arena
Neon arena survival testing wave escalation, upgrade builds, particle feedback, and boss design.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
How Claude Fable 5.1 compares
Three relevant peers, side by side.
- Claude Fable 5.1 This model
- GPT-6 Astra
- Claude Fable 5
- Claude Opus 5
6 benchmarks · 4 models
Official benchmark profile
Claude Fable 5.1 launch results for agentic coding, research, and automation.
Full official benchmark table16 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| Humanity’s Last Exam (with tools) | Reasoning | 65.0% | |
| Artificial Analysis Intelligence Index v4.1.1 | Reasoning | 65.7 | |
| FrontierMath Tier 4 (v2) | Reasoning | 87.8% | |
| DeepSWE v1.1 | Coding | 67.4% | |
| FrontierCode 1.1 Extended (score) | Coding | 63.6% | |
| FrontierCode 1.1 Main (score) | Coding | 50.9% | |
| GPQA Diamond | Science | 93.7% | |
| ARC-AGI-2 | Reasoning | 90.0% | |
| ARC-AGI-1 | Reasoning | 97.5% | |
| BenchCAD (with tools) | Professional work | 84.3% | |
| Database migration (OpenAI internal) | Coding | 57.8% | |
| Computer-use safety (OpenAI internal)Lower is better | Alignment | 9.5% | |
| Terminal-Bench-Science 0.1 | Agentic research | 52.6% | |
| Terminal-Bench 4.0 | Agentic coding | 55.8% | |
| CursorBench 3.2.0 | Agentic coding | 73.4% | |
| AutomationBench | Multi-step automation | 31.4% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Current
- API model ID
claude-fable-5-1- Context
- 1M
- Max output
- 128K
- License
- Not publicly verified
- Access route
- Not publicly verified
- Modalities
- text input, image input, text output
- Hardware
- Not publicly verified
- Claude API Standard · USD / 1M tokens
- $10 input · $0.25 cached input · $12.5 cache write · $50 output
Cache write shown is the 5-minute rate; 1-hour cache writes cost $20 per 1M tokens. Standard global routing; Batch, Fast mode where available, and regional pricing differ.
Benchmark sources
Anthropic released Claude Fable 5.1 on September 1, 2026 as the successor to Fable 5 at the same $10 / $50 per 1M token prices with cache reads cut to $0.25. The launch chart reports gains over Fable 5 and Opus 5 on terminal-based coding, scientific research, and multi-step automation.
Introducing Claude Fable 5.1Fable 5.1 was evaluated with its production safeguards enabled, so Anthropic notes these scores likely understate the model. Only rows with a published Fable 5.1 number are listed; the same-day system card is the deeper source.
Evaluation settings
- OpenAI Astra launch comparison table, checked September 7, 2026. Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
- Lower is better. OpenAI Astra launch comparison table, checked September 7, 2026. Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
- Anthropic launch chart; standard error about 3.5 to 4.5 points per model. Anthropic reproduces Opus 5 at 29.0% and Fable 5 at 24.7% in its own setup.
- Anthropic launch chart. Mythos 5.1, the same model without production safeguards, scores 60.9%.
- Anthropic launch chart; Cursor production agent harness.
- Anthropic launch chart. Safeguard interventions score zero, which Anthropic says lowers Fable 5 in particular.
- Humanity’s Last Exam (with tools)
- Source
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer. Source
- FrontierMath Tier 4 (v2)
- Source
- DeepSWE v1.1
- Source
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8). Source
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8). Source
- GPQA Diamond
- Source
- ARC-AGI-2
- Source
- ARC-AGI-1
- Source
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated. Source
- Database migration (OpenAI internal)
- Source
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ. Source
Comparison methodology
Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.
- Terminal-Bench Science 0.1
- Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- OSWorld 2.0 · offline, v2026.08.08
- Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
- ScreenSpot-Pro (no tools)
- The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
- HealthBench Professional (length-adjusted)
- OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
- ExploitBench
- Astra and Sol evaluated without production safeguards.
- ExploitGym
- The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
- ExploitBench (June–August 2026)
- OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
- SRE-Bench (single attempt)
- Cyber capability evaluation without production safeguards.
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
- ARC-AGI-3
- OpenAI Responses harness with two modified settings; see source footnote 1.
Editorial verdict
Where this model fits
Fable 5.1 replaced Fable 5 on September 1, 2026 as Anthropic's top model. It ships with the same $10 / $50 per 1M token pricing but cuts cache reads to $0.25 and enables adaptive thinking by default. The launch chart reports 55.8% on Terminal-Bench 4.0 and 73.4% on CursorBench 3.2.0 with production safeguards enabled.
Best for
Watch out
Why it is ranked here
Evidence
Superbash editorial model ranking
Takeaway: Claude Fable 5.1 is currently placed in Tier S.
The September 2026 editorial roster places Claude Fable 5.1 at rank 2.
Open source →Related guides