Anthropic

Tier S · Ship-to-prod default.

Claude Fable 5.1

Anthropic's current flagship with adaptive thinking, production safeguards, and the strongest Terminal-Bench 4.0 result published at launch.

API pricing

USD per million tokens · Claude API Standard
Input
$10
Output
$50
Cached input
$0.25
Context window
1M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

Fable 5.1 replaced Fable 5 on September 1, 2026 as Anthropic's top model. It ships with the same $10 / $50 per 1M token pricing but cuts cache reads to $0.25 and enables adaptive thinking by default. The launch chart reports 55.8% on Terminal-Bench 4.0 and 73.4% on CursorBench 3.2.0 with production safeguards enabled.

Best for

  • Difficult planning and architecture
  • Agentic coding and automation workflows
  • Terminal-based research tasks
  • Long-horizon coding sessions

Watch out

  • Production safeguards are always on and likely suppress some scores
  • Premium API pricing
  • The effort parameter controls thinking depth, not capability

Why it is ranked here

  1. Anthropic launched Fable 5.1 on September 1, 2026 as the direct successor to Fable 5 at identical pricing.
  2. Terminal-Bench 4.0 at 55.8% is the standout new number; the same harness places Fable 5 at 42.0%.
  3. The production-safeguard note means the published scores understate the model's true capability.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: Claude Fable 5.1 is currently placed in Tier S.

The September 2026 editorial roster places Claude Fable 5.1 at rank 2.

Open source →

How Claude Fable 5.1 compares

Three relevant peers, side by side.

  • Claude Fable 5.1 This model
  • GPT-6 Astra
  • Claude Fable 5
  • Claude Opus 5

6 benchmarks · 4 models

Terminal-Bench 4.0

Coding · Higher is better

Source
Claude Fable 5.155.8%
GPT-6 Astra57.9%
Claude Fable 544.5%
Claude Opus 552.6%

DeepSWE v1.1

Coding · Higher is better

Source
Claude Fable 5.167.4%
GPT-6 Astra74.1%
Claude Fable 569.9%
Claude Opus 573.7%

AutomationBench

Automation · Higher is better

Source
Claude Fable 5.131.4%
GPT-6 Astra41.4%
Claude Fable 517.4%
Claude Opus 526.9%

Terminal-Bench Science 0.1

Science · Higher is better

Source
Claude Fable 5.152.6%
GPT-6 Astra64.6%
Claude Fable 521.4%
Claude Opus 530.0%

Humanity’s Last Exam (with tools)

Reasoning · Higher is better

Source
Claude Fable 5.165.0%
GPT-6 Astra57.2%
Claude Fable 563.8%
Claude Opus 563.6%

Artificial Analysis Intelligence Index v4.1.1

Reasoning · Higher is better

Source
Claude Fable 5.165.7
GPT-6 Astra61.2
Claude Fable 562.1
Claude Opus 563.1

Official benchmark profile

Claude Fable 5.1 launch results for agentic coding, research, and automation.

Anthropic sourceSeptember 2026Anthropic model page →
Full official benchmark table16 scores and peer comparisons
BenchmarkAreaScoreComparison
Humanity’s Last Exam (with tools)Reasoning65.0%
Claude Fable 5.165.0%
Claude Fable 563.8%
Claude Opus 563.6%
GPT-6 Astra57.2%
Artificial Analysis Intelligence Index v4.1.1Reasoning65.7
FrontierMath Tier 4 (v2)Reasoning87.8%
Claude Fable 5.187.8%
Claude Fable 590.2%
GPT-5.6 Sol83.0%
GPT-6 Astra97.6%
DeepSWE v1.1Coding67.4%
Claude Fable 5.167.4%
Claude Fable 569.9%
GPT-5.6 Sol72.7%
Claude Opus 573.7%
FrontierCode 1.1 Extended (score)Coding63.6%
Claude Fable 5.163.6%
Claude Opus 563.6%
GPT-6 Astra64.5%
Claude Fable 564.9%
FrontierCode 1.1 Main (score)Coding50.9%
Claude Fable 5.150.9%
GPT-6 Astra53.3%
Claude Opus 553.4%
Claude Fable 553.5%
GPQA DiamondScience93.7%
Claude Fable 5.193.7%
Claude Opus 593.7%
GPT-5.6 Sol94.6%
Claude Fable 592.6%
ARC-AGI-2Reasoning90.0%
Claude Fable 5.190.0%
Claude Opus 590.4%
Claude Fable 589.2%
GPT-5.6 Sol92.5%
ARC-AGI-1Reasoning97.5%
Claude Fable 5.197.5%
GPT-5.6 Sol97.5%
Claude Opus 597.5%
GPT-6 Astra98.5%
BenchCAD (with tools)Professional work84.3%
Database migration (OpenAI internal)Coding57.8%
Claude Fable 5.157.8%
GPT-6 Astra63.9%
Claude Fable 550.3%
GPT-5.6 Sol42.7%
Computer-use safety (OpenAI internal)Lower is betterAlignment9.5%
Terminal-Bench-Science 0.1Agentic research52.6%
Claude Fable 5.152.6%
Claude Opus 529.0%
Claude Fable 524.7%
GPT-5.6 Sol22.4%
Terminal-Bench 4.0Agentic coding55.8%
Claude Fable 5.155.8%
Claude Mythos 5.160.9%
Claude Fable 542.0%
CursorBench 3.2.0Agentic coding73.4%
AutomationBenchMulti-step automation31.4%
Claude Fable 5.131.4%
Claude Fable 517.1%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
claude-fable-5-1
Context
1M
Max output
128K
License
Not publicly verified
Access route
Not publicly verified
Modalities
text input, image input, text output
Hardware
Not publicly verified
Claude API Standard · USD / 1M tokens
$10 input · $0.25 cached input · $12.5 cache write · $50 output

Cache write shown is the 5-minute rate; 1-hour cache writes cost $20 per 1M tokens. Standard global routing; Batch, Fast mode where available, and regional pricing differ.

Benchmark sources

Anthropic released Claude Fable 5.1 on September 1, 2026 as the successor to Fable 5 at the same $10 / $50 per 1M token prices with cache reads cut to $0.25. The launch chart reports gains over Fable 5 and Opus 5 on terminal-based coding, scientific research, and multi-step automation.

Introducing Claude Fable 5.1

Fable 5.1 was evaluated with its production safeguards enabled, so Anthropic notes these scores likely understate the model. Only rows with a published Fable 5.1 number are listed; the same-day system card is the deeper source.

Evaluation settings

  • OpenAI Astra launch comparison table, checked September 7, 2026. Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
  • Lower is better. OpenAI Astra launch comparison table, checked September 7, 2026. Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
  • Anthropic launch chart; standard error about 3.5 to 4.5 points per model. Anthropic reproduces Opus 5 at 29.0% and Fable 5 at 24.7% in its own setup.
  • Anthropic launch chart. Mythos 5.1, the same model without production safeguards, scores 60.9%.
  • Anthropic launch chart; Cursor production agent harness.
  • Anthropic launch chart. Safeguard interventions score zero, which Anthropic says lowers Fable 5 in particular.
Humanity’s Last Exam (with tools)
Source
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer. Source
FrontierMath Tier 4 (v2)
Source
DeepSWE v1.1
Source
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8). Source
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8). Source
GPQA Diamond
Source
ARC-AGI-2
Source
ARC-AGI-1
Source
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated. Source
Database migration (OpenAI internal)
Source
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ. Source

Comparison methodology

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.

Anthropic model report

Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.