Anthropic

Tier S · Ship-to-prod default.

Claude Opus 5.5

A leading choice for long coding and research jobs, with strong published results and ten inspectable visual runs.

API pricing

USD per million tokens · Claude API Standard
Input
$4
Output
$20
Cached input
$0.2
Context window
1M
Pricing details and specifications

Try the active visual tests

Open the generated scene to test it, or compare that prompt with another model.

10 runs

City Scroll Journey

Run date: 2026-09-23

Specialist test · evaluated separately

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Low-Poly Tower Defense

Run date: 2026-09-23

Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Mechanical Watch Simulator

Run date: 2026-09-23

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Run date: 2026-09-23

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly

Run date: 2026-09-23

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Editorial verdict

Where this model fits

Back to all benchmarks →

Claude Opus 5.5 is Anthropic’s September 2026 model for long-running agentic coding and knowledge work. Its Claude API rate is $4 input and $20 output per million tokens, with a one-million-token context window. Anthropic reports 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 Main at maximum effort. Ten same-prompt Superbash visual artifacts are available for inspection; eight were time-capped during validation. The S-tier placement is editorial guidance, not a normalized cross-model score.

Best for

  • Long, multi-file coding work
  • Research and knowledge-work agents
  • Complex tool-using workflows
  • Reviewed product builds

Watch out

  • Anthropic-reported scores depend on effort, harness, and production safeguards
  • Eight of ten local runs ended validation early; inspect each report
  • Always-on adaptive thinking changes migration behavior from Opus 5
  • Token pricing does not measure the cost of these local runs

Why it is ranked here

  1. Anthropic’s September 22 launch table reports gains over Opus 5 across agentic coding and knowledge-work evaluations; its Terminal-Bench 4.0 table includes GPT-6 Astra with disclosed effort differences.
  2. The documented API offers a one-million-token context window and lower input, output, and cache-read rates than Opus 5.
  3. Ten prompt-only visual builds are published with manifests and run-specific validation reports. Their quality must be judged by opening the artifacts rather than treating their count as a score.

Evidence

2026-09-23

Superbash editorial model ranking

Takeaway: Claude Opus 5.5 is currently placed in Tier S.

The September 2026 editorial roster places Claude Opus 5.5 at rank 2.

Open source →
2026-09-22

Claude Opus 5.5 launch and visual-run methodology

Takeaway: Anthropic publishes model specifications and benchmark scores; ten separate Superbash visual runs are linked from this model page.

Provider results and local same-prompt artifacts are separate evidence. The local methodology records the validation time cap.

Open source →

How Claude Opus 5.5 compares

Three relevant peers, side by side.

  • Claude Opus 5.5 This model
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra

6 benchmarks · 4 models

Terminal-Bench 4.0

Agentic coding · Higher is better

Source
Claude Opus 5.566.4%
Claude Fable 5.155.8%
Claude Opus 552.3%
GPT-6 Astra57.9%

FrontierCode v1.1 Main

Agentic coding · Higher is better

Source
Claude Opus 5.554.4%
Claude Fable 5.150.3%
Claude Opus 548.0%
GPT-6 Astra53.3%

CursorBench 4.0

Agentic coding · Higher is better

Source
Claude Opus 5.557.8%
Claude Fable 5.151.8%
Claude Opus 546.6%
GPT-6 AstraNot reported

GDPval-AA v2.1

Knowledge work · Higher is better

Source
Claude Opus 5.51846
Claude Fable 5.11735
Claude Opus 51708
GPT-6 Astra1542

AutomationBench

Business workflows · Higher is better

Source
Claude Opus 5.540.0%
Claude Fable 5.131.4%
Claude Opus 526.9%
GPT-6 Astra41.4%

Humanity's Last Exam

Reasoning · Higher is better

Source
Claude Opus 5.567.7%
Claude Fable 5.165.6%
Claude Opus 563.6%
GPT-6 Astra57.2%

Official benchmark profile

Anthropic’s published coding and knowledge-work results

Anthropic sourceSeptember 22, 2026Read the launch report →
Agentic coding

Terminal-Bench 4.0

66.4%
Agentic coding

FrontierCode v1.1 Main

54.4%
Agentic coding

CursorBench 4.0

57.8%
Full official benchmark table9 scores and peer comparisons
BenchmarkAreaScoreComparison
Terminal-Bench 4.0Agentic coding66.4%
Claude Opus 5.566.4%
GPT-6 Astra57.9%
Claude Fable 5.155.8%
Claude Opus 552.3%
FrontierCode v1.1 MainAgentic coding54.4%
Claude Opus 5.554.4%
GPT-6 Astra53.3%
Claude Fable 5.150.3%
Claude Opus 548.0%
CursorBench 4.0Agentic coding57.8%
Claude Opus 5.557.8%
Claude Fable 5.151.8%
Claude Opus 546.6%
GDPval-AA v2.1Knowledge work1846
AutomationBenchBusiness workflows40.0%
Claude Opus 5.540.0%
GPT-6 Astra41.4%
Claude Fable 5.131.4%
Claude Opus 526.9%
Humanity's Last ExamReasoning67.7%
Claude Opus 5.567.7%
Claude Fable 5.165.6%
Claude Opus 563.6%
GPT-6 Astra57.2%
Terminal-Bench-Science 0.1Scientific research58.7%
Claude Opus 5.558.7%
GPT-6 Astra64.6%
Claude Fable 5.152.6%
Claude Opus 529.0%
OSWorld 2.0Computer use81.8%
Claude Opus 5.581.8%
Claude Fable 5.180.7%
Claude Opus 574.0%
ChartographyVisual reasoning89.0%
Claude Opus 5.589.0%
Claude Fable 5.188.4%
Claude Opus 583.4%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-23 ↗
Status
Current
API model ID
claude-opus-5-5
Context
1M
Max output
128K
License
Not publicly verified
Access route
Claude API and Claude Code; also available through Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. The September 23 visual runs selected Opus 5.5 in Cursor; listed API prices are not measured run costs.
Modalities
text input, image input, text output
Hardware
Not publicly verified
Claude API Standard · USD / 1M tokens
$4 input · $0.2 cached input · $5 cache write · $20 output

Cache write shown is the 5-minute rate; 1-hour cache writes cost $8 per 1M tokens. Batch input/output is half price. Claude API Fast mode is $8 input / $40 output. Regional and tool rates may differ.

Benchmark sources

Anthropic reports Opus 5.5 at maximum adaptive-thinking effort unless noted. These provider scores are separate from the ten local prompt-only visual runs and their validation records.

Introducing Claude Opus 5.5

GPT-6 Astra figures in Anthropic’s table come from OpenAI. Terminal-Bench 4.0 uses Opus 5.5 at xhigh effort and Astra at high effort. Safeguards were enabled for Opus 5.5, with fallback models on restricted tasks. Missing GPT-6 Sol and DeepSeek V4.1 Flash rows are not zero scores.

Evaluation settings

  • Opus 5.5 xhigh effort; GPT-6 Astra high effort as reported by OpenAI. Claude Code harness; ±2.6 points standard error for Opus 5.5.
  • Anthropic launch table, maximum adaptive-thinking effort; evaluates whether agent code changes would be merged.
  • Anthropic launch table, maximum adaptive-thinking effort.
  • Artificial Analysis evaluation reported in Anthropic’s launch table; Elo score at maximum effort.
  • Zapier evaluation during early access; safeguards intervening count as failures. Peer figures from Zapier’s public leaderboard.
  • With tools; Anthropic launch table at maximum effort.
  • Anthropic launch table; ±3.5–5 points standard error per model. Astra figure as reported by OpenAI.
  • Partial score in Anthropic launch table; maximum effort.

Local visual-run methodology and validation records