Anthropic
Claude Opus 5.5
A leading choice for long coding and research jobs, with strong published results and ten inspectable visual runs.
API pricing
USD per million tokens · Claude API Standard- Input
- $4
- Output
- $20
- Cached input
- $0.2
- Context window
- 1M
Try the active visual tests
Open the generated scene to test it, or compare that prompt with another model.

City Scroll Journey
Run date: 2026-09-23
Specialist test · evaluated separately
Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Ember Glider
Run date: 2026-09-23
Sunset gliding journey testing flight energy management, checkpoint flow, and atmospheric scene craft.

Helm's Deep
Run date: 2026-09-23
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Run date: 2026-09-23
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Low-Poly Tower Defense
Run date: 2026-09-23
Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Mechanical Watch Simulator
Run date: 2026-09-23
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Neon Drift
Run date: 2026-09-23
Synthwave time-trial racing testing drift physics, lap timing, ghost replay, and unlock progression.

Starfall Arena
Run date: 2026-09-23
Neon arena survival testing wave escalation, upgrade builds, particle feedback, and boss design.

Stormwind Trebuchet Simulator
Run date: 2026-09-23
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly
Run date: 2026-09-23
Specialist test · evaluated separately
Timber assembly scene testing structure, joinery, construction order, and material clarity.
How Claude Opus 5.5 compares
Three relevant peers, side by side.
- Claude Opus 5.5 This model
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
6 benchmarks · 4 models
Official benchmark profile
Anthropic’s published coding and knowledge-work results
Terminal-Bench 4.0
66.4%FrontierCode v1.1 Main
54.4%CursorBench 4.0
57.8%Full official benchmark table9 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| Terminal-Bench 4.0 | Agentic coding | 66.4% | |
| FrontierCode v1.1 Main | Agentic coding | 54.4% | |
| CursorBench 4.0 | Agentic coding | 57.8% | |
| GDPval-AA v2.1 | Knowledge work | 1846 | |
| AutomationBench | Business workflows | 40.0% | |
| Humanity's Last Exam | Reasoning | 67.7% | |
| Terminal-Bench-Science 0.1 | Scientific research | 58.7% | |
| OSWorld 2.0 | Computer use | 81.8% | |
| Chartography | Visual reasoning | 89.0% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Current
- API model ID
claude-opus-5-5- Context
- 1M
- Max output
- 128K
- License
- Not publicly verified
- Access route
- Claude API and Claude Code; also available through Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. The September 23 visual runs selected Opus 5.5 in Cursor; listed API prices are not measured run costs.
- Modalities
- text input, image input, text output
- Hardware
- Not publicly verified
- Claude API Standard · USD / 1M tokens
- $4 input · $0.2 cached input · $5 cache write · $20 output
Cache write shown is the 5-minute rate; 1-hour cache writes cost $8 per 1M tokens. Batch input/output is half price. Claude API Fast mode is $8 input / $40 output. Regional and tool rates may differ.
Benchmark sources
Anthropic reports Opus 5.5 at maximum adaptive-thinking effort unless noted. These provider scores are separate from the ten local prompt-only visual runs and their validation records.
Introducing Claude Opus 5.5GPT-6 Astra figures in Anthropic’s table come from OpenAI. Terminal-Bench 4.0 uses Opus 5.5 at xhigh effort and Astra at high effort. Safeguards were enabled for Opus 5.5, with fallback models on restricted tasks. Missing GPT-6 Sol and DeepSeek V4.1 Flash rows are not zero scores.
Evaluation settings
- Opus 5.5 xhigh effort; GPT-6 Astra high effort as reported by OpenAI. Claude Code harness; ±2.6 points standard error for Opus 5.5.
- Anthropic launch table, maximum adaptive-thinking effort; evaluates whether agent code changes would be merged.
- Anthropic launch table, maximum adaptive-thinking effort.
- Artificial Analysis evaluation reported in Anthropic’s launch table; Elo score at maximum effort.
- Zapier evaluation during early access; safeguards intervening count as failures. Peer figures from Zapier’s public leaderboard.
- With tools; Anthropic launch table at maximum effort.
- Anthropic launch table; ±3.5–5 points standard error per model. Astra figure as reported by OpenAI.
- Partial score in Anthropic launch table; maximum effort.
Editorial verdict
Where this model fits
Claude Opus 5.5 is Anthropic’s September 2026 model for long-running agentic coding and knowledge work. Its Claude API rate is $4 input and $20 output per million tokens, with a one-million-token context window. Anthropic reports 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 Main at maximum effort. Ten same-prompt Superbash visual artifacts are available for inspection; eight were time-capped during validation. The S-tier placement is editorial guidance, not a normalized cross-model score.
Best for
Watch out
Why it is ranked here
Evidence
Superbash editorial model ranking
Takeaway: Claude Opus 5.5 is currently placed in Tier S.
The September 2026 editorial roster places Claude Opus 5.5 at rank 2.
Open source →Claude Opus 5.5 launch and visual-run methodology
Takeaway: Anthropic publishes model specifications and benchmark scores; ten separate Superbash visual runs are linked from this model page.
Provider results and local same-prompt artifacts are separate evidence. The local methodology records the validation time cap.
Open source →Related guides