Anthropic
Claude Opus 5
Still capable, but current instruction-following concerns make it a reviewed second choice.
API pricing
USD per million tokens · Claude API Standard- Input
- $5
- Output
- $25
- Cached input
- $0.5
- Context window
- 1M
Try the visual runs
Open the generated scene to test it, or compare that prompt with another model.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
How Claude Opus 5 compares
Three relevant peers, side by side.
- Claude Opus 5 This model
- Claude Fable 5.1
- GPT-6 Astra
- GPT-5.6 Sol
6 benchmarks · 4 models
Official benchmark profile
Published scores outside our visual tests
SWE-bench Pro
79.2%OSWorld 2.0
70.57%BrowseComp
90.8%Full official benchmark table12 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| SWE-bench Pro | Coding | 79.2% | |
| OSWorld 2.0 | Computer use | 70.57% | |
| BrowseComp | Web research | 90.8% | |
| SWE-bench Multilingual | Multilingual coding | 89.5% | |
| SWE-bench Multimodal | Visual coding | 59.4% | |
| DeepSWE v1.1 | Long-horizon coding | 68.8% | |
| FrontierCode 1.1 Main | Agentic coding | 53.4% | |
| FrontierBench v0.1 | Terminal work | 44.4% | |
| Humanity's Last Exam, no tools | Expert knowledge | 56.3% | |
| Humanity's Last Exam, with tools | Tool-assisted knowledge | 64.7% | |
| GDPval-AA v2 | Professional work | 1861 Elo | |
| AutomationBench | Business automation | 26.0% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Current
- API model ID
claude-opus-5- Context
- 1M
- Max output
- 128K
- License
- Not publicly verified
- Access route
- Not publicly verified
- Modalities
- Not publicly verified
- Hardware
- Not publicly verified
- Claude API Standard · USD / 1M tokens
- $5 input · $0.5 cached input · $6.25 cache write · $25 output
Cache write shown is the 5-minute rate; 1-hour cache writes cost $10 per 1M tokens. Standard global routing; Batch, Fast mode where available, and regional pricing differ.
Benchmark sources
Anthropic reports gains over Opus 4.8 in coding, computer use, web research, and professional work. Each row retains the published effort level, harness, and trial count where the system card provides them.
Claude Opus 5 System CardEvaluation settings
- Average of five trials in the standard max-effort configuration.
- First-attempt success over five runs on 1080p Ubuntu with a 500-action cap.
- Search, fetch, programmatic tools and code with a 10M-token budget and compaction.
- 300 problems in nine languages; average of five trials.
- Visual issue context in Anthropic's internal harness; average of five trials.
- 113 tasks; average of five trials.
- Best result at medium effort; Cognition mean@5 run and scoring.
- Anthropic internal run at xhigh effort over 74 tasks with mini-SWE-agent on GKE.
- Adaptive thinking, up to 1M tokens, and no compaction.
- Search, fetch, programmatic tools and code with up to 1M tokens and contamination checks.
- Independent Artificial Analysis run at max effort over 220 tasks in 44 occupations.
- Max effort on a private held-out set with deterministic workflow completion.
- BrowseComp
- Anthropic does not publish the exact Opus effort label for this row.
- FrontierBench v0.1
- Harbor's separate summary-table run reports 43.3%; this row uses Anthropic's comparable internal run.
Comparison methodology
Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.
- Terminal-Bench Science 0.1
- Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- OSWorld 2.0 · offline, v2026.08.08
- Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
- ScreenSpot-Pro (no tools)
- The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
- HealthBench Professional (length-adjusted)
- OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
- ExploitBench
- Astra and Sol evaluated without production safeguards.
- ExploitGym
- The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
- ExploitBench (June–August 2026)
- OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
- SRE-Bench (single attempt)
- Cyber capability evaluation without production safeguards.
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
- ARC-AGI-3
- OpenAI Responses harness with two modified settings; see source footnote 1.
Editorial verdict
Where this model fits
Claude Opus 5 sits in B tier. The team expected it to be a game changer, but current power-user experience has been less dependable than GPT-5.6 for following instructions and delivering the requested shape of work.
Best for
Watch out
Why it is ranked here
Evidence
Superbash editorial model ranking
Takeaway: Claude Opus 5 is currently placed in Tier B.
The September 2026 editorial roster places Claude Opus 5 at rank 12.
Open source →Related guides