Anthropic
Claude Opus 5
Version 5 benchmark runs across the shared Superbash visual prompts.
API pricing
USD per million tokens · Claude API Standard- Input
- $5
- Output
- $25
- Cached input
- $0.5
- Context window
- 1M
Inspect the original visual runs
Open the generated scene to test it, or compare that prompt with another model.

Helm's Deep
Run date: 2026-07-27
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Run date: 2026-07-27
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Mechanical Watch Simulator
Run date: 2026-07-27
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator
Run date: 2026-07-27
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly
Run date: 2026-07-27
Specialist test · evaluated separately
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Retired tests · 6 archived runs
These tests are no longer in the active suite. Their original results remain available for inspection.
How Claude Opus 5 compares
Three relevant peers, side by side.
- Claude Opus 5 This model
- Claude Fable 5.1
- GPT-6 Astra
- GPT-5.6 Sol
6 benchmarks · 4 models
Official benchmark profile
Published scores outside our visual tests
SWE-bench Pro
79.2%OSWorld 2.0
70.57%BrowseComp
90.8%Full official benchmark table12 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| SWE-bench Pro | Coding | 79.2% | |
| OSWorld 2.0 | Computer use | 70.57% | |
| BrowseComp | Web research | 90.8% | |
| SWE-bench Multilingual | Multilingual coding | 89.5% | |
| SWE-bench Multimodal | Visual coding | 59.4% | |
| DeepSWE v1.1 | Long-horizon coding | 68.8% | |
| FrontierCode 1.1 Main | Agentic coding | 53.4% | |
| FrontierBench v0.1 | Terminal work | 44.4% | |
| Humanity's Last Exam, no tools | Expert knowledge | 56.3% | |
| Humanity's Last Exam, with tools | Tool-assisted knowledge | 64.7% | |
| GDPval-AA v2 | Professional work | 1861 Elo | |
| AutomationBench | Business automation | 26.0% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Superseded in this ranking
- API model ID
claude-opus-5- Context
- 1M
- Max output
- 128K
- License
- Not publicly verified
- Access route
- Not publicly verified
- Modalities
- Not publicly verified
- Hardware
- Not publicly verified
- Claude API Standard · USD / 1M tokens
- $5 input · $0.5 cached input · $6.25 cache write · $25 output
Cache write shown is the 5-minute rate; 1-hour cache writes cost $10 per 1M tokens. Standard global routing; Batch, Fast mode where available, and regional pricing differ.
Benchmark sources
Anthropic reports gains over Opus 4.8 in coding, computer use, web research, and professional work. Each row retains the published effort level, harness, and trial count where the system card provides them.
Claude Opus 5 System CardEvaluation settings
- Average of five trials in the standard max-effort configuration.
- First-attempt success over five runs on 1080p Ubuntu with a 500-action cap.
- Search, fetch, programmatic tools and code with a 10M-token budget and compaction.
- 300 problems in nine languages; average of five trials.
- Visual issue context in Anthropic's internal harness; average of five trials.
- 113 tasks; average of five trials.
- Best result at medium effort; Cognition mean@5 run and scoring.
- Anthropic internal run at xhigh effort over 74 tasks with mini-SWE-agent on GKE.
- Adaptive thinking, up to 1M tokens, and no compaction.
- Search, fetch, programmatic tools and code with up to 1M tokens and contamination checks.
- Independent Artificial Analysis run at max effort over 220 tasks in 44 occupations.
- Max effort on a private held-out set with deterministic workflow completion.
- BrowseComp
- Anthropic does not publish the exact Opus effort label for this row.
- FrontierBench v0.1
- Harbor's separate summary-table run reports 43.3%; this row uses Anthropic's comparable internal run.
Comparison methodology
Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.
- Terminal-Bench Science 0.1
- Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- OSWorld 2.0 · offline, v2026.08.08
- Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
- ScreenSpot-Pro (no tools)
- The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
- HealthBench Professional (length-adjusted)
- OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
- ExploitBench
- Astra and Sol evaluated without production safeguards.
- ExploitGym
- The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
- ExploitBench (June–August 2026)
- OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
- SRE-Bench (single attempt)
- Cyber capability evaluation without production safeguards.
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
- ARC-AGI-3
- OpenAI Responses harness with two modified settings; see source footnote 1.