Anthropic
Claude Opus 5
Still capable, but current instruction-following concerns make it a reviewed second choice.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
claude-opus-5- Context
- 1M
- Max output
- 128K
- API price / 1M tokens
- $5 input · $25 output
Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How Claude Opus 5 scores beyond our visual tests.
Anthropic reports gains over Opus 4.8 in coding, computer use, web research, and professional work. Each row retains the published effort level, harness, and trial count where the system card provides them.
SWE-bench Pro
79.2%OSWorld 2.0
70.57%BrowseComp
90.8%Full official benchmark table12 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| SWE-bench Pro | Coding | 79.2% | Average of five trials in the standard max-effort configuration. |
| OSWorld 2.0 | Computer use | 70.57% | First-attempt success over five runs on 1080p Ubuntu with a 500-action cap. |
| BrowseComp | Web research | 90.8% | Search, fetch, programmatic tools and code with a 10M-token budget and compaction.Anthropic does not publish the exact Opus effort label for this row. |
| SWE-bench Multilingual | Multilingual coding | 89.5% | 300 problems in nine languages; average of five trials. |
| SWE-bench Multimodal | Visual coding | 59.4% | Visual issue context in Anthropic's internal harness; average of five trials. |
| DeepSWE v1.1 | Long-horizon coding | 68.8% | 113 tasks; average of five trials. |
| FrontierCode 1.1 Main | Agentic coding | 53.4% | Best result at medium effort; Cognition mean@5 run and scoring. |
| FrontierBench v0.1 | Terminal work | 44.4% | Anthropic internal run at xhigh effort over 74 tasks with mini-SWE-agent on GKE.Harbor's separate summary-table run reports 43.3%; this row uses Anthropic's comparable internal run. |
| Humanity's Last Exam, no tools | Expert knowledge | 56.3% | Adaptive thinking, up to 1M tokens, and no compaction. |
| Humanity's Last Exam, with tools | Tool-assisted knowledge | 64.7% | Search, fetch, programmatic tools and code with up to 1M tokens and contamination checks. |
| GDPval-AA v2 | Professional work | 1861 Elo | Independent Artificial Analysis run at max effort over 220 tasks in 44 occupations. |
| AutomationBench | Business automation | 26.0% | Max effort on a private held-out set with deterministic workflow completion. |
Superbash commentary
Our take
Claude Opus 5 sits in B tier. The team expected it to be a game changer, but current power-user experience has been less dependable than GPT-5.6 for following instructions and delivering the requested shape of work.
Best for
Watch out
Why it is ranked here
Evidence and commentary
Superbash editorial model ranking
Takeaway: Claude Opus 5 is currently placed in Tier B.
The August 2026 editorial roster places Claude Opus 5 at rank 11.
Open source →Keep learning