Anthropic

Tier B · Specialist or second-choice models.

Claude Opus 5

Still capable, but current instruction-following concerns make it a reviewed second choice.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
claude-opus-5
Context
1M
Max output
128K
API price / 1M tokens
$5 input · $25 output

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

11 runs

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Superbash commentary

Our take

Back to the full tier list →

Claude Opus 5 sits in B tier. The team expected it to be a game changer, but current power-user experience has been less dependable than GPT-5.6 for following instructions and delivering the requested shape of work.

Best for

  • Reviewed Claude workflows
  • Continuation from a strong plan
  • Tasks where Anthropic tooling is already established

Watch out

  • Instruction following can feel inconsistent
  • Power-user experience may differ from launch expectations
  • Avoid a yearly lock-in while the model race is changing this quickly

Why it is ranked here

  1. B rather than A reflects the gap between its raw capability and the team’s current delivery experience.
  2. The team’s main concern is practical instruction-following and delivery, not a claim that the model lacks raw capability.
  3. Speculation about why behavior changed is not treated as evidence; the placement reflects the observed workflow experience only.

Evidence and commentary

2026-08-17

Superbash editorial model ranking

Takeaway: Claude Opus 5 is currently placed in Tier B.

The August 2026 editorial roster places Claude Opus 5 at rank 11.

Open source →

Official benchmark profile

How Claude Opus 5 scores beyond our visual tests.

Anthropic reports gains over Opus 4.8 in coding, computer use, web research, and professional work. Each row retains the published effort level, harness, and trial count where the system card provides them.

Anthropic sourceJuly 2026Source report →
Coding

SWE-bench Pro

79.2%
Computer use

OSWorld 2.0

70.57%
Web research

BrowseComp

90.8%
Full official benchmark table12 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
SWE-bench ProCoding79.2%Average of five trials in the standard max-effort configuration.
Claude Fable 580.0%
Claude Opus 579.2%
Claude Opus 4.869.2%
GPT-5.6 Sol64.6%
OSWorld 2.0Computer use70.57%First-attempt success over five runs on 1080p Ubuntu with a 500-action cap.
Claude Opus 570.57%
Claude Fable 566.1%
GPT-5.6 Sol62.6%
Claude Opus 4.855.7%
BrowseCompWeb research90.8%Search, fetch, programmatic tools and code with a 10M-token budget and compaction.Anthropic does not publish the exact Opus effort label for this row.
Claude Opus 590.8%
GPT-5.6 Sol90.4%
Claude Fable 587.4%
Claude Opus 4.884.3%
SWE-bench MultilingualMultilingual coding89.5%300 problems in nine languages; average of five trials.
Claude Opus 589.5%
Claude Fable 586.6%
Claude Opus 4.884.4%
SWE-bench MultimodalVisual coding59.4%Visual issue context in Anthropic's internal harness; average of five trials.
Claude Opus 559.4%
Claude Fable 554.1%
Claude Opus 4.838.4%
DeepSWE v1.1Long-horizon coding68.8%113 tasks; average of five trials.
GPT-5.6 Sol72.7%
Claude Fable 569.7%
Claude Opus 568.8%
Claude Opus 4.859.0%
FrontierCode 1.1 MainAgentic coding53.4%Best result at medium effort; Cognition mean@5 run and scoring.
Claude Fable 553.5%
Claude Opus 553.4%
GPT-5.6 Sol47.5%
Claude Opus 4.846.5%
FrontierBench v0.1Terminal work44.4%Anthropic internal run at xhigh effort over 74 tasks with mini-SWE-agent on GKE.Harbor's separate summary-table run reports 43.3%; this row uses Anthropic's comparable internal run.
Claude Opus 544.4%
GPT-5.6 Sol max37.5%
Claude Fable 5 max33.7%
Claude Opus 4.818.7%
Humanity's Last Exam, no toolsExpert knowledge56.3%Adaptive thinking, up to 1M tokens, and no compaction.
Claude Fable 556.5%
Claude Opus 556.3%
Claude Opus 4.849.8%
Humanity's Last Exam, with toolsTool-assisted knowledge64.7%Search, fetch, programmatic tools and code with up to 1M tokens and contamination checks.
Claude Opus 564.7%
Claude Fable 563.9%
Claude Opus 4.857.9%
GDPval-AA v2Professional work1861 EloIndependent Artificial Analysis run at max effort over 220 tasks in 44 occupations.
AutomationBenchBusiness automation26.0%Max effort on a private held-out set with deterministic workflow completion.
Claude Opus 526.0%
GPT-5.6 Sol18.1%
Claude Fable 517.4%
Claude Opus 4.817.0%