Anthropic

Tier B · Specialist or second-choice models.

Claude Opus 5

Still capable, but current instruction-following concerns make it a reviewed second choice.

API pricing

USD per million tokens · Claude API Standard
Input
$5
Output
$25
Cached input
$0.5
Context window
1M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

11 runs

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

Claude Opus 5 sits in B tier. The team expected it to be a game changer, but current power-user experience has been less dependable than GPT-5.6 for following instructions and delivering the requested shape of work.

Best for

  • Reviewed Claude workflows
  • Continuation from a strong plan
  • Tasks where Anthropic tooling is already established

Watch out

  • Instruction following can feel inconsistent
  • Power-user experience may differ from launch expectations
  • Avoid a yearly lock-in while the model race is changing this quickly

Why it is ranked here

  1. B rather than A reflects the gap between its raw capability and the team’s current delivery experience.
  2. The team’s main concern is practical instruction-following and delivery, not a claim that the model lacks raw capability.
  3. Speculation about why behavior changed is not treated as evidence; the placement reflects the observed workflow experience only.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: Claude Opus 5 is currently placed in Tier B.

The September 2026 editorial roster places Claude Opus 5 at rank 12.

Open source →

How Claude Opus 5 compares

Three relevant peers, side by side.

  • Claude Opus 5 This model
  • Claude Fable 5.1
  • GPT-6 Astra
  • GPT-5.6 Sol

6 benchmarks · 4 models

Terminal-Bench 4.0

Coding · Higher is better

Source
Claude Opus 552.6%
Claude Fable 5.155.8%
GPT-6 Astra57.9%
GPT-5.6 Sol37.3%

Agents’ Last Exam

Computer use · Higher is better

Source
Claude Opus 555.5%
Claude Fable 5.1Not reported
GPT-6 Astra59.3%
GPT-5.6 Sol53.6%

DeepSWE v1.1

Coding · Higher is better

Source
Claude Opus 573.7%
Claude Fable 5.167.4%
GPT-6 Astra74.1%
GPT-5.6 Sol72.7%

AutomationBench

Automation · Higher is better

Source
Claude Opus 526.9%
Claude Fable 5.131.4%
GPT-6 Astra41.4%
GPT-5.6 Sol18.1%

BrowseComp

Research · Higher is better

Source
Claude Opus 590.8%
Claude Fable 5.1Not reported
GPT-6 Astra91.5%
GPT-5.6 Sol90.4%

Humanity’s Last Exam (with tools)

Reasoning · Higher is better

Source
Claude Opus 563.6%
Claude Fable 5.165.0%
GPT-6 Astra57.2%
GPT-5.6 SolNot reported

Official benchmark profile

Published scores outside our visual tests

Anthropic sourceJuly 2026Source report →
Coding

SWE-bench Pro

79.2%
Computer use

OSWorld 2.0

70.57%
Web research

BrowseComp

90.8%
Full official benchmark table12 scores and peer comparisons
BenchmarkAreaScoreComparison
SWE-bench ProCoding79.2%
Claude Opus 579.2%
Claude Fable 580.0%
Claude Opus 4.869.2%
GPT-5.6 Sol64.6%
OSWorld 2.0Computer use70.57%
Claude Opus 570.57%
Claude Fable 566.1%
GPT-5.6 Sol62.6%
Claude Opus 4.855.7%
BrowseCompWeb research90.8%
Claude Opus 590.8%
GPT-5.6 Sol90.4%
Claude Fable 587.4%
Claude Opus 4.884.3%
SWE-bench MultilingualMultilingual coding89.5%
Claude Opus 589.5%
Claude Fable 586.6%
Claude Opus 4.884.4%
SWE-bench MultimodalVisual coding59.4%
Claude Opus 559.4%
Claude Fable 554.1%
Claude Opus 4.838.4%
DeepSWE v1.1Long-horizon coding68.8%
Claude Opus 568.8%
Claude Fable 569.7%
GPT-5.6 Sol72.7%
Claude Opus 4.859.0%
FrontierCode 1.1 MainAgentic coding53.4%
Claude Opus 553.4%
Claude Fable 553.5%
GPT-5.6 Sol47.5%
Claude Opus 4.846.5%
FrontierBench v0.1Terminal work44.4%
Claude Opus 544.4%
GPT-5.6 Sol max37.5%
Claude Fable 5 max33.7%
Claude Opus 4.818.7%
Humanity's Last Exam, no toolsExpert knowledge56.3%
Claude Opus 556.3%
Claude Fable 556.5%
Claude Opus 4.849.8%
Humanity's Last Exam, with toolsTool-assisted knowledge64.7%
Claude Opus 564.7%
Claude Fable 563.9%
Claude Opus 4.857.9%
GDPval-AA v2Professional work1861 Elo
AutomationBenchBusiness automation26.0%
Claude Opus 526.0%
GPT-5.6 Sol18.1%
Claude Fable 517.4%
Claude Opus 4.817.0%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
claude-opus-5
Context
1M
Max output
128K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
Claude API Standard · USD / 1M tokens
$5 input · $0.5 cached input · $6.25 cache write · $25 output

Cache write shown is the 5-minute rate; 1-hour cache writes cost $10 per 1M tokens. Standard global routing; Batch, Fast mode where available, and regional pricing differ.

Benchmark sources

Anthropic reports gains over Opus 4.8 in coding, computer use, web research, and professional work. Each row retains the published effort level, harness, and trial count where the system card provides them.

Claude Opus 5 System Card

Evaluation settings

  • Average of five trials in the standard max-effort configuration.
  • First-attempt success over five runs on 1080p Ubuntu with a 500-action cap.
  • Search, fetch, programmatic tools and code with a 10M-token budget and compaction.
  • 300 problems in nine languages; average of five trials.
  • Visual issue context in Anthropic's internal harness; average of five trials.
  • 113 tasks; average of five trials.
  • Best result at medium effort; Cognition mean@5 run and scoring.
  • Anthropic internal run at xhigh effort over 74 tasks with mini-SWE-agent on GKE.
  • Adaptive thinking, up to 1M tokens, and no compaction.
  • Search, fetch, programmatic tools and code with up to 1M tokens and contamination checks.
  • Independent Artificial Analysis run at max effort over 220 tasks in 44 occupations.
  • Max effort on a private held-out set with deterministic workflow completion.
BrowseComp
Anthropic does not publish the exact Opus effort label for this row.
FrontierBench v0.1
Harbor's separate summary-table run reports 43.3%; this row uses Anthropic's comparable internal run.

Comparison methodology

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.

Anthropic model report

Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.