Compare benchmark scores.

Choose two models and a workload to see where each leads.

September 2026 comparison

Provider-published results · checked

Read OpenAI’s source table

6 selected benchmarks · GPT-6 Astra vs Claude Fable 5.1 · pp = percentage points

Terminal-Bench 4.0

Coding · Higher is better
Astra57.9%
Fable 5.155.8%
2.1 ppAstra ahead

Terminal-Bench Science 0.1

Science · Higher is better
Astra64.6%
Fable 5.152.6%
12.0 ppAstra ahead

AutomationBench

Automation · Higher is better
Astra41.4%
Fable 5.131.4%
10.0 ppAstra ahead

Humanity’s Last Exam (with tools)

Reasoning · Higher is better
Astra57.2%
Fable 5.165.0%
7.8 ppFable 5.1 ahead

Artificial Analysis Intelligence Index v4.1.1

Reasoning · Higher is better · index points
Astra61.2
Fable 5.165.7
4.5 pointsFable 5.1 ahead

FrontierMath Tier 4 (v2)

Reasoning · Higher is better
Astra97.6%
Fable 5.187.8%
9.8 ppAstra ahead
All 40 benchmarks · 6 models View source snapshot
OpenAI launch table, checked 7 September 2026. — means not reported. * indicates a methodology note. Higher is better unless marked ↓.
BenchmarkAstraSolFable 5.1Fable 5Opus 5Gemini 3.8 Flash
Terminal-Bench 4.057.9%37.3%55.8%44.5%52.6%19.1%
Terminal-Bench Science 0.1 *64.6%22.4%52.6%21.4%30.0%
AutomationBench41.4%18.1%31.4%17.4%26.9%
Humanity’s Last Exam (with tools)57.2%65.0%63.8%63.6%
Artificial Analysis Intelligence Index v4.1.1 *61.260.965.762.163.158.7
FrontierMath Tier 4 (v2)97.6%83.0%87.8%90.2%73.2%
DeepSWE v1.174.1%72.7%67.4%69.9%73.7%73.8%
FrontierCode 1.1 Extended (score) *64.5%60.6%63.6%64.9%63.6%56.3%
FrontierCode 1.1 Main (score) *53.3%47.5%50.9%53.5%53.4%43.6%
GPQA Diamond96.0%94.6%93.7%92.6%93.7%95.3%
ARC-AGI-295.0%92.5%90.0%89.2%90.4%
ARC-AGI-198.5%97.5%97.5%98.5%97.5%
Agents’ Last Exam59.3%53.6%48.7%55.5%
OSWorld 2.0 · offline, v2026.08.08 *72.6%65.7%70.2%
ScreenSpot-Pro (no tools) *92.7%76.9%
BenchCAD (with tools) *95.9%83.3%84.3%67.5%82.1%
BrowseComp91.5%90.4%87.4%90.8%
OpenScore String Quartets (1 − OMR-NED)0.840.19
Design tasks (OpenAI internal)50.0%47.4%35.8%
Data science tasks (OpenAI internal)40.9%30.5%34.7%
Database migration (OpenAI internal)63.9%42.7%57.8%50.3%
Artificial Analysis Coding Agent Index v1.467.065.167.268.161.2
GeneBench Pro37.1%32.3%
MedChemBench (OpenAI internal)49.3%47.4%
LifeSciBench60.3%59.9%
HealthBench Professional (length-adjusted) *63.4%60.5%58.1%60.9%56.4%52.1%
ExploitBench *100.0%78.5%70.0%
ExploitGym *42.4%30.3%22.0%
ExploitBench (June–August 2026) *39.0%5.5%
SRE-Bench (single attempt) *88.0%55.9%12.5%
SEC-Bench Pro85.4%79.1%
Computer-use safety (OpenAI internal) ↓ *2.4%22.0%9.5%18.3%11.5%
Computer-use safety with AutoReview ↓1.8%4.3%
Circumvention (OpenAI internal) ↓0.0%0.29%
ExploitGym honeypot ↓0.0%48.2%
Impossible ExploitGym100.0%
Hallucination (OpenAI internal) ↓4.2%12.2%
MRCR v2 · 8 needles · 256K–512K100.0%91.5%
MRCR v2 · 8 needles · 512K–1M96.3%73.8%
ARC-AGI-3 *99.9%7.8%30.2%
Sources and methodology

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.

Gaps are differences in reported scores, not statistical significance. Percentage bars use 0–100; index bars use a 0–100 display scale.

Anthropic’s model report
Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.
7models on shared source tables
6models on the same evaluation
1unranked for now

The September launch snapshot above and earlier independent measurements below retain their own sources and versions.

Independent comparison

Earlier independent scores across 9 models.

This retained Artificial Analysis Index v4.1 snapshot combines nine evaluations across coding, agentic work, science, knowledge, physics, and long-context reasoning. Higher is better. Astra and Fable 5.1’s v4.1.1 figures appear above; they are not mixed into this older index.

Top index61Grok 4.6
Fastest output178GPT-5.6 Luna, tokens per second
Lowest blended price$0.174GPT-5.6 Luna, per 1M tokens
  1. 61AA source ↗
  2. Claude Fable 5Max, Opus 4.8 fallback
    60AA source ↗
  3. Claude Opus 5Adaptive reasoning, xhigh
    60AA source ↗
  4. 59AA source ↗
  5. Kimi K3Reasoning
    57AA source ↗
  6. 55AA source ↗
  7. GPT-5.5xhigh
    55AA source ↗
  8. 54AA source ↗
  9. 46AA source ↗

Configuration matters. Fable 5 includes its published Opus 4.8 fallback. GPT and Claude scores use the effort level shown beside each model.

Cost explorer

Estimate your token costs.

Set total input and output tokens for a workload. This estimate uses the displayed short-context API rates, including active promotions, across all requests.

Example workloads

Astra and Fable 5.1 share $10 input / $50 output standard rates per million tokens. Astra requests above 272K input tokens use $20 input / $75 output for the full request; this calculator applies that threshold. Cache reads, cache writes, tools, and Fast mode are excluded.

ModelAA IndexInput / 1MOutput / 1MSpeedYour task
Claude Fable 5Anthropic · Max, Opus 4.8 fallback60Index v4.1$10.00$50.0069.7tokens/sec$6.00AA data ↗Price ↗
Claude Opus 5Anthropic · Adaptive reasoning, xhigh60Index v4.1$5.00$25.0053.1tokens/sec$3.00AA data ↗Price ↗
GPT-5.6 SolOpenAI · Max59Index v4.1$4.00$20.0077.1tokens/sec$2.40AA data ↗Price ↗
Kimi K3Moonshot AI · Reasoning57Index v4.1$3.00$15.0032.0tokens/sec$1.80AA data ↗Price ↗
GPT-5.6 TerraOpenAI · Max55Index v4.1$2.00$12.00134.5tokens/sec$1.36AA data ↗Price ↗
GPT-5.5OpenAI · xhigh55Index v4.1$5.00$30.0072.5tokens/sec$3.40AA data ↗Price ↗
Grok 4.5Xai · High54Index v4.1$2.00$6.0067.1tokens/sec$0.88AA data ↗Price ↗
Grok 4.6Xai · High61Index v4.1$2.00$6.0065.1tokens/sec$0.88AA data ↗Price ↗
GPT-5.6 LunaOpenAI · High46Index v4.1$0.20$1.20178.0tokens/sec$0.14AA data ↗Price ↗
GPT-6 AstraOpenAI · Standard API · pricing checked September 7, 2026See v4.1.1 above$10.00$50.00Not measuredNo speed estimate$6.00Price ↗
Claude Fable 5.1Anthropic · Standard API · pricing checked September 7, 2026See v4.1.1 above$10.00$50.00Not measuredNo speed estimate$6.00Price ↗

What is included: uncached input and generated output at short-context API rates verified 2026-09-08.

What is not included: long-context premiums, cache writes, tool calls, storage, batch discounts, and provider markups. A long prompt may cost more than this estimate; check the linked pricing source.

Blended reference: Calculated from the current catalog using a 7:2:1 cache-hit/input/output mix. The lowest in this set is GPT-5.6 Luna at $0.17 per 1M blended tokens.

Comparable score lanes

Keep incompatible scores separate.

Each plot covers one benchmark. Orange marks one source table. Outlined marks use the same named evaluation from another publisher and are not assumed to be identical runs.

Agentic coding

SWE-bench Pro

GPT-5.6 comparison table ↗

Public tasks. Scores below are published in cross-model tables; effort and harness are retained in each source.

  1. Claude Fable 5
    80.0%Shared table
  2. Grok 4.5
    64.7%Same evaluation ↗
  3. GPT-5.6 Sol
    64.6%Shared table
  4. GPT-5.6 Terra
    63.4%Shared table
  5. GPT-5.6 Luna
    62.7%Shared table
  6. GPT-5.5
    59.4%Shared table
  7. Inkling
    54.3%Same evaluation ↗

Published July 2026. Open the source before using these figures for a purchasing decision.

Terminal agents

Terminal-Bench 2.1

GPT-5.6 comparison table ↗

Command-line task completion. The primary figures use OpenAI’s shared table; the other reported scores are labelled separately rather than normalized.

  1. GPT-5.6 Sol
    88.8%Shared table
  2. GPT-5.6 Terra
    87.4%Shared table
  3. GPT-5.5
    85.6%Shared table
  4. GPT-5.6 Luna
    84.7%Shared table
  5. Grok 4.5
    83.3%Same evaluation ↗
  6. Claude Fable 5
    83.1%Shared table
  7. Inkling
    63.8%Same evaluation ↗

Published July 2026. Open the source before using these figures for a purchasing decision.

Science reasoning

GPQA Diamond

Inkling model card cross-model table ↗

No-tools science questions. This lane keeps the reported configuration in view because reasoning budgets differ by provider.

  1. GPT-5.6 Sol
    94.1%Shared table
  2. Claude Fable 5
    92.6%Shared table
  3. Inkling
    87.2%Shared table

Published July 2026. Open the source before using these figures for a purchasing decision.

Methodology

Limits of this comparison

Artificial Analysis Index results come from one independent harness. The exact model configuration remains attached to every score.

Token cost scenarios multiply published input and output prices by the workload you enter. They exclude tools, storage, cache writes, and provider discounts.

A shared-table mark means the source published the compared models in one table. A same-eval mark means the benchmark name matches, but the source or configuration differs.

Missing data is not a low score. Models without a compatible public score remain unranked until a reproducible result is published.

Our Superbash visual runs are a separate, same-prompt evidence layer. Use the live run comparison to inspect build quality directly.

Inspect the same-prompt visual runs →