Compare benchmark scores.

Choose two models and a workload to see where each leads.

September 2026 comparison

Provider-published results · checked

Read OpenAI’s source table

6 selected benchmarks · GPT-6 Astra vs Claude Fable 5.1 · pp = percentage points

Terminal-Bench 4.0

Coding · Higher is better
Astra57.9%
Fable 5.155.8%
2.1 ppAstra ahead

Terminal-Bench Science 0.1

Science · Higher is better
Astra64.6%
Fable 5.152.6%
12.0 ppAstra ahead

AutomationBench

Automation · Higher is better
Astra41.4%
Fable 5.131.4%
10.0 ppAstra ahead

Humanity’s Last Exam (with tools)

Reasoning · Higher is better
Astra57.2%
Fable 5.165.0%
7.8 ppFable 5.1 ahead

Artificial Analysis Intelligence Index v4.1.1

Reasoning · Higher is better · index points
Astra61.2
Fable 5.165.7
4.5 pointsFable 5.1 ahead

FrontierMath Tier 4 (v2)

Reasoning · Higher is better
Astra97.6%
Fable 5.187.8%
9.8 ppAstra ahead
All 40 benchmarks · 3 models View source snapshot
OpenAI launch table, checked 7 September 2026. — means not reported. * indicates a methodology note. Higher is better unless marked ↓.
BenchmarkAstraFable 5.1Gemini 3.8 Flash
Terminal-Bench 4.057.9%55.8%19.1%
Terminal-Bench Science 0.1 *64.6%52.6%—
AutomationBench41.4%31.4%—
Humanity’s Last Exam (with tools)57.2%65.0%—
Artificial Analysis Intelligence Index v4.1.1 *61.265.758.7
FrontierMath Tier 4 (v2)97.6%87.8%—
DeepSWE v1.174.1%67.4%73.8%
FrontierCode 1.1 Extended (score) *64.5%63.6%56.3%
FrontierCode 1.1 Main (score) *53.3%50.9%43.6%
GPQA Diamond96.0%93.7%95.3%
ARC-AGI-295.0%90.0%—
ARC-AGI-198.5%97.5%—
Agents’ Last Exam59.3%——
OSWorld 2.0 · offline, v2026.08.08 *72.6%——
ScreenSpot-Pro (no tools) *92.7%——
BenchCAD (with tools) *95.9%84.3%—
BrowseComp91.5%——
OpenScore String Quartets (1 − OMR-NED)0.84——
Design tasks (OpenAI internal)50.0%——
Data science tasks (OpenAI internal)40.9%——
Database migration (OpenAI internal)63.9%57.8%—
Artificial Analysis Coding Agent Index v1.467.0—61.2
GeneBench Pro37.1%——
MedChemBench (OpenAI internal)49.3%——
LifeSciBench60.3%——
HealthBench Professional (length-adjusted) *63.4%58.1%52.1%
ExploitBench *100.0%——
ExploitGym *42.4%——
ExploitBench (June–August 2026) *39.0%——
SRE-Bench (single attempt) *88.0%——
SEC-Bench Pro85.4%——
Computer-use safety (OpenAI internal) ↓ *2.4%9.5%—
Computer-use safety with AutoReview ↓1.8%——
Circumvention (OpenAI internal) ↓0.0%——
ExploitGym honeypot ↓0.0%——
Impossible ExploitGym100.0%——
Hallucination (OpenAI internal) ↓4.2%——
MRCR v2 · 8 needles · 256K–512K100.0%——
MRCR v2 · 8 needles · 512K–1M96.3%——
ARC-AGI-3 *99.9%——
Sources and methodology

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.

Gaps are differences in reported scores, not statistical significance. Percentage bars use 0–100; index bars use a 0–100 display scale.

Anthropic’s model report
Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.
3models on shared source tables
5models on the same evaluation
2unranked for now

The September launch snapshot above and earlier independent measurements below retain their own sources and versions.

Independent comparison

Earlier independent scores across 1 models.

This retained Artificial Analysis Index v4.1 snapshot combines nine evaluations across coding, agentic work, science, knowledge, physics, and long-context reasoning. Higher is better. Astra and Fable 5.1’s v4.1.1 figures appear above; they are not mixed into this older index.

Top index57Kimi K3
Fastest output32Kimi K3, tokens per second
Lowest blended price$2.31Kimi K3, per 1M tokens
  1. Kimi K3Reasoning
    57AA source ↗

Configuration matters. This earlier index retains each model’s published effort setting and source.

Cost explorer

Estimate your token costs.

Set total input and output tokens for a workload. This estimate uses the displayed short-context API rates, including active promotions, across all requests.

Example workloads

Opus 5.5 is $4 input / $20 output; Astra and Fable 5.1 are $10 / $50; Sol is $2 / $10. Astra and Sol requests above 272K input tokens use their higher rates for the full request. Grok 4.7 requests at or above 200K prompt tokens use $4 / $12. DeepSeek uses peak rates here; off-peak rates are half. Cache reads, cache writes, tools, regional premiums, and Fast mode are excluded.

ModelAA IndexInput / 1MOutput / 1MSpeedYour task
Kimi K3Moonshot AI · Reasoning57Index v4.1$3.00$15.0032.0tokens/sec$1.80AA data ↗Price ↗
Claude Opus 5.5Anthropic · Standard API · pricing checked 2026-09-23—See v4.1.1 above$4.00$20.00Not measuredNo speed estimate$2.40Price ↗
GPT-6 AstraOpenAI · Standard API · pricing checked 2026-09-08—See v4.1.1 above$10.00$50.00Not measuredNo speed estimate$6.00Price ↗
GPT-6 SolOpenAI · Standard API · pricing checked 2026-09-23—See v4.1.1 above$2.00$10.00Not measuredNo speed estimate$1.20Price ↗
DeepSeek V4.1 FlashDeepSeek · Standard API · pricing checked 2026-09-21—See v4.1.1 above$0.30$1.20Not measuredNo speed estimate$0.16Price ↗
Claude Fable 5.1Anthropic · Standard API · pricing checked 2026-09-08—See v4.1.1 above$10.00$50.00Not measuredNo speed estimate$6.00Price ↗
Grok 4.7SpaceXAI · Standard API · pricing checked 2026-09-22—See v4.1.1 above$2.00$6.00Not measuredNo speed estimate$0.88Price ↗

What is included: uncached input and generated output at short-context API rates. The baseline full-catalog audit was 2026-09-08; newer additions retain their own verification date in the catalog.

What is not included: long-context premiums, cache writes, tool calls, storage, batch discounts, and provider markups. A long prompt may cost more than this estimate; check the linked pricing source.

Blended reference: Calculated from the current catalog using a 7:2:1 cache-hit/input/output mix. The lowest in this set is Kimi K3 at $2.31 per 1M blended tokens.

Comparable score lanes

Keep incompatible scores separate.

Each plot covers one benchmark. Orange marks one source table. Outlined marks use the same named evaluation from another publisher and are not assumed to be identical runs.

Agentic coding

SWE-bench Pro

GPT-5.6 comparison table ↗

Public tasks. Scores below are published in cross-model tables; effort and harness are retained in each source.

  1. Inkling
    54.3%Same evaluation ↗

Published July 2026. Open the source before using these figures for a purchasing decision.

Terminal agents

Terminal-Bench 2.1

GPT-5.6 comparison table ↗

Command-line task completion. The primary figures use OpenAI’s shared table; the other reported scores are labelled separately rather than normalized.

  1. Inkling
    63.8%Same evaluation ↗

Published July 2026. Open the source before using these figures for a purchasing decision.

Science reasoning

GPQA Diamond

Inkling model card cross-model table ↗

No-tools science questions. This lane keeps the reported configuration in view because reasoning budgets differ by provider.

  1. Inkling
    87.2%Shared table

Published July 2026. Open the source before using these figures for a purchasing decision.

Methodology

Limits of this comparison

Artificial Analysis Index results come from one independent harness. The exact model configuration remains attached to every score.

Token cost scenarios multiply published input and output prices by the workload you enter. They exclude tools, storage, cache writes, and provider discounts.

A shared-table mark means the source published the compared models in one table. A same-eval mark means the benchmark name matches, but the source or configuration differs.

Missing data is not a low score. Models without a compatible public score remain unranked until a reproducible result is published.

Our Superbash visual runs are a separate, same-prompt evidence layer. Use the live run comparison to inspect build quality directly.

Inspect the same-prompt visual runs →