Compare benchmark scores.
Choose two models and a workload to see where each leads.
Provider-published results · checked
Read OpenAI’s source table6 selected benchmarks · GPT-6 Astra vs Claude Fable 5.1 · pp = percentage points
Terminal-Bench 4.0
Coding · Higher is betterTerminal-Bench Science 0.1
Science · Higher is betterAutomationBench
Automation · Higher is betterHumanity’s Last Exam (with tools)
Reasoning · Higher is betterArtificial Analysis Intelligence Index v4.1.1
Reasoning · Higher is better · index pointsFrontierMath Tier 4 (v2)
Reasoning · Higher is betterDeepSWE v1.1
Coding · Higher is betterFrontierCode 1.1 Extended (score)
Coding · Higher is betterFrontierCode 1.1 Main (score)
Coding · Higher is betterGPQA Diamond
Science · Higher is betterARC-AGI-2
Reasoning · Higher is betterARC-AGI-1
Reasoning · Higher is betterAgents’ Last Exam
Computer use · Higher is betterOSWorld 2.0 · offline, v2026.08.08
Computer use · Higher is betterScreenSpot-Pro (no tools)
Computer use · Higher is betterBenchCAD (with tools)
Professional work · Higher is betterBrowseComp
Research · Higher is betterOpenScore String Quartets (1 − OMR-NED)
Professional work · Higher is betterDesign tasks (OpenAI internal)
Professional work · Higher is betterData science tasks (OpenAI internal)
Science · Higher is betterDatabase migration (OpenAI internal)
Coding · Higher is betterArtificial Analysis Coding Agent Index v1.4
Coding · Higher is better · index pointsGeneBench Pro
Science · Higher is betterMedChemBench (OpenAI internal)
Science · Higher is betterLifeSciBench
Science · Higher is betterHealthBench Professional (length-adjusted)
Science · Higher is betterExploitBench
Cybersecurity · Higher is betterExploitGym
Cybersecurity · Higher is betterExploitBench (June–August 2026)
Cybersecurity · Higher is betterSRE-Bench (single attempt)
Cybersecurity · Higher is betterSEC-Bench Pro
Cybersecurity · Higher is betterComputer-use safety (OpenAI internal)
Alignment · Lower is betterComputer-use safety with AutoReview
Alignment · Lower is betterCircumvention (OpenAI internal)
Alignment · Lower is betterExploitGym honeypot
Alignment · Lower is betterImpossible ExploitGym
Alignment · Higher is betterHallucination (OpenAI internal)
Alignment · Lower is betterMRCR v2 · 8 needles · 256K–512K
Long context · Higher is betterMRCR v2 · 8 needles · 512K–1M
Long context · Higher is betterARC-AGI-3
Reasoning · Higher is betterAll 40 benchmarks · 6 models View source snapshot
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 44.5% | 52.6% | 19.1% |
| Terminal-Bench Science 0.1 * | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | — |
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | — |
| Humanity’s Last Exam (with tools) | 57.2% | — | 65.0% | 63.8% | 63.6% | — |
| Artificial Analysis Intelligence Index v4.1.1 * | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 90.2% | 73.2% | — |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% |
| FrontierCode 1.1 Extended (score) * | 64.5% | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% |
| FrontierCode 1.1 Main (score) * | 53.3% | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 89.2% | 90.4% | — |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 98.5% | 97.5% | — |
| Agents’ Last Exam | 59.3% | 53.6% | — | 48.7% | 55.5% | — |
| OSWorld 2.0 · offline, v2026.08.08 * | 72.6% | 65.7% | — | — | 70.2% | — |
| ScreenSpot-Pro (no tools) * | 92.7% | 76.9% | — | — | — | — |
| BenchCAD (with tools) * | 95.9% | 83.3% | 84.3% | 67.5% | 82.1% | — |
| BrowseComp | 91.5% | 90.4% | — | 87.4% | 90.8% | — |
| OpenScore String Quartets (1 − OMR-NED) | 0.84 | 0.19 | — | — | — | — |
| Design tasks (OpenAI internal) | 50.0% | 47.4% | — | 35.8% | — | — |
| Data science tasks (OpenAI internal) | 40.9% | 30.5% | — | 34.7% | — | — |
| Database migration (OpenAI internal) | 63.9% | 42.7% | 57.8% | 50.3% | — | — |
| Artificial Analysis Coding Agent Index v1.4 | 67.0 | 65.1 | — | 67.2 | 68.1 | 61.2 |
| GeneBench Pro | 37.1% | 32.3% | — | — | — | — |
| MedChemBench (OpenAI internal) | 49.3% | 47.4% | — | — | — | — |
| LifeSciBench | 60.3% | 59.9% | — | — | — | — |
| HealthBench Professional (length-adjusted) * | 63.4% | 60.5% | 58.1% | 60.9% | 56.4% | 52.1% |
| ExploitBench * | 100.0% | 78.5% | — | — | 70.0% | — |
| ExploitGym * | 42.4% | 30.3% | — | — | 22.0% | — |
| ExploitBench (June–August 2026) * | 39.0% | 5.5% | — | — | — | — |
| SRE-Bench (single attempt) * | 88.0% | 55.9% | — | — | 12.5% | — |
| SEC-Bench Pro | 85.4% | 79.1% | — | — | — | — |
| Computer-use safety (OpenAI internal) ↓ * | 2.4% | 22.0% | 9.5% | 18.3% | 11.5% | — |
| Computer-use safety with AutoReview ↓ | 1.8% | 4.3% | — | — | — | — |
| Circumvention (OpenAI internal) ↓ | 0.0% | 0.29% | — | — | — | — |
| ExploitGym honeypot ↓ | 0.0% | 48.2% | — | — | — | — |
| Impossible ExploitGym | 100.0% | — | — | — | — | — |
| Hallucination (OpenAI internal) ↓ | 4.2% | 12.2% | — | — | — | — |
| MRCR v2 · 8 needles · 256K–512K | 100.0% | 91.5% | — | — | — | — |
| MRCR v2 · 8 needles · 512K–1M | 96.3% | 73.8% | — | — | — | — |
| ARC-AGI-3 * | 99.9% | 7.8% | — | — | 30.2% | — |
Astra scores and visual runsFable 5.1 scores and visual runsCompare API costs
Sources and methodology
Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
Gaps are differences in reported scores, not statistical significance. Percentage bars use 0–100; index bars use a 0–100 display scale.
Anthropic’s model report- Terminal-Bench Science 0.1
- Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- OSWorld 2.0 · offline, v2026.08.08
- Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
- ScreenSpot-Pro (no tools)
- The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
- HealthBench Professional (length-adjusted)
- OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
- ExploitBench
- Astra and Sol evaluated without production safeguards.
- ExploitGym
- The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
- ExploitBench (June–August 2026)
- OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
- SRE-Bench (single attempt)
- Cyber capability evaluation without production safeguards.
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
- ARC-AGI-3
- OpenAI Responses harness with two modified settings; see source footnote 1.
The September launch snapshot above and earlier independent measurements below retain their own sources and versions.
Independent comparison
Earlier independent scores across 9 models.
This retained Artificial Analysis Index v4.1 snapshot combines nine evaluations across coding, agentic work, science, knowledge, physics, and long-context reasoning. Higher is better. Astra and Fable 5.1’s v4.1.1 figures appear above; they are not mixed into this older index.
- Grok 4.6High61AA source ↗
- Claude Fable 5Max, Opus 4.8 fallback60AA source ↗
- Claude Opus 5Adaptive reasoning, xhigh60AA source ↗
- GPT-5.6 SolMax59AA source ↗
- Kimi K3Reasoning57AA source ↗
- 55AA source ↗
- GPT-5.5xhigh55AA source ↗
- Grok 4.5High54AA source ↗
- GPT-5.6 LunaHigh46AA source ↗
Configuration matters. Fable 5 includes its published Opus 4.8 fallback. GPT and Claude scores use the effort level shown beside each model.
Cost explorer
Estimate your token costs.
Set total input and output tokens for a workload. This estimate uses the displayed short-context API rates, including active promotions, across all requests.
Astra and Fable 5.1 share $10 input / $50 output standard rates per million tokens. Astra requests above 272K input tokens use $20 input / $75 output for the full request; this calculator applies that threshold. Cache reads, cache writes, tools, and Fast mode are excluded.
| Model | AA Index | Input / 1M | Output / 1M | Speed | Your task |
|---|---|---|---|---|---|
| Claude Fable 5Anthropic · Max, Opus 4.8 fallback | 60Index v4.1 | $10.00 | $50.00 | 69.7tokens/sec | AA data ↗Price ↗ |
| Claude Opus 5Anthropic · Adaptive reasoning, xhigh | 60Index v4.1 | $5.00 | $25.00 | 53.1tokens/sec | AA data ↗Price ↗ |
| GPT-5.6 SolOpenAI · Max | 59Index v4.1 | $4.00 | $20.00 | 77.1tokens/sec | AA data ↗Price ↗ |
| Kimi K3Moonshot AI · Reasoning | 57Index v4.1 | $3.00 | $15.00 | 32.0tokens/sec | AA data ↗Price ↗ |
| GPT-5.6 TerraOpenAI · Max | 55Index v4.1 | $2.00 | $12.00 | 134.5tokens/sec | AA data ↗Price ↗ |
| GPT-5.5OpenAI · xhigh | 55Index v4.1 | $5.00 | $30.00 | 72.5tokens/sec | AA data ↗Price ↗ |
| Grok 4.5Xai · High | 54Index v4.1 | $2.00 | $6.00 | 67.1tokens/sec | AA data ↗Price ↗ |
| Grok 4.6Xai · High | 61Index v4.1 | $2.00 | $6.00 | 65.1tokens/sec | AA data ↗Price ↗ |
| GPT-5.6 LunaOpenAI · High | 46Index v4.1 | $0.20 | $1.20 | 178.0tokens/sec | AA data ↗Price ↗ |
| GPT-6 AstraOpenAI · Standard API · pricing checked September 7, 2026 | —See v4.1.1 above | $10.00 | $50.00 | Not measuredNo speed estimate | Price ↗ |
| Claude Fable 5.1Anthropic · Standard API · pricing checked September 7, 2026 | —See v4.1.1 above | $10.00 | $50.00 | Not measuredNo speed estimate | Price ↗ |
What is included: uncached input and generated output at short-context API rates verified 2026-09-08.
What is not included: long-context premiums, cache writes, tool calls, storage, batch discounts, and provider markups. A long prompt may cost more than this estimate; check the linked pricing source.
Blended reference: Calculated from the current catalog using a 7:2:1 cache-hit/input/output mix. The lowest in this set is GPT-5.6 Luna at $0.17 per 1M blended tokens.
Methodology
Limits of this comparison
Artificial Analysis Index results come from one independent harness. The exact model configuration remains attached to every score.
Token cost scenarios multiply published input and output prices by the workload you enter. They exclude tools, storage, cache writes, and provider discounts.
A shared-table mark means the source published the compared models in one table. A same-eval mark means the benchmark name matches, but the source or configuration differs.
Missing data is not a low score. Models without a compatible public score remain unranked until a reproducible result is published.
Our Superbash visual runs are a separate, same-prompt evidence layer. Use the live run comparison to inspect build quality directly.