Current comparison
Four models, one set of briefs.
Compare Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol, and DeepSeek V4.1 Flash using their published API facts and inspectable same-prompt visual runs. Run counts are evidence coverage, not quality scores.
| Model | Context | Input / output | Published visual runs | Source |
|---|---|---|---|---|
| Claude Opus 5.5Anthropic | 1M | $4 / $20USD per 1M tokens · Claude API Standard | 10Inspect artifacts → | API facts ↗ |
| GPT-6 AstraOpenAI | 1.05M | $10 / $50USD per 1M tokens · OpenAI API Standard (≤272K input tokens) | 20Inspect artifacts → | API facts ↗ |
| GPT-6 SolOpenAI | 1.05M | $2 / $10USD per 1M tokens · OpenAI API Standard (≤272K input tokens) | 10Inspect artifacts → | API facts ↗ |
| DeepSeek V4.1 FlashDeepSeek | 1M | $0.3 / $1.2USD per 1M tokens · DeepSeek API peak rate | 16Inspect artifacts → | API facts ↗ |
DeepSeek prices above use its peak rate; off-peak rates are half. OpenAI prices shown are for prompts up to 272K input tokens. Anthropic reports Opus 5.5 versus Astra on named evaluations; Sol and DeepSeek lack a compatible published result in that table. Open the four-model Starfall Arena comparison →
Compare benchmark scores.
Choose two models and a workload to see where each leads.
Provider-published results · checked
Read OpenAI’s source table6 selected benchmarks · GPT-6 Astra vs Claude Fable 5.1 · pp = percentage points
Terminal-Bench 4.0
Coding · Higher is betterTerminal-Bench Science 0.1
Science · Higher is betterAutomationBench
Automation · Higher is betterHumanity’s Last Exam (with tools)
Reasoning · Higher is betterArtificial Analysis Intelligence Index v4.1.1
Reasoning · Higher is better · index pointsFrontierMath Tier 4 (v2)
Reasoning · Higher is betterDeepSWE v1.1
Coding · Higher is betterFrontierCode 1.1 Extended (score)
Coding · Higher is betterFrontierCode 1.1 Main (score)
Coding · Higher is betterGPQA Diamond
Science · Higher is betterARC-AGI-2
Reasoning · Higher is betterARC-AGI-1
Reasoning · Higher is betterAgents’ Last Exam
Computer use · Higher is betterOSWorld 2.0 · offline, v2026.08.08
Computer use · Higher is betterScreenSpot-Pro (no tools)
Computer use · Higher is betterBenchCAD (with tools)
Professional work · Higher is betterBrowseComp
Research · Higher is betterOpenScore String Quartets (1 − OMR-NED)
Professional work · Higher is betterDesign tasks (OpenAI internal)
Professional work · Higher is betterData science tasks (OpenAI internal)
Science · Higher is betterDatabase migration (OpenAI internal)
Coding · Higher is betterArtificial Analysis Coding Agent Index v1.4
Coding · Higher is better · index pointsGeneBench Pro
Science · Higher is betterMedChemBench (OpenAI internal)
Science · Higher is betterLifeSciBench
Science · Higher is betterHealthBench Professional (length-adjusted)
Science · Higher is betterExploitBench
Cybersecurity · Higher is betterExploitGym
Cybersecurity · Higher is betterExploitBench (June–August 2026)
Cybersecurity · Higher is betterSRE-Bench (single attempt)
Cybersecurity · Higher is betterSEC-Bench Pro
Cybersecurity · Higher is betterComputer-use safety (OpenAI internal)
Alignment · Lower is betterComputer-use safety with AutoReview
Alignment · Lower is betterCircumvention (OpenAI internal)
Alignment · Lower is betterExploitGym honeypot
Alignment · Lower is betterImpossible ExploitGym
Alignment · Higher is betterHallucination (OpenAI internal)
Alignment · Lower is betterMRCR v2 · 8 needles · 256K–512K
Long context · Higher is betterMRCR v2 · 8 needles · 512K–1M
Long context · Higher is betterARC-AGI-3
Reasoning · Higher is betterAll 40 benchmarks · 3 models View source snapshot
| Benchmark | Astra | Fable 5.1 | Gemini 3.8 Flash |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 55.8% | 19.1% |
| Terminal-Bench Science 0.1 * | 64.6% | 52.6% | — |
| AutomationBench | 41.4% | 31.4% | — |
| Humanity’s Last Exam (with tools) | 57.2% | 65.0% | — |
| Artificial Analysis Intelligence Index v4.1.1 * | 61.2 | 65.7 | 58.7 |
| FrontierMath Tier 4 (v2) | 97.6% | 87.8% | — |
| DeepSWE v1.1 | 74.1% | 67.4% | 73.8% |
| FrontierCode 1.1 Extended (score) * | 64.5% | 63.6% | 56.3% |
| FrontierCode 1.1 Main (score) * | 53.3% | 50.9% | 43.6% |
| GPQA Diamond | 96.0% | 93.7% | 95.3% |
| ARC-AGI-2 | 95.0% | 90.0% | — |
| ARC-AGI-1 | 98.5% | 97.5% | — |
| Agents’ Last Exam | 59.3% | — | — |
| OSWorld 2.0 · offline, v2026.08.08 * | 72.6% | — | — |
| ScreenSpot-Pro (no tools) * | 92.7% | — | — |
| BenchCAD (with tools) * | 95.9% | 84.3% | — |
| BrowseComp | 91.5% | — | — |
| OpenScore String Quartets (1 − OMR-NED) | 0.84 | — | — |
| Design tasks (OpenAI internal) | 50.0% | — | — |
| Data science tasks (OpenAI internal) | 40.9% | — | — |
| Database migration (OpenAI internal) | 63.9% | 57.8% | — |
| Artificial Analysis Coding Agent Index v1.4 | 67.0 | — | 61.2 |
| GeneBench Pro | 37.1% | — | — |
| MedChemBench (OpenAI internal) | 49.3% | — | — |
| LifeSciBench | 60.3% | — | — |
| HealthBench Professional (length-adjusted) * | 63.4% | 58.1% | 52.1% |
| ExploitBench * | 100.0% | — | — |
| ExploitGym * | 42.4% | — | — |
| ExploitBench (June–August 2026) * | 39.0% | — | — |
| SRE-Bench (single attempt) * | 88.0% | — | — |
| SEC-Bench Pro | 85.4% | — | — |
| Computer-use safety (OpenAI internal) ↓ * | 2.4% | 9.5% | — |
| Computer-use safety with AutoReview ↓ | 1.8% | — | — |
| Circumvention (OpenAI internal) ↓ | 0.0% | — | — |
| ExploitGym honeypot ↓ | 0.0% | — | — |
| Impossible ExploitGym | 100.0% | — | — |
| Hallucination (OpenAI internal) ↓ | 4.2% | — | — |
| MRCR v2 · 8 needles · 256K–512K | 100.0% | — | — |
| MRCR v2 · 8 needles · 512K–1M | 96.3% | — | — |
| ARC-AGI-3 * | 99.9% | — | — |
Astra scores and visual runsFable 5.1 scores and visual runsCompare API costs
Sources and methodology
Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
Gaps are differences in reported scores, not statistical significance. Percentage bars use 0–100; index bars use a 0–100 display scale.
Anthropic’s model report- Terminal-Bench Science 0.1
- Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- OSWorld 2.0 · offline, v2026.08.08
- Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
- ScreenSpot-Pro (no tools)
- The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
- HealthBench Professional (length-adjusted)
- OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
- ExploitBench
- Astra and Sol evaluated without production safeguards.
- ExploitGym
- The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
- ExploitBench (June–August 2026)
- OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
- SRE-Bench (single attempt)
- Cyber capability evaluation without production safeguards.
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
- ARC-AGI-3
- OpenAI Responses harness with two modified settings; see source footnote 1.
The September launch snapshot above and earlier independent measurements below retain their own sources and versions.
Independent comparison
Earlier independent scores across 1 models.
This retained Artificial Analysis Index v4.1 snapshot combines nine evaluations across coding, agentic work, science, knowledge, physics, and long-context reasoning. Higher is better. Astra and Fable 5.1’s v4.1.1 figures appear above; they are not mixed into this older index.
- Kimi K3Reasoning57AA source ↗
Configuration matters. This earlier index retains each model’s published effort setting and source.
Cost explorer
Estimate your token costs.
Set total input and output tokens for a workload. This estimate uses the displayed short-context API rates, including active promotions, across all requests.
Opus 5.5 is $4 input / $20 output; Astra and Fable 5.1 are $10 / $50; Sol is $2 / $10. Astra and Sol requests above 272K input tokens use their higher rates for the full request. Grok 4.7 requests at or above 200K prompt tokens use $4 / $12. DeepSeek uses peak rates here; off-peak rates are half. Cache reads, cache writes, tools, regional premiums, and Fast mode are excluded.
| Model | AA Index | Input / 1M | Output / 1M | Speed | Your task |
|---|---|---|---|---|---|
| Kimi K3Moonshot AI · Reasoning | 57Index v4.1 | $3.00 | $15.00 | 32.0tokens/sec | AA data ↗Price ↗ |
| Claude Opus 5.5Anthropic · Standard API · pricing checked 2026-09-23 | —See v4.1.1 above | $4.00 | $20.00 | Not measuredNo speed estimate | Price ↗ |
| GPT-6 AstraOpenAI · Standard API · pricing checked 2026-09-08 | —See v4.1.1 above | $10.00 | $50.00 | Not measuredNo speed estimate | Price ↗ |
| GPT-6 SolOpenAI · Standard API · pricing checked 2026-09-23 | —See v4.1.1 above | $2.00 | $10.00 | Not measuredNo speed estimate | Price ↗ |
| DeepSeek V4.1 FlashDeepSeek · Standard API · pricing checked 2026-09-21 | —See v4.1.1 above | $0.30 | $1.20 | Not measuredNo speed estimate | Price ↗ |
| Claude Fable 5.1Anthropic · Standard API · pricing checked 2026-09-08 | —See v4.1.1 above | $10.00 | $50.00 | Not measuredNo speed estimate | Price ↗ |
| Grok 4.7SpaceXAI · Standard API · pricing checked 2026-09-22 | —See v4.1.1 above | $2.00 | $6.00 | Not measuredNo speed estimate | Price ↗ |
What is included: uncached input and generated output at short-context API rates. The baseline full-catalog audit was 2026-09-08; newer additions retain their own verification date in the catalog.
What is not included: long-context premiums, cache writes, tool calls, storage, batch discounts, and provider markups. A long prompt may cost more than this estimate; check the linked pricing source.
Blended reference: Calculated from the current catalog using a 7:2:1 cache-hit/input/output mix. The lowest in this set is Kimi K3 at $2.31 per 1M blended tokens.
Methodology
Limits of this comparison
Artificial Analysis Index results come from one independent harness. The exact model configuration remains attached to every score.
Token cost scenarios multiply published input and output prices by the workload you enter. They exclude tools, storage, cache writes, and provider discounts.
A shared-table mark means the source published the compared models in one table. A same-eval mark means the benchmark name matches, but the source or configuration differs.
Missing data is not a low score. Models without a compatible public score remain unranked until a reproducible result is published.
Our Superbash visual runs are a separate, same-prompt evidence layer. Use the live run comparison to inspect build quality directly.