OpenAI

Tier S · Ship-to-prod default.

GPT-6 Astra

S-tier for end-to-end agent work, backed by a complete set of 16 published Superbash visual benchmark runs.

API pricing

USD per million tokens · OpenAI API Standard (≤272K input tokens)
Input
$10
Output
$50
Cached input
$1
Context window
1.05M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

GPT-6 Astra completed all 16 Superbash visual benchmark briefs in independent prompt-only runs. Every published artifact is linked from the benchmark page with its manifest and validation record. S tier reflects the completed, inspectable run set and the editorial assessment of Astra as the strongest current choice for difficult multi-step product work; it is not a numeric leaderboard claim.

Best for

  • Difficult multi-step product work
  • Agent-led implementation with review
  • Browser and UI verification
  • Complex coding tasks that need persistent follow-through

Watch out

  • The visual suite is local evidence, not a normalized cross-provider score
  • Public API limits, pricing, and context specifications are not asserted by this profile
  • Review scope and visual decisions before shipping

Why it is ranked here

  1. Astra has a complete 16-run Superbash visual benchmark set, covering the full published prompt suite rather than a partial sample.
  2. Each run has an inspectable artifact, manifest, and validation record; the method deliberately keeps those local results separate from provider or third-party numeric score tables.
  3. S tier is an editorial routing decision for demanding end-to-end work, not a claim that Astra wins every benchmark or workload.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: GPT-6 Astra is currently placed in Tier S.

The September 2026 editorial roster places GPT-6 Astra at rank 1.

Open source →
2026-09-07

Superbash visual benchmark suite: GPT-6 Astra

Takeaway: All 16 prompt-only Astra runs are published with manifests, artifact hashes, and validation records.

The methodology record describes how the local runs were isolated and validated; it is evidence of this site’s visual suite, not a provider benchmark or independent universal score.

Open source →

How GPT-6 Astra compares

Three relevant peers, side by side.

  • GPT-6 Astra This model
  • Claude Fable 5.1
  • GPT-5.6 Sol
  • Claude Opus 5

6 benchmarks · 4 models

Terminal-Bench 4.0

Coding · Higher is better

Source
GPT-6 Astra57.9%
Claude Fable 5.155.8%
GPT-5.6 Sol37.3%
Claude Opus 552.6%

Agents’ Last Exam

Computer use · Higher is better

Source
GPT-6 Astra59.3%
Claude Fable 5.1Not reported
GPT-5.6 Sol53.6%
Claude Opus 555.5%

AutomationBench

Automation · Higher is better

Source
GPT-6 Astra41.4%
Claude Fable 5.131.4%
GPT-5.6 Sol18.1%
Claude Opus 526.9%

Terminal-Bench Science 0.1

Science · Higher is better

Source
GPT-6 Astra64.6%
Claude Fable 5.152.6%
GPT-5.6 Sol22.4%
Claude Opus 530.0%

DeepSWE v1.1

Coding · Higher is better

Source
GPT-6 Astra74.1%
Claude Fable 5.167.4%
GPT-5.6 Sol72.7%
Claude Opus 573.7%

Humanity’s Last Exam (with tools)

Reasoning · Higher is better

Source
GPT-6 Astra57.2%
Claude Fable 5.165.0%
GPT-5.6 SolNot reported
Claude Opus 563.6%

Official benchmark profile

Astra’s published benchmark results

OpenAI sourceChecked September 7, 2026Source report →
Full official benchmark table40 scores and peer comparisons
BenchmarkAreaScoreComparison
Terminal-Bench 4.0Coding57.9%
GPT-6 Astra57.9%
Claude Fable 5.155.8%
Claude Opus 552.6%
Claude Fable 544.5%
Terminal-Bench Science 0.1Science64.6%
GPT-6 Astra64.6%
Claude Fable 5.152.6%
Claude Opus 530.0%
GPT-5.6 Sol22.4%
AutomationBenchAutomation41.4%
GPT-6 Astra41.4%
Claude Fable 5.131.4%
Claude Opus 526.9%
GPT-5.6 Sol18.1%
Humanity’s Last Exam (with tools)Reasoning57.2%
GPT-6 Astra57.2%
Claude Opus 563.6%
Claude Fable 563.8%
Claude Fable 5.165.0%
Artificial Analysis Intelligence Index v4.1.1Reasoning61.2
FrontierMath Tier 4 (v2)Reasoning97.6%
GPT-6 Astra97.6%
Claude Fable 590.2%
Claude Fable 5.187.8%
GPT-5.6 Sol83.0%
DeepSWE v1.1Coding74.1%
GPT-6 Astra74.1%
Gemini 3.8 Flash73.8%
Claude Opus 573.7%
GPT-5.6 Sol72.7%
FrontierCode 1.1 Extended (score)Coding64.5%
GPT-6 Astra64.5%
Claude Fable 564.9%
Claude Fable 5.163.6%
Claude Opus 563.6%
FrontierCode 1.1 Main (score)Coding53.3%
GPT-6 Astra53.3%
Claude Opus 553.4%
Claude Fable 553.5%
Claude Fable 5.150.9%
GPQA DiamondScience96.0%
GPT-6 Astra96.0%
Gemini 3.8 Flash95.3%
GPT-5.6 Sol94.6%
Claude Fable 5.193.7%
ARC-AGI-2Reasoning95.0%
GPT-6 Astra95.0%
GPT-5.6 Sol92.5%
Claude Opus 590.4%
Claude Fable 5.190.0%
ARC-AGI-1Reasoning98.5%
GPT-6 Astra98.5%
Claude Fable 598.5%
GPT-5.6 Sol97.5%
Claude Fable 5.197.5%
Agents’ Last ExamComputer use59.3%
GPT-6 Astra59.3%
Claude Opus 555.5%
GPT-5.6 Sol53.6%
Claude Fable 548.7%
OSWorld 2.0 · offline, v2026.08.08Computer use72.6%
ScreenSpot-Pro (no tools)Computer use92.7%
BenchCAD (with tools)Professional work95.9%
BrowseCompResearch91.5%
GPT-6 Astra91.5%
Claude Opus 590.8%
GPT-5.6 Sol90.4%
Claude Fable 587.4%
OpenScore String Quartets (1 − OMR-NED)Professional work0.84
Design tasks (OpenAI internal)Professional work50.0%
GPT-6 Astra50.0%
GPT-5.6 Sol47.4%
Claude Fable 535.8%
Data science tasks (OpenAI internal)Science40.9%
GPT-6 Astra40.9%
Claude Fable 534.7%
GPT-5.6 Sol30.5%
Database migration (OpenAI internal)Coding63.9%
GPT-6 Astra63.9%
Claude Fable 5.157.8%
Claude Fable 550.3%
GPT-5.6 Sol42.7%
Artificial Analysis Coding Agent Index v1.4Coding67.0
GeneBench ProScience37.1%
GPT-6 Astra37.1%
GPT-5.6 Sol32.3%
MedChemBench (OpenAI internal)Science49.3%
GPT-6 Astra49.3%
GPT-5.6 Sol47.4%
LifeSciBenchScience60.3%
GPT-6 Astra60.3%
GPT-5.6 Sol59.9%
HealthBench Professional (length-adjusted)Science63.4%
ExploitBenchCybersecurity100.0%
GPT-6 Astra100.0%
GPT-5.6 Sol78.5%
Claude Opus 570.0%
ExploitGymCybersecurity42.4%
ExploitBench (June–August 2026)Cybersecurity39.0%
SRE-Bench (single attempt)Cybersecurity88.0%
GPT-6 Astra88.0%
GPT-5.6 Sol55.9%
Claude Opus 512.5%
SEC-Bench ProCybersecurity85.4%
GPT-6 Astra85.4%
GPT-5.6 Sol79.1%
Computer-use safety (OpenAI internal)Lower is betterAlignment2.4%
Computer-use safety with AutoReviewLower is betterAlignment1.8%
GPT-6 Astra1.8%
GPT-5.6 Sol4.3%
Circumvention (OpenAI internal)Lower is betterAlignment0.0%
GPT-6 Astra0.0%
GPT-5.6 Sol0.29%
ExploitGym honeypotLower is betterAlignment0.0%
GPT-6 Astra0.0%
GPT-5.6 Sol48.2%
Impossible ExploitGymAlignment100.0%
Hallucination (OpenAI internal)Lower is betterAlignment4.2%
GPT-6 Astra4.2%
GPT-5.6 Sol12.2%
MRCR v2 · 8 needles · 256K–512KLong context100.0%
GPT-6 Astra100.0%
GPT-5.6 Sol91.5%
MRCR v2 · 8 needles · 512K–1MLong context96.3%
GPT-6 Astra96.3%
GPT-5.6 Sol73.8%
ARC-AGI-3Reasoning99.9%
GPT-6 Astra99.9%
Claude Opus 530.2%
GPT-5.6 Sol7.8%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
gpt-6-astra
Context
1.05M
Max output
128K
License
Not publicly verified
Access route
OpenAI API; rolling out through ChatGPT and Codex, Microsoft Azure, and AWS Bedrock. Access depends on account and rollout. The September 7 visual benchmarks used Codex subagents; API prices are not the recorded cost of those runs.
Modalities
text input, image input, text output
Hardware
Not publicly verified
OpenAI API Standard (≤272K input tokens) · USD / 1M tokens
$10 input · $1 cached input · $12.5 cache write · $50 output

Above 272K input tokens, the full request costs $20 input, $2 cached input, $25 cache write, and $75 output per 1M tokens. Batch, Flex, Fast mode, and regional pricing differ.

Benchmark sources

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.

OpenAI Astra launch comparison table

Local visual runs and their validation records are separate from these provider-published scores.

Evaluation settings

  • Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
  • Lower is better. Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards.
Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.

Comparison methodology

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.

Anthropic model report

Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.

Local visual-run methodology and validation records