OpenAI
GPT-5.6 Sol
Version 5.6-high benchmark runs across the shared Superbash visual prompts.
API pricing
USD per million tokens · OpenAI API Standard (≤272K input tokens)- Input
- $4
- Output
- $20
- Cached input
- $0.4
- Context window
- 1.05M
Inspect the original visual runs
Open the generated scene to test it, or compare that prompt with another model.

Helm's Deep
Run date: 2026-07-10
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Run date: 2026-07-10
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Stormwind Trebuchet Simulator
Run date: 2026-07-24
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly
Run date: 2026-07-10
Specialist test · evaluated separately
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Retired tests · 6 archived runs
These tests are no longer in the active suite. Their original results remain available for inspection.
How GPT-5.6 Sol compares
Three relevant peers, side by side.
- GPT-5.6 Sol This model
- GPT-6 Astra
- Claude Fable 5.1
- Claude Opus 5
6 benchmarks · 4 models
Official benchmark profile
Published scores outside our visual tests
SWE-bench Pro
64.6%Terminal-Bench 2.1
88.8%OSWorld 2.0
62.6%Full official benchmark table9 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| SWE-bench Pro | Coding | 64.6% | |
| Terminal-Bench 2.1 | Agentic coding | 88.8% | |
| OSWorld 2.0 | Computer use | 62.6% | |
| BrowseComp | Tool use | 90.4% | |
| BenchCAD | Computer-aided design | 70.6% | |
| BenchCAD with Python tool | Tool use | 83.4% | |
| GPQA Diamond | Academic reasoning | 94.6% | |
| FrontierMath Tier 1-3 v2 | Math | 89.0% | |
| FrontierMath Tier 4 v2 | Math | 83.0% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Superseded in this ranking
- API model ID
gpt-5.6-sol- Context
- 1.05M
- Max output
- 128K
- License
- Not publicly verified
- Access route
- Not publicly verified
- Modalities
- Not publicly verified
- Hardware
- Not publicly verified
- OpenAI API Standard (≤272K input tokens) · USD / 1M tokens
- $4 input · $0.4 cached input · $5 cache write · $20 output
OpenAI API promotional rates available through at least November 21, 2026. Above 272K input tokens, the full request costs $8 input, $0.8 cached input, $10 cache write, and $30 output per 1M tokens. Batch, Flex, Fast mode, and regional pricing differ.
Benchmark sources
OpenAI presents Sol as the flagship GPT-5.6 model. These rows use the shared GPT-5.6 comparison table so the scores can be read directly against Terra, Luna, GPT-5.5, Claude, and Gemini where published.
GPT-5.6: Frontier intelligence that scales with your ambitionEvaluation settings
- Shared OpenAI GPT-5.6 comparison table.
- High-effort computer-use comparison table.
- High-effort score; OpenAI separately reports Sol Ultra at 92.2%.
- Vision2Code score without Python tool.
- Vision2Code score with Python tool.
- Shared reasoning benchmark table.
Comparison methodology
Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.
- Terminal-Bench Science 0.1
- Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
- Artificial Analysis Intelligence Index v4.1.1
- Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
- FrontierCode 1.1 Extended (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- FrontierCode 1.1 Main (score)
- Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
- OSWorld 2.0 · offline, v2026.08.08
- Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
- ScreenSpot-Pro (no tools)
- The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
- BenchCAD (with tools)
- Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
- HealthBench Professional (length-adjusted)
- OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
- ExploitBench
- Astra and Sol evaluated without production safeguards.
- ExploitGym
- The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
- ExploitBench (June–August 2026)
- OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
- SRE-Bench (single attempt)
- Cyber capability evaluation without production safeguards.
- Computer-use safety (OpenAI internal)
- Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
- ARC-AGI-3
- OpenAI Responses harness with two modified settings; see source footnote 1.