OpenAI

Tier A · Strong with clear operating rules.

GPT-5.6 Sol

A dependable heavy-lifting builder: practical, capable, and willing to finish.

API pricing

USD per million tokens · OpenAI API Standard (≤272K input tokens)
Input
$4
Output
$20
Cached input
$0.4
Context window
1.05M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

10 runs

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

GPT-5.6 Sol is an A-tier choice for demanding coding and product work. It remains a dependable builder for real app work and migrations, but Astra now holds the S-tier recommendation for the hardest end-to-end agent tasks.

Best for

  • Heavy coding work
  • App builds and migrations
  • Long implementation sessions
  • A dependable frontier daily driver

Watch out

  • High-end usage can still be expensive
  • Subscription limits and resets can change
  • Review production changes before shipping

Why it is ranked here

  1. The team uses Sol for much of its own heavy coding work and described it as smart enough for most serious builds.
  2. Its advantage is execution discipline: it “shuts up and works” instead of turning every task into a debate.
  3. A tier preserves Sol as a strong, dependable option while reserving S tier for Astra’s complete current visual run set and end-to-end editorial recommendation.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: GPT-5.6 Sol is currently placed in Tier A.

The September 2026 editorial roster places GPT-5.6 Sol at rank 5.

Open source →

How GPT-5.6 Sol compares

Three relevant peers, side by side.

  • GPT-5.6 Sol This model
  • GPT-6 Astra
  • Claude Fable 5.1
  • Claude Opus 5

6 benchmarks · 4 models

Terminal-Bench 4.0

Coding · Higher is better

Source
GPT-5.6 Sol37.3%
GPT-6 Astra57.9%
Claude Fable 5.155.8%
Claude Opus 552.6%

Agents’ Last Exam

Computer use · Higher is better

Source
GPT-5.6 Sol53.6%
GPT-6 Astra59.3%
Claude Fable 5.1Not reported
Claude Opus 555.5%

DeepSWE v1.1

Coding · Higher is better

Source
GPT-5.6 Sol72.7%
GPT-6 Astra74.1%
Claude Fable 5.167.4%
Claude Opus 573.7%

AutomationBench

Automation · Higher is better

Source
GPT-5.6 Sol18.1%
GPT-6 Astra41.4%
Claude Fable 5.131.4%
Claude Opus 526.9%

BrowseComp

Research · Higher is better

Source
GPT-5.6 Sol90.4%
GPT-6 Astra91.5%
Claude Fable 5.1Not reported
Claude Opus 590.8%

Terminal-Bench Science 0.1

Science · Higher is better

Source
GPT-5.6 Sol22.4%
GPT-6 Astra64.6%
Claude Fable 5.152.6%
Claude Opus 530.0%

Official benchmark profile

Published scores outside our visual tests

OpenAI sourceJuly 2026Source report →
Coding

SWE-bench Pro

64.6%
Agentic coding

Terminal-Bench 2.1

88.8%
Computer use

OSWorld 2.0

62.6%
Full official benchmark table9 scores and peer comparisons
BenchmarkAreaScoreComparison
SWE-bench ProCoding64.6%
GPT-5.6 Sol64.6%
GPT-5.6 Terra63.4%
GPT-5.6 Luna62.7%
Claude Opus 4.869.2%
Terminal-Bench 2.1Agentic coding88.8%
GPT-5.6 Sol88.8%
GPT-5.6 Terra87.4%
GPT-5.585.6%
GPT-5.6 Luna84.7%
OSWorld 2.0Computer use62.6%
GPT-5.6 Sol62.6%
Claude Opus 4.854.8%
GPT-5.6 Terra50.2%
GPT-5.547.5%
BrowseCompTool use90.4%
GPT-5.6 Sol90.4%
GPT-5.6 Sol Ultra92.2%
Claude Mythos 588.0%
Claude Mythos Preview87.9%
BenchCADComputer-aided design70.6%
GPT-5.6 Sol70.6%
GPT-5.6 Luna63.1%
GPT-5.6 Terra62.3%
GPT-5.544.4%
BenchCAD with Python toolTool use83.4%
GPT-5.6 Sol83.4%
GPT-5.6 Terra78.2%
GPT-5.6 Luna73.9%
GPT-5.559.4%
GPQA DiamondAcademic reasoning94.6%
GPT-5.6 Sol94.6%
Gemini 3.1 Pro94.3%
Claude Mythos 594.1%
GPT-5.593.6%
FrontierMath Tier 1-3 v2Math89.0%
GPT-5.6 Sol89.0%
Claude Fable 587.0%
GPT-5.585.3%
GPT-5.6 Terra84.9%
FrontierMath Tier 4 v2Math83.0%
GPT-5.6 Sol83.0%
Claude Fable 587.8%
GPT-5.572.5%
GPT-5.6 Terra68.3%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
gpt-5.6-sol
Context
1.05M
Max output
128K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
OpenAI API Standard (≤272K input tokens) · USD / 1M tokens
$4 input · $0.4 cached input · $5 cache write · $20 output

OpenAI API promotional rates available through at least November 21, 2026. Above 272K input tokens, the full request costs $8 input, $0.8 cached input, $10 cache write, and $30 output per 1M tokens. Batch, Flex, Fast mode, and regional pricing differ.

Benchmark sources

OpenAI presents Sol as the flagship GPT-5.6 model. These rows use the shared GPT-5.6 comparison table so the scores can be read directly against Terra, Luna, GPT-5.5, Claude, and Gemini where published.

GPT-5.6: Frontier intelligence that scales with your ambition

Evaluation settings

  • Shared OpenAI GPT-5.6 comparison table.
  • High-effort computer-use comparison table.
  • High-effort score; OpenAI separately reports Sol Ultra at 92.2%.
  • Vision2Code score without Python tool.
  • Vision2Code score with Python tool.
  • Shared reasoning benchmark table.

Comparison methodology

Maximum reported score at any effort; research/API environments. A shared table does not guarantee identical harnesses, budgets, or safeguards. Gaps describe reported scores, not statistical significance. Percentage bars run from 0 to 100; index bars use a 0–100 display scale.

Anthropic model report

Terminal-Bench Science 0.1
Anthropic reports standard error of about 3.5–4.5 points per model; small gaps may be noise.
Artificial Analysis Intelligence Index v4.1.1
Independent index as reproduced in OpenAI’s launch table; separate from the older v4.1 observations in the cost explorer.
FrontierCode 1.1 Extended (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
FrontierCode 1.1 Main (score)
Astra uses additional developer guidance about keeping code and tests scoped (source footnote 8).
OSWorld 2.0 · offline, v2026.08.08
Partial score on the offline subset. Anthropic’s modified tasks and grading are incompatible; no Fable 5.1 value is supplied here.
ScreenSpot-Pro (no tools)
The source labels Mythos’s 87.3 as Fable 5 (footnote 17). It is excluded from the Fable column here because the model’s safeguards differ.
BenchCAD (with tools)
Claude results use three evaluation modifications. Scores shown for context; no direct gap is calculated.
HealthBench Professional (length-adjusted)
OpenAI re-evaluation with GPT-5.4 grading. Fable 5.1 uses Opus 5 fallback for refusals; not a standalone Fable result.
ExploitBench
Astra and Sol evaluated without production safeguards.
ExploitGym
The source’s Fable columns are Mythos 5.1 (30.4) and Mythos 5 (28.4), excluded here. Astra/Sol have no six-hour limit and no production safeguards. Not a production-model comparison.
ExploitBench (June–August 2026)
OpenAI internal set, without production safeguards. Sol’s 5.5 is affected by a 300-turn cap; source also reports 11.5 with fewer limits.
SRE-Bench (single attempt)
Cyber capability evaluation without production safeguards.
Computer-use safety (OpenAI internal)
Lower is better. Research harness without production confirmations; provider safeguards and tools still differ.
ARC-AGI-3
OpenAI Responses harness with two modified settings; see source footnote 1.