OpenAI

Tier B · Specialist or second-choice models.

GPT-5.5

A very good delivery model that has been overtaken by cheaper Terra.

API pricing

USD per million tokens · OpenAI API Standard (≤272K input tokens)
Input
$5
Output
$30
Cached input
$0.5
Context window
1.05M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

9 runs

Editorial verdict

Where this model fits

Back to all benchmarks →

GPT-5.5 still writes clean code and fits OpenAI tooling well, but the August recording treats it as a relic of the previous generation: Terra is roughly half the API price in the current catalog and is now the team’s routine daily driver.

Best for

  • Plan mode
  • Architecture drafts
  • OpenAI-native workflows
  • Shorter coding tasks

Watch out

  • Reasoning overhead can fill context
  • Can drift without explicit goals
  • May add features you did not ask for

Why it is ranked here

  1. The video praises its plan mode as one of the best tested options.
  2. For long-horizon vibe coding, the model can overcomplicate work and enter the context-window danger zone.
  3. The newer Terra option now undercuts its practical role, so GPT-5.5 remains in B rather than moving up.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: GPT-5.5 is currently placed in Tier B.

The September 2026 editorial roster places GPT-5.5 at rank 13.

Open source →
2026-06-17

NEW AI Model Tier List for Vibe Coding!

Takeaway: Use GPT-5.5 in plan mode rather than as a long-horizon coding default.

The source video says it is strong and clean, but overthinks and spends too much context on extended builds.

Open source →

How GPT-5.5 compares

Three relevant peers, side by side.

  • GPT-5.5 This model
  • Claude Opus 4.8
  • GPT-5.6 Luna
  • GPT-5.6 Sol

6 benchmarks · 4 models

SWE-bench Pro

Coding · Higher is better

Source
GPT-5.558.6%
Claude Opus 4.869.2%
GPT-5.6 Luna62.7%
GPT-5.6 Sol64.6%

Terminal-Bench 2.1

Agentic coding · Higher is better

Source
GPT-5.585.6%
Claude Opus 4.878.9%
GPT-5.6 Luna84.7%
GPT-5.6 Sol88.8%

OSWorld 2.0

Computer use · Higher is better

Source
GPT-5.547.5%
Claude Opus 4.854.8%
GPT-5.6 Luna45.6%
GPT-5.6 Sol62.6%

OSWorld-Verified

Computer use · Higher is better

Source
GPT-5.578.7%
Claude Opus 4.883.4%
GPT-5.6 LunaNot reported
GPT-5.6 SolNot reported

BrowseComp

Tool use · Higher is better

Source
GPT-5.584.4%
Claude Opus 4.884.3%
GPT-5.6 Luna83.3%
GPT-5.6 Sol90.4%

BenchCAD

Computer-aided design · Higher is better

Source
GPT-5.544.4%
Claude Opus 4.827.3%
GPT-5.6 Luna63.1%
GPT-5.6 Sol70.6%

Official benchmark profile

Published scores outside our visual tests

OpenAI sourceMay 2026Source report →
Coding

SWE-bench Pro

58.6%
Agentic coding

Terminal-Bench 2.1

85.6%
Computer use

OSWorld 2.0

47.5%
Full official benchmark table12 scores and peer comparisons
BenchmarkAreaScoreComparison
SWE-bench ProCoding58.6%
GPT-5.558.6%
GPT-5.6 Luna62.7%
Gemini 3.1 Pro54.2%
GPT-5.6 Terra63.4%
Terminal-Bench 2.1Agentic coding85.6%
GPT-5.585.6%
GPT-5.6 Luna84.7%
GPT-5.6 Terra87.4%
Claude Fable 583.1%
OSWorld 2.0Computer use47.5%
GPT-5.547.5%
GPT-5.6 Luna45.6%
GPT-5.6 Terra50.2%
Claude Opus 4.854.8%
OSWorld-VerifiedComputer use78.7%
GPT-5.578.7%
Gemini 3.5 Flash78.4%
Claude Opus 4.883.4%
Claude Fable 585.0%
BrowseCompTool use84.4%
GPT-5.584.4%
Claude Opus 4.884.3%
GPT-5.6 Luna83.3%
Gemini 3.1 Pro85.9%
BenchCADComputer-aided design44.4%
GPT-5.544.4%
Claude Mythos 538.4%
Claude Mythos Preview35.5%
Claude Opus 4.827.3%
BenchCAD with Python toolTool use59.4%
GPT-5.559.4%
Claude Mythos 556.7%
Claude Mythos Preview56.4%
Claude Opus 4.848.1%
GPQA DiamondAcademic reasoning93.6%
GPT-5.593.6%
Claude Mythos 594.1%
GPT-5.6 Terra92.9%
Gemini 3.1 Pro94.3%
FrontierMath Tier 1-3 v2Math85.3%
GPT-5.585.3%
GPT-5.6 Terra84.9%
Claude Fable 587.0%
GPT-5.6 Sol89.0%
FrontierMath Tier 4 v2Math72.5%
GPT-5.572.5%
GPT-5.6 Terra68.3%
GPT-5.6 Sol83.0%
GPT-5.6 Luna58.5%
GDPvalProfessional work84.9%
CyberGymCybersecurity81.8%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
gpt-5.5
Context
1.05M
Max output
128K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
OpenAI API Standard (≤272K input tokens) · USD / 1M tokens
$5 input · $0.5 cached input · $30 output

Above 272K input tokens, input is $10 and output is $45 per 1M tokens for the full session. Regional processing adds 10%.

Benchmark sources

GPT-5.5 rows combine OpenAI’s launch benchmarks with shared GPT-5.6 comparison rows where those make the score directly comparable to newer models.

Introducing GPT-5.5

Evaluation settings

  • OpenAI GPT-5.5 launch score; the GPT-5.6 shared comparison table reports 59.4% on its comparison run.
  • Shared GPT-5.6 comparison table.
  • OpenAI GPT-5.5 launch score.
  • Shared GPT-5.6 comparison table and GPT-5.5 launch result.
  • Vision2Code score without Python tool from the shared GPT-5.6 table.
  • Vision2Code score with Python tool from the shared GPT-5.6 table.
  • OpenAI science reasoning result and shared GPT-5.6 table.