OpenAI

Tier C · Situational tools.

GPT-5.6 Luna

A delivery fallback the team rarely reaches for while Terra remains accessible.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
gpt-5.6-luna
Context
1.05M
Max output
128K
API price / 1M tokens
$1 input · $6 output

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

9 runs

Superbash commentary

Our take

Back to the full tier list →

GPT-5.6 Luna sits in C tier. The team sees it as a delivery model, but in practice uses Terra on low reasoning for the same boring work and has not found a strong reason to choose Luna first.

Best for

  • Fallback delivery work
  • Low-stakes bounded tasks
  • Cases where Luna is the available OpenAI option

Watch out

  • Little firsthand team usage so far
  • Terra on low reasoning currently overlaps its role
  • Do not route important decisions to it without review

Why it is ranked here

  1. Luna lands in C because its intended delivery role overlaps too heavily with low-reasoning Terra.
  2. The team said it is rarely forced to use Luna because Terra remains available and handles routine work well.
  3. This is a usage-based placement, not a claim that Luna failed a complete benchmark suite.

Evidence and commentary

2026-08-17

Superbash editorial model ranking

Takeaway: GPT-5.6 Luna is currently placed in Tier C.

The August 2026 editorial roster places GPT-5.6 Luna at rank 13.

Open source →

Official benchmark profile

How GPT-5.6 Luna scores beyond our visual tests.

OpenAI describes Luna as the fastest and most cost-efficient GPT-5.6 variant. These rows use the shared GPT-5.6 comparison table so scores line up against Sol, Terra, GPT-5.5, Claude, and Gemini where published.

OpenAI sourceJuly 2026Source report →
Coding

SWE-bench Pro

62.7%
Agentic coding

Terminal-Bench 2.1

84.7%
Computer use

OSWorld 2.0

45.6%
Full official benchmark table9 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
SWE-bench ProCoding62.7%Shared OpenAI GPT-5.6 comparison table.
Claude Fable 580.0%
Claude Opus 4.869.2%
GPT-5.6 Sol64.6%
GPT-5.6 Terra63.4%
GPT-5.6 Luna62.7%
GPT-5.559.4%
Gemini 3.1 Pro54.2%
Terminal-Bench 2.1Agentic coding84.7%Shared OpenAI GPT-5.6 comparison table.
GPT-5.6 Sol88.8%
GPT-5.6 Terra87.4%
GPT-5.585.6%
GPT-5.6 Luna84.7%
Claude Fable 583.1%
Claude Opus 4.878.9%
Gemini 3.1 Pro70.7%
OSWorld 2.0Computer use45.6%High-effort computer-use comparison table.
GPT-5.6 Sol62.6%
Claude Opus 4.854.8%
GPT-5.6 Terra50.2%
GPT-5.547.5%
GPT-5.6 Luna45.6%
BrowseCompTool use83.3%High-effort browsing-agent comparison table.
GPT-5.6 Sol Ultra92.2%
GPT-5.6 Sol90.4%
Claude Mythos 588.0%
Claude Mythos Preview87.9%
GPT-5.6 Terra87.5%
Gemini 3.1 Pro85.9%
GPT-5.584.4%
Claude Opus 4.884.3%
BenchCADComputer-aided design63.1%Vision2Code score without Python tool.
GPT-5.6 Sol70.6%
GPT-5.6 Luna63.1%
GPT-5.6 Terra62.3%
GPT-5.544.4%
Claude Mythos 538.4%
Claude Mythos Preview35.5%
Claude Opus 4.827.3%
BenchCAD with Python toolTool use73.9%Vision2Code score with Python tool.
GPT-5.6 Sol83.4%
GPT-5.6 Terra78.2%
GPT-5.6 Luna73.9%
GPT-5.559.4%
Claude Mythos 556.7%
Claude Mythos Preview56.4%
Claude Opus 4.848.1%
GPQA DiamondAcademic reasoning92.3%Shared reasoning benchmark table.
GPT-5.6 Sol94.6%
Gemini 3.1 Pro94.3%
Claude Mythos 594.1%
GPT-5.593.6%
GPT-5.6 Terra92.9%
Claude Fable 592.6%
GPT-5.6 Luna92.3%
Claude Opus 4.892.0%
FrontierMath Tier 1-3 v2Math78.6%Shared reasoning benchmark table.
GPT-5.6 Sol89.0%
Claude Fable 587.0%
GPT-5.585.3%
GPT-5.6 Terra84.9%
Claude Opus 4.880.0%
GPT-5.6 Luna78.6%
Gemini 3.1 Pro59.6%
FrontierMath Tier 4 v2Math58.5%Shared reasoning benchmark table.
Claude Fable 587.8%
GPT-5.6 Sol83.0%
GPT-5.572.5%
GPT-5.6 Terra68.3%
GPT-5.6 Luna58.5%
Claude Opus 4.856.1%