OpenAI

GPT-5.6 Terra

Version 5.6 benchmark runs across the shared Superbash visual prompts.

API pricing

USD per million tokens · OpenAI API Standard (≤272K input tokens)
Input
$2
Output
$12
Cached input
$0.2
Context window
1.05M
Pricing details and specifications

Inspect the original visual runs

Open the generated scene to test it, or compare that prompt with another model.

3 runs

Yingzao Fashi Assembly

Run date: 2026-07-10

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

How GPT-5.6 Terra compares

Three relevant peers, side by side.

  • GPT-5.6 Terra This model
  • Claude Opus 4.8
  • GPT-5.5
  • GPT-5.6 Luna

6 benchmarks · 4 models

SWE-bench Pro

Coding · Higher is better

Source
GPT-5.6 Terra63.4%
Claude Opus 4.869.2%
GPT-5.559.4%
GPT-5.6 Luna62.7%

Terminal-Bench 2.1

Agentic coding · Higher is better

Source
GPT-5.6 Terra87.4%
Claude Opus 4.878.9%
GPT-5.585.6%
GPT-5.6 Luna84.7%

OSWorld 2.0

Computer use · Higher is better

Source
GPT-5.6 Terra50.2%
Claude Opus 4.854.8%
GPT-5.547.5%
GPT-5.6 Luna45.6%

BrowseComp

Tool use · Higher is better

Source
GPT-5.6 Terra87.5%
Claude Opus 4.884.3%
GPT-5.584.4%
GPT-5.6 Luna83.3%

BenchCAD

Computer-aided design · Higher is better

Source
GPT-5.6 Terra62.3%
Claude Opus 4.827.3%
GPT-5.544.4%
GPT-5.6 Luna63.1%

BenchCAD with Python tool

Tool use · Higher is better

Source
GPT-5.6 Terra78.2%
Claude Opus 4.848.1%
GPT-5.559.4%
GPT-5.6 Luna73.9%

Official benchmark profile

Published scores outside our visual tests

OpenAI sourceJuly 2026Source report →
Coding

SWE-bench Pro

63.4%
Agentic coding

Terminal-Bench 2.1

87.4%
Computer use

OSWorld 2.0

50.2%
Full official benchmark table9 scores and peer comparisons
BenchmarkAreaScoreComparison
SWE-bench ProCoding63.4%
GPT-5.6 Terra63.4%
GPT-5.6 Luna62.7%
GPT-5.6 Sol64.6%
GPT-5.559.4%
Terminal-Bench 2.1Agentic coding87.4%
GPT-5.6 Terra87.4%
GPT-5.6 Sol88.8%
GPT-5.585.6%
GPT-5.6 Luna84.7%
OSWorld 2.0Computer use50.2%
GPT-5.6 Terra50.2%
GPT-5.547.5%
Claude Opus 4.854.8%
GPT-5.6 Luna45.6%
BrowseCompTool use87.5%
GPT-5.6 Terra87.5%
Claude Mythos Preview87.9%
Claude Mythos 588.0%
Gemini 3.1 Pro85.9%
BenchCADComputer-aided design62.3%
GPT-5.6 Terra62.3%
GPT-5.6 Luna63.1%
GPT-5.6 Sol70.6%
GPT-5.544.4%
BenchCAD with Python toolTool use78.2%
GPT-5.6 Terra78.2%
GPT-5.6 Luna73.9%
GPT-5.6 Sol83.4%
GPT-5.559.4%
GPQA DiamondAcademic reasoning92.9%
GPT-5.6 Terra92.9%
Claude Fable 592.6%
GPT-5.6 Luna92.3%
GPT-5.593.6%
FrontierMath Tier 1-3 v2Math84.9%
GPT-5.6 Terra84.9%
GPT-5.585.3%
Claude Fable 587.0%
GPT-5.6 Sol89.0%
FrontierMath Tier 4 v2Math68.3%
GPT-5.6 Terra68.3%
GPT-5.572.5%
GPT-5.6 Luna58.5%
Claude Opus 4.856.1%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Superseded in this ranking
API model ID
gpt-5.6-terra
Context
1.05M
Max output
128K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
OpenAI API Standard (≤272K input tokens) · USD / 1M tokens
$2 input · $0.2 cached input · $2.5 cache write · $12 output

Above 272K input tokens, the full request costs $4 input, $0.4 cached input, $5 cache write, and $18 output per 1M tokens. Batch, Flex, Fast mode, and regional pricing differ.

Benchmark sources

OpenAI positions Terra as the capable lower-cost GPT-5.6 option. These rows use the shared GPT-5.6 comparison table so the scores line up against Sol, Luna, GPT-5.5, Claude, and Gemini where published.

GPT-5.6: Frontier intelligence that scales with your ambition

Evaluation settings

  • Shared OpenAI GPT-5.6 comparison table.
  • High-effort computer-use comparison table.
  • High-effort browsing-agent comparison table.
  • Vision2Code score without Python tool.
  • Vision2Code score with Python tool.
  • Shared reasoning benchmark table.