Cursor
Composer 2.5
Version 2.5 benchmark runs across the shared Superbash visual prompts.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
- Not publicly verified
- Context
- Not publicly verified
- Max output
- Not publicly verified
- API price / 1M tokens
- $0.5 input · $2.5 output
Standard mode; Cursor also publishes a faster $3 input / $15 output mode.
Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How Composer 2.5 performs on coding benchmarks.
Cursor continued training Kimi K2.5 for Composer 2.5 and reports gains over Composer 2 across agentic, multilingual, and harder coding tasks. The release table publishes three scores.
Terminal-Bench 2.0
69.3%SWE-bench Multilingual
79.8%CursorBench v3.1 harder tasks
63.2%Full official benchmark table3 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| Terminal-Bench 2.0 | Agentic coding | 69.3% | Cursor release-table result; no Composer effort label or detailed harness setting is published. |
| SWE-bench Multilingual | Multilingual coding | 79.8% | Cursor release-table result; no Composer effort label is published. |
| CursorBench v3.1 harder tasks | Hard agentic coding | 63.2% | Cursor's own harder-task benchmark; no Composer effort label is published.Cursor marks the comparison figures for Opus 4.7 and GPT-5.5 as self-reported. |