Cursor

Composer 2.5

Version 2.5 benchmark runs across the shared Superbash visual prompts.

API pricing

USD per million tokens · Cursor Standard mode
Input
$0.5
Output
$2.5
Cached input
$0.2
Pricing details and specifications

Try the active visual tests

Open the generated scene to test it, or compare that prompt with another model.

3 runs

Yingzao Fashi Assembly

Run date: 2026-07-22

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

How Composer 2.5 compares

Three relevant peers, side by side.

  • Composer 2.5 This model
  • Composer 2
  • Claude Opus 4.7
  • GPT-5.5

3 benchmarks · 4 models

Terminal-Bench 2.0

Agentic coding · Higher is better

Source
Composer 2.569.3%
Composer 261.7%
Claude Opus 4.769.4%
GPT-5.582.7%

SWE-bench Multilingual

Multilingual coding · Higher is better

Source
Composer 2.579.8%
Composer 273.7%
Claude Opus 4.780.5%
GPT-5.577.8%

CursorBench v3.1 harder tasks

Hard agentic coding · Higher is better

Source
Composer 2.563.2%
Composer 252.2%
Claude Opus 4.7Not reported
GPT-5.5Not reported

Official benchmark profile

How Composer 2.5 performs on coding benchmarks.

Cursor sourceMay 2026Source report →
Agentic coding

Terminal-Bench 2.0

69.3%
Multilingual coding

SWE-bench Multilingual

79.8%
Hard agentic coding

CursorBench v3.1 harder tasks

63.2%
Full official benchmark table3 scores and peer comparisons
BenchmarkAreaScoreComparison
Terminal-Bench 2.0Agentic coding69.3%
Composer 2.569.3%
Claude Opus 4.769.4%
Composer 261.7%
GPT-5.582.7%
SWE-bench MultilingualMultilingual coding79.8%
Composer 2.579.8%
Claude Opus 4.780.5%
GPT-5.577.8%
Composer 273.7%
CursorBench v3.1 harder tasksHard agentic coding63.2%
Composer 2.563.2%
GPT-5.5 xhigh64.3%
Claude Opus 4.7 max64.8%
Composer 252.2%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
Not publicly verified
Context
Not publicly verified
Max output
Not publicly verified
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
Cursor Standard mode · USD / 1M tokens
$0.5 input · $0.2 cached input · $2.5 output

Fast mode costs $3 input, $0.50 cached input, and $15 output per 1M tokens. These are model usage rates inside Cursor, not a standalone public API or subscription price.

Benchmark sources

Cursor continued training Kimi K2.5 for Composer 2.5 and reports gains over Composer 2 across agentic, multilingual, and harder coding tasks. The release table publishes three scores.

Composer 2.5 product page

Evaluation settings

  • Cursor release-table result; no Composer effort label or detailed harness setting is published.
  • Cursor release-table result; no Composer effort label is published.
  • Cursor's own harder-task benchmark; no Composer effort label is published.
CursorBench v3.1 harder tasks
Cursor marks the comparison figures for Opus 4.7 and GPT-5.5 as self-reported.