Xai

Grok 4.6

Version 4.6 benchmark runs across the shared Superbash visual prompts.

API pricing

USD per million tokens · xAI API (<200K input tokens)
Input
$2
Output
$6
Cached input
$0.5
Context window
500K
Pricing details and specifications

Inspect the original visual runs

Open the generated scene to test it, or compare that prompt with another model.

10 runs

City Scroll Journey

Run date: 2026-08-14

Specialist test · evaluated separately

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Low-Poly Tower Defense

Run date: 2026-08-14

Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Mechanical Watch Simulator

Run date: 2026-08-14

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Run date: 2026-08-14

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly

Run date: 2026-08-14

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

How Grok 4.6 compares

Three relevant peers, side by side.

  • Grok 4.6 This model
  • Fable 5 Max
  • Grok 4.5
  • GPT-5.6 Sol

6 benchmarks · 4 models

AA Intelligence Index

Composite capability · Higher is better

Source
Grok 4.661
Fable 5 Max62
Grok 4.556
GPT-5.6 Sol61

GDPVal-AA v2

Knowledge work · Higher is better

Source
Grok 4.61753
Fable 5 Max1741
Grok 4.51526
GPT-5.6 Sol1728

CursorBench v3.2

Agentic coding · Higher is better

Source
Grok 4.669.9%
Fable 5 Max70.5%
Grok 4.566.7%
GPT-5.6 Sol67.2%

DeepSWE v1.1

Software engineering · Higher is better

Source
Grok 4.665.9%
Fable 5 Max70.0%
Grok 4.554.0%
GPT-5.6 Sol73.0%

FrontierCode v1.1 (Extended)

Long-horizon coding · Higher is better

Source
Grok 4.661.3%
Fable 5 Max63.6%
Grok 4.556.6%
GPT-5.6 Sol60.6%

APEX-Agents

Agent reliability · Higher is better

Source
Grok 4.657.5%
Fable 5 Max59.2%
Grok 4.547.1%
GPT-5.6 Sol56.7%

Official benchmark profile

Grok 4.6’s published agentic coding and knowledge-work results.

SpaceXAI sourceAugust 2026xAI launch report →
Composite capability

AA Intelligence Index

61
Knowledge work

GDPVal-AA v2

1753
Agentic coding

CursorBench v3.2

69.9%
Full official benchmark table10 scores and peer comparisons
BenchmarkAreaScoreComparison
AA Intelligence IndexComposite capability61
GDPVal-AA v2Knowledge work1753
CursorBench v3.2Agentic coding69.9%
Grok 4.669.9%
Fable 5 Max70.5%
GPT-5.6 Sol67.2%
Grok 4.566.7%
DeepSWE v1.1Software engineering65.9%
Grok 4.665.9%
Fable 5 Max70.0%
GPT-5.6 Sol73.0%
Grok 4.554.0%
FrontierCode v1.1 (Extended)Long-horizon coding61.3%
Grok 4.661.3%
GPT-5.6 Sol60.6%
Fable 5 Max63.6%
Grok 4.556.6%
Terminal-Bench v3.0Command-line agents26.0%
APEX-AgentsAgent reliability57.5%
Grok 4.657.5%
GPT-5.6 Sol56.7%
Fable 5 Max59.2%
Grok 4.547.1%
APEX-SWESoftware engineering56.4%
Grok 4.656.4%
Fable 5 Max58.8%
Grok 4.553.6%
AA-BriefcaseKnowledge work1577
Harvey LAB (Vals)Legal work15.8%
Grok 4.615.8%
Grok 4.512.9%
Fable 5 Max11.3%
GPT-5.6 Sol2.5%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Superseded in this ranking
API model ID
grok-4.6
Context
500K
Max output
Not publicly verified
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
xAI API (<200K input tokens) · USD / 1M tokens
$2 input · $0.5 cached input · $6 output

At 200K input tokens or more, all tokens in the request use $4 input, $1 cached input, and $12 output per 1M tokens.

Benchmark sources

xAI reports Grok 4.6 on a mix of composite, knowledge-work, and agentic-coding evaluations. The rows below retain xAI’s exact benchmark labels and only compare the named configurations in the same launch table.

Introducing Grok 4.6

These are xAI-reported launch-table values. Compare configurations and effort modes before treating cross-provider rows as like-for-like.

Evaluation settings

  • xAI launch-table result; the index aggregates nine benchmarks.
  • xAI launch-table rating; higher is better.
  • xAI launch-table result for the named CursorBench configuration.
  • xAI launch-table result for its DeepSWE v1.1 setup.
  • xAI launch-table result for the Extended variant.
  • xAI launch-table result. This newer benchmark is not interchangeable with Terminal-Bench 2.x results elsewhere on the site.
  • xAI launch-table result for the named APEX-Agents evaluation.
  • xAI launch-table result for the named APEX-SWE evaluation.
  • xAI launch-table result for the Harvey LAB Vals evaluation.