Xai

Tier S · Ship-to-prod default.

Grok 4.6

Choose Grok 4.6 for fast agent loops. For difficult coding quality alone, treat it as closer to A tier.

Verified model facts

Identity, limits, and pricing

Provider source · checked 2026-09-02 ↗
Status
Current
API model ID
grok-4.6
Context
500K
Max output
Not publicly verified
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
API price / 1M tokens
$2 input · $0.3 cached input · $6 output

Prompts of 200K tokens or more use the provider's higher long-context rate.

Same-prompt visual tests

Inspect the actual runs

Open the generated scene to test it, or compare that prompt with another model.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

Grok 4.6 is fast enough to change multi-step agent work and has become a practical Hermes option. It can still drift during a task, and its coding quality does not consistently match the strongest planners and builders.

Best for

  • Fast agent loops
  • Hermes and parallel agent work
  • Ambitious interactive builds
  • Tasks where latency changes the workflow

Watch out

  • Can still drift halfway through a task
  • Coding quality feels A-tier even when agent usefulness is S-tier
  • Review generated work before shipping

Why it is ranked here

  1. The editorial split is S for agent usefulness and A for coding quality, with speed deciding the overall tier.
  2. The team uses Grok in Hermes because lower latency improves multi-step work.
  3. Sixteen visible Superbash runs support inspection, but this placement is still editorial guidance rather than a universal score.

Evidence

2026-08-27

Superbash editorial model ranking

Takeaway: Grok 4.6 is currently placed in Tier S.

The August 2026 editorial roster places Grok 4.6 at rank 3.

Open source →
2026-08-12

Introducing Grok 4.6

Takeaway: Grok 4.6 is placed in S tier for the current visual-build and agentic-coding evidence set.

SpaceXAI reports the launch table; Superbash keeps the runs and source settings visible rather than flattening them into one score.

Open source →

Official benchmark profile

Grok 4.6’s published agentic coding and knowledge-work results.

xAI reports Grok 4.6 on a mix of composite, knowledge-work, and agentic-coding evaluations. The rows below retain xAI’s exact benchmark labels and only compare the named configurations in the same launch table.

SpaceXAI sourceAugust 2026xAI launch report →
Composite capability

AA Intelligence Index

61
Knowledge work

GDPVal-AA v2

1753
Agentic coding

CursorBench v3.2

69.9%
Full official benchmark table10 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
AA Intelligence IndexComposite capability61xAI launch-table result; the index aggregates nine benchmarks.
GDPVal-AA v2Knowledge work1753xAI launch-table rating; higher is better.
CursorBench v3.2Agentic coding69.9%xAI launch-table result for the named CursorBench configuration.
Fable 5 Max70.5%
Grok 4.669.9%
GPT-5.6 Sol67.2%
Grok 4.566.7%
DeepSWE v1.1Software engineering65.9%xAI launch-table result for its DeepSWE v1.1 setup.
GPT-5.6 Sol73.0%
Fable 5 Max70.0%
Grok 4.665.9%
Grok 4.554.0%
FrontierCode v1.1 (Extended)Long-horizon coding61.3%xAI launch-table result for the Extended variant.
Fable 5 Max63.6%
Grok 4.661.3%
GPT-5.6 Sol60.6%
Grok 4.556.6%
Terminal-Bench v3.0Command-line agents26.0%xAI launch-table result. This newer benchmark is not interchangeable with Terminal-Bench 2.x results elsewhere on the site.
APEX-AgentsAgent reliability57.5%xAI launch-table result for the named APEX-Agents evaluation.
Fable 5 Max59.2%
Grok 4.657.5%
GPT-5.6 Sol56.7%
Grok 4.547.1%
APEX-SWESoftware engineering56.4%xAI launch-table result for the named APEX-SWE evaluation.
Fable 5 Max58.8%
Grok 4.656.4%
Grok 4.553.6%
AA-BriefcaseKnowledge work1577xAI launch-table rating; higher is better.
Harvey LAB (Vals)Legal work15.8%xAI launch-table result for the Harvey LAB Vals evaluation.
Grok 4.615.8%
Grok 4.512.9%
Fable 5 Max11.3%
GPT-5.6 Sol2.5%