SpaceXAI

Tier A · Strong with clear operating rules.

Grok 4.7

A-tier price-performance candidate for long-running coding and knowledge work, with a full published Superbash visual suite.

API pricing

USD per million tokens · xAI API (<200K prompt tokens)
Input
$2
Output
$6
Cached input
$0.5
Context window
500K
Pricing details and specifications

Try the active visual tests

Open the generated scene to test it, or compare that prompt with another model.

10 runs

City Scroll Journey

Run date: 2026-09-22

Specialist test · evaluated separately

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Low-Poly Tower Defense

Run date: 2026-09-22

Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Mechanical Watch Simulator

Run date: 2026-09-22

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Run date: 2026-09-22

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly

Run date: 2026-09-22

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

Editorial verdict

Where this model fits

Back to all benchmarks →

Grok 4.7 is SpaceXAI's September flagship for coding, agentic tasks, and knowledge work. The launch data shows meaningful gains over 4.6 on long-running coding, terminal, engineering, and professional-work evaluations at the same standard API price. All 16 same-prompt Superbash visual runs are published and stay separate from those first-party scores.

Best for

  • Long-running coding and agent tasks
  • Professional documents and presentations
  • Tool-heavy workflows with a large context
  • Price-sensitive frontier-model evaluation

Watch out

  • Most current task scores are first-party launch results
  • Visual-run quality is prompt-specific; open the live artifacts
  • xhigh can be slower and more verbose than lower reasoning settings
  • Long-context and regional requests cost more

Why it is ranked here

  1. The xAI launch table reports gains over Grok 4.6 across all seven published task evaluations, including CursorBench 4.0, Terminal-Bench 4.0, and AA Briefcase v1.1.
  2. The public API keeps the $2 input, $0.50 cached-input, and $6 output rates below 200K prompt tokens, making it inexpensive relative to the named launch-table peers.
  3. The 16 same-prompt visual runs are published for direct inspection and are not folded into the provider scores.

Evidence

2026-09-23

Superbash editorial model ranking

Takeaway: Grok 4.7 is currently placed in Tier A.

The September 2026 editorial roster places Grok 4.7 at rank 4.

Open source →
2026-09-21

Introducing Grok 4.7

Takeaway: Grok 4.7 replaces 4.6 in the current editorial roster, with its first-party scores kept separate from the visual benchmark suite.

SpaceXAI publishes the launch configuration and peer table. The Superbash visual runs are a separate, same-prompt evidence layer.

Open source →

How Grok 4.7 compares

Three relevant peers, side by side.

  • Grok 4.7 This model
  • Fable 5.1 max
  • GPT-5.6 Sol max
  • Grok 4.6 high

6 benchmarks · 4 models

CursorBench 4.0

Software engineering · Higher is better

Source
Grok 4.746.3%
Fable 5.1 max51.8%
GPT-5.6 Sol max41.7%
Grok 4.6 high40.4%

DeepSWE v1.1

Software engineering · Higher is better

Source
Grok 4.771.0%
Fable 5.1 max70.0%
GPT-5.6 Sol max72.7%
Grok 4.6 high65.2%

EEBench

Electrical engineering · Higher is better

Source
Grok 4.764.0%
Fable 5.1 max56.4%
GPT-5.6 Sol max39.4%
Grok 4.6 high53.0%

Terminal-Bench 4.0

Multi-hour terminal work · Higher is better

Source
Grok 4.738.0%
Fable 5.1 max57.9%
GPT-5.6 Sol max37.3%
Grok 4.6 high20.3%

Harvey Legal Agent Benchmark

Legal work · Higher is better

Source
Grok 4.719.6%
Fable 5.1 max6.7%
GPT-5.6 Sol max2.5%
Grok 4.6 high15.8%

HealthBench Professional

Clinical reasoning · Higher is better

Source
Grok 4.756.7%
Fable 5.1 max62.1%
GPT-5.6 Sol max60.5%
Grok 4.6 high48.5%

Official benchmark profile

Grok 4.7’s published long-horizon coding and knowledge-work results.

SpaceXAI sourceSeptember 2026SpaceXAI launch report →
Software engineering

CursorBench 4.0

46.3%
Software engineering

DeepSWE v1.1

71.0%
Electrical engineering

EEBench

64.0%
Full official benchmark table9 scores and peer comparisons
BenchmarkAreaScoreComparison
CursorBench 4.0Software engineering46.3%
Grok 4.746.3%
GPT-5.6 Sol max41.7%
Fable 5.1 max51.8%
Grok 4.6 high40.4%
DeepSWE v1.1Software engineering71.0%
Grok 4.771.0%
Fable 5.1 max70.0%
GPT-5.6 Sol max72.7%
Grok 4.6 high65.2%
EEBenchElectrical engineering64.0%
Grok 4.764.0%
Fable 5.1 max56.4%
Grok 4.6 high53.0%
GPT-5.6 Sol max39.4%
AA Briefcase v1.1Multi-hour office work1,657
Terminal-Bench 4.0Multi-hour terminal work38.0%
Grok 4.738.0%
GPT-5.6 Sol max37.3%
Grok 4.6 high20.3%
Fable 5.1 max57.9%
Harvey Legal Agent BenchmarkLegal work19.6%
Grok 4.719.6%
Grok 4.6 high15.8%
Fable 5.1 max6.7%
GPT-5.6 Sol max2.5%
HealthBench ProfessionalClinical reasoning56.7%
Grok 4.756.7%
GPT-5.6 Sol max60.5%
Fable 5.1 max62.1%
Grok 4.6 high48.5%
LatchBio biosafety benchmarkBiosafety62.4%
HackerBench v0.3 risky prompts allowedLower is betterCybersecurity safety3.3%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-22 ↗
Status
Current
API model ID
grok-4.7
Context
500K
Max output
Not publicly verified
License
Not publicly verified
Access route
Public xAI API, Grok Build, Cursor, and supported model gateways. Grok 4.7 Fast is the same model on faster serving at twice the token rates; it is available only in Cursor and Grok Build, not through the public xAI API.
Modalities
text input, image input, text output
Hardware
Not publicly verified
xAI API (<200K prompt tokens) · USD / 1M tokens
$2 input · $0.5 cached input · $6 output

At 200K prompt tokens or more, all tokens in the request use $4 input, $1 cached input, and $12 output per 1M tokens. Grok 4.7 Fast costs twice the standard token rates and is not available on the public xAI API. The US regional endpoint adds a 10% premium.

Benchmark sources

SpaceXAI reports Grok 4.7 on seven coding, terminal, professional-work, health, and engineering evaluations. Superbash keeps these provider results separate from the published same-prompt visual runs.

Introducing Grok 4.7

These are first-party launch-table values. Grok 4.7 uses xhigh effort except for the starred DeepSWE v1.1 result, which uses high effort; peer configurations are retained in the comparison labels.

Evaluation settings

  • SpaceXAI launch-table result; Grok 4.7 xhigh.
  • SpaceXAI launch-table result; Grok 4.7 high rather than xhigh.
  • SpaceXAI launch-table rating; Grok 4.7 xhigh; higher is better.
  • SpaceXAI-reported launch result for benign utility and safe refusal.
  • SpaceXAI-reported launch result; lower is better for risky or malicious prompts allowed.
DeepSWE v1.1
The launch table marks only this Grok 4.7 score with an asterisk for high effort.

Local visual-run methodology and validation records