Grok 4.7 Is Cheap. Can It Actually Build?
What we tested
We ran 16 visual coding benchmarks on Cursor with Grok 4.7: games, simulations, a city generator, an interactive watch, and more. The goal was to see how far the model gets us toward something we can actually use.
The builds were promising. Grok made a tower defense game with working towers, gold, and a frost tower that slows enemies. That gives you enough to try a strategy. An Ember Glider sim let you fly through rings, but the landing was broken. Pitch input stopped working on the island, so you needed the separate R key to relaunch. The flight foundation worked, but one advertised control still needed fixing.
That pattern repeated. The interactive watch looked impressive, but the seconds wheel turned 60 times faster than its hand. Grok's physics was wrong in the trebuchet sim: double the stone speed should make air resistance four times stronger, but Grok made it only twice as strong. The predicted shot still matched the actual shot because both used the same wrong calculation. That agreement doesn't prove the physics is correct. If you use a simulator to understand how changing a machine affects a shot, you need to check the relationship being simulated, even when the animation looks convincing.
Astra and Fable 5.1 both did better on those requirements. Astra kept the watch's wheel and hand on the same time scale, and its trebuchet got the air resistance relationship right. Fable's watch failed more basically: a missing color value stopped the drawing code on the first frame.
The cost picture
Cursor's models allowance showed 9% used on 50 million tokens for the full session, all included in our $20 Pro plan. About 83% of those tokens were cache reads, so 50 million doesn't mean 50 million tokens of new code. Cursor labeled roughly 9 million as Grok and the other 41 million as auto, and it doesn't tell us which model handled those auto calls. So we can't credit every token to Grok.
If you're using the API, Grok starts at $2 per million input tokens and $6 for output. Astra and Fable are both $10 for input and $50 for output at standard rates. Grok is much cheaper for fresh input and output, while Fable charges less for cache reads. Your total depends on that mix and how many attempts the job takes. For prototyping, this is a useful result: 9% of our allowance got us 16 builds to explore. Last time we used Fable on this, it blew past 100%.
Where Grok fits
Grok 4.7 is not our main model. We use it the same way as 4.6: for speed, and for workflows that pull data from X. It's not bad at creative coding, but we're still using GPT-6 Astra for the main builds and exploring Jev, the new hot tool that works very differently from how LLMs work. Check the previous video for that.
So if you need to try an idea quickly, Grok 4.7 is a cheap way to get a prototype. Just check the behavior your project depends on, especially timing, physics, and flight controls. The briefs already asked for those things, and the builds missed them.
