Kimi K3 good at 3D? We Put It to the Test
Suggested for this guide
Kimi
Start with the current Kimi plan and confirm that it includes the model and coding access this guide needs.
Best for: Claude Code setups and cost-aware model work
Check Kimi plansPartner link. It supports Superbash Learn at no extra cost to you.Kimi K3 cleared all four 3D stress tests — including a poem-to-game conversion most frontier models fail — with only fixable bugs. The claim that Kimi K3 is good at 3D held up. The creator ran four generation tests, and every one of them earned a pass, two of them on the first iteration with no edits.
Helm's Deep — the strongest result
This is the creator's current favorite benchmark, and it's playable at boxmining aai.com if you want to run it yourself. Kimi K3's output was smooth, with clean movement and — critically — it wasn't too dark. Most other models, including GPT 5.6 Soul in its initial runs, render Helm's Deep so dark you can't see what's happening. The creator's bar: "I don't want it like Game of Thrones dark during the long night episode."
The cloud and lighting detail stood out, with light visibly radiating through the scene. There's a landmark indicator in the bottom-right corner that updates as you move — the demo showed "court of the Hornberg," correctly placed where the elf archers stood in the film. The castle and keep sit up top with rocks as a placeholder, but the placeholder is subtle enough that the creator says he wouldn't have noticed it as one.
This was the first iteration, zero changes made. Two flaws: he managed to scale a wall, hit an invisible wall, and got stuck. Next step for this build is adding combat elements to stage the actual Helm's Deep battle. Verdict on the test: "big big pass."
Hogwarts broom flight simulator
An autoflight demo — no player input, the broom flies itself. The flight movement is dynamic with constant motion, and the game has ring logic: it knows to fly through the hoops.
The main failure is scale. The castle is far too small — "that is certainly not Hogwarts." The creator flags size as a persistent weakness across frontier models generally, not just K3. Design quality is cleaner than Grok 4.5 and on par with GPT 5.6 Soul. Lighting and sky detail impressed, though one light source was over-bright to the point he couldn't tell if it was the moon or the sun. The flight path runs through the dark forest. Verdict: "big pass," with "a lot of potential."
Jabberwock — the poem-to-game test
This is the test most frontier models "have a big problem with": converting the poem into a playable game. Kimi K3 did it on the second commit — the only debug needed was inverted A/D strafing, which was fixed.
The result has real game structure: instructions display top-left, quoted directly from the poem, and there's actual player progression. You cross the roots into the tulgey wood, the scene darkens, and the Jabberwock appears as a dragon. Combat works — the player takes damage, there's a health bar, and you can slay it. Post-kill, the sky turns gold and you "galumph" back — galloping, mounted on the Jabberwock's head, which the creator calls a "very cool detail." The loop closes with "Dream the poem again," completing one full game loop.
He did hit another invisible wall during the return leg. Verdict: "big pass. Very impressed."
Low-poly Minecraft-style world
Generated from a single one-shot prompt. Movement quality matched GPT 5.6 Soul, and unlike the comparison, K3 "definitely completed rendering the entire game world." Core mechanics work: you can mine stone and wood, then craft at the workbench.
Two glitches: a floating rock in the world, and a recurring crafting bug where whatever you craft spawns where you're standing — in his words, "it spawned right up my ass." Verdict: happy with it for a first one-shot.
How K3 stacks up
The pattern across all four tests: Kimi K3 matches or beats GPT 5.6 Soul and Grok 4.5 on the things that usually break 3D generation — lighting you can actually see, complete world rendering, and game logic that closes a loop. Its weaknesses are shared frontier-model problems (scale, invisible walls) plus small fixable bugs (inverted strafe, craft spawn position). Nothing failed outright.
Practical takeaway
If you're evaluating models for 3D scene or game generation, K3 belongs in your test rotation — the Helm's Deep demo is public, so you can benchmark it against your own prompts before committing. Expect to spend a commit or two on movement bugs and physics boundaries, but expect the core scene and game logic to land on the first or second shot.

