DeepSeek V4.1 Flash Is WAY Better Than I Expected

Published
Sep 23, 2026
Duration
8:16
Click to load the YouTube player

Why the split-brain design matters

DeepSeek V4.1 Flash uses a causal encoder-decoder setup where the model's 40 layers are split into two groups of 20. The first half reads your prompt and prepares information. For long prompts, most of that can skip the second half, which only processes the recent bits. When the model writes a response, both halves kick in. DeepSeek says this means about 8 billion active parameters while reading and 16 billion while writing.

That matters for coding agents that read a lot between writing steps. The model also uses less memory for context. DeepSeek's own numbers put 4.1 Flash around Opus 5 on a software engineering test called Deep Suite, but behind Opus and GPT-5.6 Sol on Terminal Bench 3. Take those with salt since they're DeepSeek's own evals. The useful part is that long input loops get cheaper.

What it built for us

We ran it through our usual visual tests. The Helm's Deep demo surprised us. Most models make a castle that looks vaguely like Helm's Deep, but this one got the long wall, the surrounding mountains, and the causeway right. That's only the second time we've seen a model nail all three. There were glitches walking up the causeway, but the inside had real detail and felt like a walkable prototype. The code grouped static scenery by material to reduce draw calls, which is a sensible optimization for a browser scene.

The mechanical watch came out clean and polished. The gears are actual 3D shapes, and their motion runs off a shared simulated clock rather than collision checks between gear teeth. That's a good trade-off for an animated explainer. The controls were awkward when we tried to move the camera, but the visuals held up.

The Stormwind trebuchet run also impressed. DeepSeek separated physics from screen refresh and calculated motion in 240 steps per simulated second, with air resistance on the projectile. None of the other runs in this batch matched those three, and a few were just bad, so it's not a universal tool.

The roof test and where this model fits

Our long-running 3D Chinese assembly test usually trips models on the roof. Earlier models would try to manage tens of thousands of individual tile objects and fall apart. DeepSeek V4.1 Flash used shaped surfaces with tile ridges built into the geometry, and it actually completed a roof. First flash model we've seen do that. It's a practical answer to the tile-count problem.

For our current setup, Gemini 6 Astra remains the main builder. We use 4.1 Flash for cheap first passes on expensive ideas, then hand the prototype to Astra to finish. If you had to pick one low-cost model for prototyping, this is the one, but it's not replacing our main workflow.