Testing DeepSeek v4 Pro + NEW DeepSeek Harness
What went wrong on the new Harness
We ran DeepSeek V4 Pro through the Helm's Deep benchmark on their brand-new Harness, and the visual results were rough. The city scroll landing page came out vanilla. Ember Glider, the flight sim, was so dark we couldn't see the plane. Jabberw Walk gave us no poem instructions and no monster to fight. Even the 3D Chinese architecture assembly couldn't finish a roof, which the old V4 Pro handled fine.
That last detail is the thing. We know V4 Pro is a stronger model than the old one. So what changed? The Harness itself. The old test of V4 Flash ran on Cursor, not the new Harness, and it did okay. Not amazing, but you could see the plane and the movement worked. Put the stronger Pro model on the new Harness, and everything got darker, emptier, and less playable. We think the Harness is the problem, not the model.
The Harness is a blank slate
DeepSeek's new Harness is still in developer preview. It ships plugin-heavy, which means it's basically empty until you install what you need. If you expect to one-shot a prompt and get a good result on your first try, forget it. You need to spend time installing plugins, modifying settings, and customizing the environment before your agent can do much.
That's a real turnoff if you just want to vibe-code a front end. But it's not a dealbreaker for everyone. If you've already got a bunch of plugins set up, the Harness might work well for you. The key is that the Harness and the model are two different things. You can't blame V4 Pro for visual garbage when the Harness isn't wired up yet. And you can't praise the Harness when the model does all the thinking.
Where V4 Pro actually wins
Even after that ugly benchmark run, we kept V4 Pro on our tier list. Why? Because it's five times cheaper than Kimi K3 for running our YouTube dashboard, and it's much faster.
Our dashboard pipeline syncs YouTube API data, comments, new videos, and shorts. With Kimi K3, a quick sync took two or three minutes. With Miniax M3 before that, same story. DeepSeek V4 Pro does the whole run in about 5 seconds. It's not showing off. It just goes straight to the point, and the cost is tiny.
The reason is structural. Top frontier models charge a lot because they bundle multimodal generation. V4 Pro is strong at text and strong at backend agentic work, but you pay for what you get. For front-end visual work, Cursor is still the better place to run it. For text-heavy, data-heavy backend pipelines, V4 Pro on the Harness is cheap and fast once you've set it up.
So the call is simple. Don't make V4 Pro your default for visual work, especially not on the new Harness. But if you're building backend tools and agents, it's a bargain. Five times cheaper than the alternative, and we've seen it sync our whole dashboard in seconds.
