DeepSeek
DeepSeek V4 Flash
The lightweight B-tier cost champion and the team’s new default burner.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
deepseek-v4-flash- Context
- 1M
- Max output
- 384K
- API price / 1M tokens
- $0.14 input · $0.28 output
Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How DeepSeek V4 Flash scores beyond our visual tests.
DeepSeek V4 Flash activates 13B of 284B parameters with a 1M-token context. These are the official V4-Flash Max results, which use the largest published reasoning budget and differ from the non-thinking and high settings.
MMLU-Pro
86.2%GPQA Diamond
88.1%LiveCodeBench
91.6%Full official benchmark table12 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| MMLU-Pro | Knowledge | 86.2% | Exact match at V4-Flash Max reasoning effort. |
| GPQA Diamond | STEM reasoning | 88.1% | Pass@1 at V4-Flash Max reasoning effort. |
| LiveCodeBench | Code reasoning | 91.6% | Pass@1 at V4-Flash Max reasoning effort. |
| HMMT February 2026 | Math | 94.8% | Pass@1 at V4-Flash Max reasoning effort. |
| IMOAnswerBench | Math reasoning | 88.4% | Pass@1 at V4-Flash Max reasoning effort. |
| MRCR 1M | Long context | 78.7% | Mean matching rate at the 1M-token setting. |
| Terminal-Bench 2.0 | Agentic coding | 56.9% | Accuracy in the official DeepSeek comparison table. |
| SWE-bench Verified | Agentic coding | 79.0% | Resolved rate in the official DeepSeek comparison table. |
| SWE-bench Pro | Agentic coding | 52.6% | Resolved rate in the official DeepSeek comparison table. |
| BrowseComp | Web research | 73.2% | Pass@1 in the official DeepSeek comparison table. |
| MCP Atlas | Tool use | 69.0% | Pass@1 in the official DeepSeek comparison table. |
| Toolathlon | Tool use | 47.8% | Pass@1 in the official DeepSeek comparison table. |
Superbash commentary
Our take
DeepSeek V4 Flash is the standout budget model in the August roster. The team spent about $1 on the complete shared visual benchmark set—versus roughly $25 to $40 for some other models—and now reaches for DeepSeek when it needs cheap, immediately refillable capacity.
Best for
Watch out
Why it is ranked here
Evidence and commentary
Superbash editorial model ranking
Takeaway: DeepSeek V4 Flash is currently placed in Tier B.
The August 2026 editorial roster places DeepSeek V4 Flash at rank 10.
Open source →NEW AI Model Tier List for Vibe Coding!
Takeaway: Use DeepSeek V4 Flash for low-cost planning and scaffolding.
The historical source emphasized hand-holding and review; the current editorial roster raises it to B tier.
Open source →Keep learning