DeepSeek

Tier B · Specialist or second-choice models.

DeepSeek V4 Flash

The lightweight B-tier cost champion and the team’s new default burner.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
deepseek-v4-flash
Context
1M
Max output
384K
API price / 1M tokens
$0.14 input · $0.28 output

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

11 runs

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Superbash commentary

Our take

Back to the full tier list →

DeepSeek V4 Flash is the standout budget model in the August roster. The team spent about $1 on the complete shared visual benchmark set—versus roughly $25 to $40 for some other models—and now reaches for DeepSeek when it needs cheap, immediately refillable capacity.

Best for

  • Cheap daily implementation
  • Plan-mode brainstorming
  • Scaffolds and bulk rough drafts
  • Workloads that may need instant credit top-ups

Watch out

  • Visible visual bugs appeared in benchmark output
  • The lightweight model remains below Pro on capability
  • Requires review before production

Why it is ranked here

  1. The team’s recorded spend was about $1 for all shared benchmark runs; exact costs will vary with prompts, retries, and provider pricing.
  2. The outputs still contained obvious visual defects, including broken spatial composition in the Helm’s Deep run.
  3. Its B-tier value comes from extreme efficiency and flexible top-ups, not a claim that it matches flagship quality.

Evidence and commentary

2026-08-17

Superbash editorial model ranking

Takeaway: DeepSeek V4 Flash is currently placed in Tier B.

The August 2026 editorial roster places DeepSeek V4 Flash at rank 10.

Open source →
2026-06-17

NEW AI Model Tier List for Vibe Coding!

Takeaway: Use DeepSeek V4 Flash for low-cost planning and scaffolding.

The historical source emphasized hand-holding and review; the current editorial roster raises it to B tier.

Open source →

Official benchmark profile

How DeepSeek V4 Flash scores beyond our visual tests.

DeepSeek V4 Flash activates 13B of 284B parameters with a 1M-token context. These are the official V4-Flash Max results, which use the largest published reasoning budget and differ from the non-thinking and high settings.

DeepSeek sourceApril 2026Source report →
Knowledge

MMLU-Pro

86.2%
STEM reasoning

GPQA Diamond

88.1%
Code reasoning

LiveCodeBench

91.6%
Full official benchmark table12 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
MMLU-ProKnowledge86.2%Exact match at V4-Flash Max reasoning effort.
Gemini 3.1 Pro High91.0%
Claude Opus 4.6 Max89.1%
DeepSeek V4 Pro Max87.5%
GPT-5.4 xHigh87.5%
DeepSeek V4 Flash86.2%
GPQA DiamondSTEM reasoning88.1%Pass@1 at V4-Flash Max reasoning effort.
Gemini 3.1 Pro High94.3%
GPT-5.4 xHigh93.0%
Claude Opus 4.6 Max91.3%
DeepSeek V4 Pro Max90.1%
DeepSeek V4 Flash88.1%
LiveCodeBenchCode reasoning91.6%Pass@1 at V4-Flash Max reasoning effort.
DeepSeek V4 Pro Max93.5%
Gemini 3.1 Pro High91.7%
DeepSeek V4 Flash91.6%
Kimi K2.6 Thinking89.6%
Claude Opus 4.6 Max88.8%
HMMT February 2026Math94.8%Pass@1 at V4-Flash Max reasoning effort.
GPT-5.4 xHigh97.7%
Claude Opus 4.6 Max96.2%
DeepSeek V4 Pro Max95.2%
DeepSeek V4 Flash94.8%
IMOAnswerBenchMath reasoning88.4%Pass@1 at V4-Flash Max reasoning effort.
GPT-5.4 xHigh91.4%
DeepSeek V4 Pro Max89.8%
DeepSeek V4 Flash88.4%
Kimi K2.6 Thinking86.0%
MRCR 1MLong context78.7%Mean matching rate at the 1M-token setting.
Claude Opus 4.6 Max92.9%
DeepSeek V4 Pro Max83.5%
DeepSeek V4 Flash78.7%
Gemini 3.1 Pro High76.3%
Terminal-Bench 2.0Agentic coding56.9%Accuracy in the official DeepSeek comparison table.
GPT-5.4 xHigh75.1%
Gemini 3.1 Pro High68.5%
DeepSeek V4 Pro Max67.9%
Kimi K2.6 Thinking66.7%
DeepSeek V4 Flash56.9%
SWE-bench VerifiedAgentic coding79.0%Resolved rate in the official DeepSeek comparison table.
Claude Opus 4.6 Max80.8%
DeepSeek V4 Pro Max80.6%
Gemini 3.1 Pro High80.6%
DeepSeek V4 Flash79.0%
SWE-bench ProAgentic coding52.6%Resolved rate in the official DeepSeek comparison table.
Kimi K2.6 Thinking58.6%
GLM 5.1 Thinking58.4%
GPT-5.4 xHigh57.7%
DeepSeek V4 Pro Max55.4%
DeepSeek V4 Flash52.6%
BrowseCompWeb research73.2%Pass@1 in the official DeepSeek comparison table.
Gemini 3.1 Pro High85.9%
Claude Opus 4.6 Max83.7%
DeepSeek V4 Pro Max83.4%
DeepSeek V4 Flash73.2%
MCP AtlasTool use69.0%Pass@1 in the official DeepSeek comparison table.
DeepSeek V4 Pro Max73.6%
GLM 5.1 Thinking71.8%
Gemini 3.1 Pro High69.2%
DeepSeek V4 Flash69.0%
ToolathlonTool use47.8%Pass@1 in the official DeepSeek comparison table.
GPT-5.4 xHigh54.6%
DeepSeek V4 Pro Max51.8%
Kimi K2.6 Thinking50.0%
DeepSeek V4 Flash47.8%