DeepSeek

DeepSeek V4 Flash

Version 1.0 benchmark runs across the shared Superbash visual prompts.

API pricing

USD per million tokens · DeepSeek API peak rate
Input
$0.44
Output
$1.32
Cached input
$0.014
Context window
1M
Pricing details and specifications

Inspect the original visual runs

Open the generated scene to test it, or compare that prompt with another model.

5 runs

Mechanical Watch Simulator

Run date: 2026-08-04

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Run date: 2026-08-04

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly

Run date: 2026-08-04

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

How DeepSeek V4 Flash compares

Three relevant peers, side by side.

  • DeepSeek V4 Flash This model
  • DeepSeek V4 Pro Max
  • Gemini 3.1 Pro High
  • Claude Opus 4.6 Max

6 benchmarks · 4 models

MMLU-Pro

Knowledge · Higher is better

Source
DeepSeek V4 Flash86.2%
DeepSeek V4 Pro Max87.5%
Gemini 3.1 Pro High91.0%
Claude Opus 4.6 Max89.1%

GPQA Diamond

STEM reasoning · Higher is better

Source
DeepSeek V4 Flash88.1%
DeepSeek V4 Pro Max90.1%
Gemini 3.1 Pro High94.3%
Claude Opus 4.6 Max91.3%

LiveCodeBench

Code reasoning · Higher is better

Source
DeepSeek V4 Flash91.6%
DeepSeek V4 Pro Max93.5%
Gemini 3.1 Pro High91.7%
Claude Opus 4.6 Max88.8%

HMMT February 2026

Math · Higher is better

Source
DeepSeek V4 Flash94.8%
DeepSeek V4 Pro Max95.2%
Gemini 3.1 Pro HighNot reported
Claude Opus 4.6 Max96.2%

IMOAnswerBench

Math reasoning · Higher is better

Source
DeepSeek V4 Flash88.4%
DeepSeek V4 Pro Max89.8%
Gemini 3.1 Pro HighNot reported
Claude Opus 4.6 MaxNot reported

MRCR 1M

Long context · Higher is better

Source
DeepSeek V4 Flash78.7%
DeepSeek V4 Pro Max83.5%
Gemini 3.1 Pro High76.3%
Claude Opus 4.6 Max92.9%

Official benchmark profile

Published scores outside our visual tests

DeepSeek sourceApril 2026Source report →
Knowledge

MMLU-Pro

86.2%
STEM reasoning

GPQA Diamond

88.1%
Code reasoning

LiveCodeBench

91.6%
Full official benchmark table12 scores and peer comparisons
BenchmarkAreaScoreComparison
MMLU-ProKnowledge86.2%
DeepSeek V4 Flash86.2%
DeepSeek V4 Pro Max87.5%
GPT-5.4 xHigh87.5%
Claude Opus 4.6 Max89.1%
GPQA DiamondSTEM reasoning88.1%
DeepSeek V4 Flash88.1%
DeepSeek V4 Pro Max90.1%
Claude Opus 4.6 Max91.3%
GPT-5.4 xHigh93.0%
LiveCodeBenchCode reasoning91.6%
DeepSeek V4 Flash91.6%
Gemini 3.1 Pro High91.7%
DeepSeek V4 Pro Max93.5%
Kimi K2.6 Thinking89.6%
HMMT February 2026Math94.8%
DeepSeek V4 Flash94.8%
DeepSeek V4 Pro Max95.2%
Claude Opus 4.6 Max96.2%
GPT-5.4 xHigh97.7%
IMOAnswerBenchMath reasoning88.4%
DeepSeek V4 Flash88.4%
DeepSeek V4 Pro Max89.8%
Kimi K2.6 Thinking86.0%
GPT-5.4 xHigh91.4%
MRCR 1MLong context78.7%
DeepSeek V4 Flash78.7%
Gemini 3.1 Pro High76.3%
DeepSeek V4 Pro Max83.5%
Claude Opus 4.6 Max92.9%
Terminal-Bench 2.0Agentic coding56.9%
DeepSeek V4 Flash56.9%
Kimi K2.6 Thinking66.7%
DeepSeek V4 Pro Max67.9%
Gemini 3.1 Pro High68.5%
SWE-bench VerifiedAgentic coding79.0%
DeepSeek V4 Flash79.0%
DeepSeek V4 Pro Max80.6%
Gemini 3.1 Pro High80.6%
Claude Opus 4.6 Max80.8%
SWE-bench ProAgentic coding52.6%
DeepSeek V4 Flash52.6%
DeepSeek V4 Pro Max55.4%
GPT-5.4 xHigh57.7%
GLM 5.1 Thinking58.4%
BrowseCompWeb research73.2%
DeepSeek V4 Flash73.2%
DeepSeek V4 Pro Max83.4%
Claude Opus 4.6 Max83.7%
Gemini 3.1 Pro High85.9%
MCP AtlasTool use69.0%
DeepSeek V4 Flash69.0%
Gemini 3.1 Pro High69.2%
GLM 5.1 Thinking71.8%
DeepSeek V4 Pro Max73.6%
ToolathlonTool use47.8%
DeepSeek V4 Flash47.8%
Kimi K2.6 Thinking50.0%
DeepSeek V4 Pro Max51.8%
GPT-5.4 xHigh54.6%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Superseded in this ranking
API model ID
deepseek-v4-flash
Context
1M
Max output
384K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
DeepSeek API peak rate · USD / 1M tokens
$0.44 input · $0.014 cached input · $1.32 output

Peak hours: Monday–Friday, 01:00–04:00 and 06:00–10:00 UTC. All other hours use half-price off-peak rates: $0.22 input, $0.007 cached input, and $0.66 output per 1M tokens.

Benchmark sources

DeepSeek V4 Flash activates 13B of 284B parameters with a 1M-token context. These are the official V4-Flash Max results, which use the largest published reasoning budget and differ from the non-thinking and high settings.

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Evaluation settings

  • Exact match at V4-Flash Max reasoning effort.
  • Pass@1 at V4-Flash Max reasoning effort.
  • Mean matching rate at the 1M-token setting.
  • Accuracy in the official DeepSeek comparison table.
  • Resolved rate in the official DeepSeek comparison table.
  • Pass@1 in the official DeepSeek comparison table.