DeepSeek

Tier A · Strong with clear operating rules.

DeepSeek V4 Pro

The A-tier value flagship, pending more direct testing of the new Pro release.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
deepseek-v4-pro
Context
1M
Max output
384K
API price / 1M tokens
$0.435 input · $0.87 output

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Superbash commentary

Our take

Back to the full tier list →

DeepSeek V4 Pro moves into A tier on the strength of the new architecture and DeepSeek’s exceptional cost efficiency. The team had not yet completed a direct V4 Pro evaluation when recording, so this is a provisional flagship placement informed partly by hands-on V4 Flash use.

Best for

  • Cost-efficient daily work
  • Reasoning and backend thinking
  • Workloads that need on-demand credit top-ups

Watch out

  • Weak visual polish
  • Do not make it the default for front-end vibe coding

Why it is ranked here

  1. The team explicitly separates the flagship Pro at A from the lightweight Flash at B.
  2. Direct experience with the new Pro release was still limited at recording time, so the placement is not presented as a completed benchmark verdict.
  3. DeepSeek has become a team default because it combines low cost with the flexibility to top up credits immediately.

Evidence and commentary

2026-08-17

Superbash editorial model ranking

Takeaway: DeepSeek V4 Pro is currently placed in Tier A.

The August 2026 editorial roster places DeepSeek V4 Pro at rank 5.

Open source →
2026-06-17

NEW AI Model Tier List for Vibe Coding!

Takeaway: Use DeepSeek V4 Pro for cost-conscious reasoning, then review the UI carefully.

The source video’s front-end caveat remains relevant even with the current A-tier placement.

Open source →

Official benchmark profile

How DeepSeek V4 Pro scores beyond our visual tests.

DeepSeek V4 Pro activates 49B of 1.6T parameters with a 1M-token context. These are the official V4-Pro Max results, which use the largest published reasoning budget and differ from the non-thinking and high settings.

DeepSeek sourceApril 2026Source report →
Knowledge

MMLU-Pro

87.5%
STEM reasoning

GPQA Diamond

90.1%
Code reasoning

LiveCodeBench

93.5%
Full official benchmark table12 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
MMLU-ProKnowledge87.5%Exact match at V4-Pro Max reasoning effort.
Gemini 3.1 Pro High91.0%
Claude Opus 4.6 Max89.1%
DeepSeek V4 Pro87.5%
DeepSeek V4 Flash Max86.2%
GPQA DiamondSTEM reasoning90.1%Pass@1 at V4-Pro Max reasoning effort.
Gemini 3.1 Pro High94.3%
GPT-5.4 xHigh93.0%
Claude Opus 4.6 Max91.3%
DeepSeek V4 Pro90.1%
DeepSeek V4 Flash Max88.1%
LiveCodeBenchCode reasoning93.5%Pass@1 at V4-Pro Max reasoning effort.
DeepSeek V4 Pro93.5%
Gemini 3.1 Pro High91.7%
DeepSeek V4 Flash Max91.6%
Kimi K2.6 Thinking89.6%
Claude Opus 4.6 Max88.8%
HMMT February 2026Math95.2%Pass@1 at V4-Pro Max reasoning effort.
GPT-5.4 xHigh97.7%
Claude Opus 4.6 Max96.2%
DeepSeek V4 Pro95.2%
DeepSeek V4 Flash Max94.8%
IMOAnswerBenchMath reasoning89.8%Pass@1 at V4-Pro Max reasoning effort.
GPT-5.4 xHigh91.4%
DeepSeek V4 Pro89.8%
DeepSeek V4 Flash Max88.4%
Kimi K2.6 Thinking86.0%
MRCR 1MLong context83.5%Mean matching rate at the 1M-token setting.
Claude Opus 4.6 Max92.9%
DeepSeek V4 Pro83.5%
DeepSeek V4 Flash Max78.7%
Gemini 3.1 Pro High76.3%
Terminal-Bench 2.0Agentic coding67.9%Accuracy in the official DeepSeek comparison table.
GPT-5.4 xHigh75.1%
Gemini 3.1 Pro High68.5%
DeepSeek V4 Pro67.9%
Kimi K2.6 Thinking66.7%
DeepSeek V4 Flash Max56.9%
SWE-bench VerifiedAgentic coding80.6%Resolved rate in the official DeepSeek comparison table.
Claude Opus 4.6 Max80.8%
DeepSeek V4 Pro80.6%
Gemini 3.1 Pro High80.6%
DeepSeek V4 Flash Max79.0%
SWE-bench ProAgentic coding55.4%Resolved rate in the official DeepSeek comparison table.
Kimi K2.6 Thinking58.6%
GLM 5.1 Thinking58.4%
GPT-5.4 xHigh57.7%
DeepSeek V4 Pro55.4%
DeepSeek V4 Flash Max52.6%
BrowseCompWeb research83.4%Pass@1 in the official DeepSeek comparison table.
Gemini 3.1 Pro High85.9%
Claude Opus 4.6 Max83.7%
DeepSeek V4 Pro83.4%
DeepSeek V4 Flash Max73.2%
MCP AtlasTool use73.6%Pass@1 in the official DeepSeek comparison table.
Claude Opus 4.6 Max73.8%
DeepSeek V4 Pro73.6%
GLM 5.1 Thinking71.8%
Gemini 3.1 Pro High69.2%
DeepSeek V4 Flash Max69.0%
ToolathlonTool use51.8%Pass@1 in the official DeepSeek comparison table.
GPT-5.4 xHigh54.6%
DeepSeek V4 Pro51.8%
Kimi K2.6 Thinking50.0%
DeepSeek V4 Flash Max47.8%