DeepSeek

Tier A · Strong with clear operating rules.

DeepSeek V4 Pro

The A-tier value flagship, pending more direct testing of the new Pro release.

API pricing

USD per million tokens · DeepSeek API peak rate
Input
$1.32
Output
$3.96
Cached input
$0.044
Context window
1M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

DeepSeek V4 Pro moves into A tier on the strength of the new architecture and DeepSeek’s exceptional cost efficiency. The team had not yet completed a direct V4 Pro evaluation when recording, so this is a provisional flagship placement informed partly by hands-on V4 Flash use.

Best for

  • Cost-efficient daily work
  • Reasoning and backend thinking
  • Workloads that need on-demand credit top-ups

Watch out

  • Weak visual polish
  • Do not make it the default for front-end vibe coding

Why it is ranked here

  1. The team explicitly separates the flagship Pro at A from the lightweight Flash at B.
  2. Direct experience with the new Pro release was still limited at recording time, so the placement is not presented as a completed benchmark verdict.
  3. DeepSeek has become a team default because it combines low cost with the flexibility to top up credits immediately.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: DeepSeek V4 Pro is currently placed in Tier A.

The September 2026 editorial roster places DeepSeek V4 Pro at rank 5.

Open source →
2026-06-17

NEW AI Model Tier List for Vibe Coding!

Takeaway: Use DeepSeek V4 Pro for cost-conscious reasoning, then review the UI carefully.

The source video’s front-end caveat remains relevant even with the current A-tier placement.

Open source →

How DeepSeek V4 Pro compares

Three relevant peers, side by side.

  • DeepSeek V4 Pro This model
  • DeepSeek V4 Flash Max
  • Claude Opus 4.6 Max
  • Gemini 3.1 Pro High

6 benchmarks · 4 models

MMLU-Pro

Knowledge · Higher is better

Source
DeepSeek V4 Pro87.5%
DeepSeek V4 Flash Max86.2%
Claude Opus 4.6 Max89.1%
Gemini 3.1 Pro High91.0%

GPQA Diamond

STEM reasoning · Higher is better

Source
DeepSeek V4 Pro90.1%
DeepSeek V4 Flash Max88.1%
Claude Opus 4.6 Max91.3%
Gemini 3.1 Pro High94.3%

LiveCodeBench

Code reasoning · Higher is better

Source
DeepSeek V4 Pro93.5%
DeepSeek V4 Flash Max91.6%
Claude Opus 4.6 Max88.8%
Gemini 3.1 Pro High91.7%

HMMT February 2026

Math · Higher is better

Source
DeepSeek V4 Pro95.2%
DeepSeek V4 Flash Max94.8%
Claude Opus 4.6 Max96.2%
Gemini 3.1 Pro HighNot reported

IMOAnswerBench

Math reasoning · Higher is better

Source
DeepSeek V4 Pro89.8%
DeepSeek V4 Flash Max88.4%
Claude Opus 4.6 MaxNot reported
Gemini 3.1 Pro HighNot reported

MRCR 1M

Long context · Higher is better

Source
DeepSeek V4 Pro83.5%
DeepSeek V4 Flash Max78.7%
Claude Opus 4.6 Max92.9%
Gemini 3.1 Pro High76.3%

Official benchmark profile

Published scores outside our visual tests

DeepSeek sourceApril 2026Source report →
Knowledge

MMLU-Pro

87.5%
STEM reasoning

GPQA Diamond

90.1%
Code reasoning

LiveCodeBench

93.5%
Full official benchmark table12 scores and peer comparisons
BenchmarkAreaScoreComparison
MMLU-ProKnowledge87.5%
DeepSeek V4 Pro87.5%
DeepSeek V4 Flash Max86.2%
Claude Opus 4.6 Max89.1%
Gemini 3.1 Pro High91.0%
GPQA DiamondSTEM reasoning90.1%
DeepSeek V4 Pro90.1%
Claude Opus 4.6 Max91.3%
DeepSeek V4 Flash Max88.1%
GPT-5.4 xHigh93.0%
LiveCodeBenchCode reasoning93.5%
DeepSeek V4 Pro93.5%
Gemini 3.1 Pro High91.7%
DeepSeek V4 Flash Max91.6%
Kimi K2.6 Thinking89.6%
HMMT February 2026Math95.2%
DeepSeek V4 Pro95.2%
DeepSeek V4 Flash Max94.8%
Claude Opus 4.6 Max96.2%
GPT-5.4 xHigh97.7%
IMOAnswerBenchMath reasoning89.8%
DeepSeek V4 Pro89.8%
DeepSeek V4 Flash Max88.4%
GPT-5.4 xHigh91.4%
Kimi K2.6 Thinking86.0%
MRCR 1MLong context83.5%
DeepSeek V4 Pro83.5%
DeepSeek V4 Flash Max78.7%
Gemini 3.1 Pro High76.3%
Claude Opus 4.6 Max92.9%
Terminal-Bench 2.0Agentic coding67.9%
DeepSeek V4 Pro67.9%
Gemini 3.1 Pro High68.5%
Kimi K2.6 Thinking66.7%
GPT-5.4 xHigh75.1%
SWE-bench VerifiedAgentic coding80.6%
DeepSeek V4 Pro80.6%
Gemini 3.1 Pro High80.6%
Claude Opus 4.6 Max80.8%
DeepSeek V4 Flash Max79.0%
SWE-bench ProAgentic coding55.4%
DeepSeek V4 Pro55.4%
GPT-5.4 xHigh57.7%
DeepSeek V4 Flash Max52.6%
GLM 5.1 Thinking58.4%
BrowseCompWeb research83.4%
DeepSeek V4 Pro83.4%
Claude Opus 4.6 Max83.7%
Gemini 3.1 Pro High85.9%
DeepSeek V4 Flash Max73.2%
MCP AtlasTool use73.6%
DeepSeek V4 Pro73.6%
Claude Opus 4.6 Max73.8%
GLM 5.1 Thinking71.8%
Gemini 3.1 Pro High69.2%
ToolathlonTool use51.8%
DeepSeek V4 Pro51.8%
Kimi K2.6 Thinking50.0%
GPT-5.4 xHigh54.6%
DeepSeek V4 Flash Max47.8%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
deepseek-v4-pro
Context
1M
Max output
384K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
DeepSeek API peak rate · USD / 1M tokens
$1.32 input · $0.044 cached input · $3.96 output

Peak hours: Monday–Friday, 01:00–04:00 and 06:00–10:00 UTC. All other hours use half-price off-peak rates: $0.66 input, $0.022 cached input, and $1.98 output per 1M tokens.

Benchmark sources

DeepSeek V4 Pro activates 49B of 1.6T parameters with a 1M-token context. These are the official V4-Pro Max results, which use the largest published reasoning budget and differ from the non-thinking and high settings.

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Evaluation settings

  • Exact match at V4-Pro Max reasoning effort.
  • Pass@1 at V4-Pro Max reasoning effort.
  • Mean matching rate at the 1M-token setting.
  • Accuracy in the official DeepSeek comparison table.
  • Resolved rate in the official DeepSeek comparison table.
  • Pass@1 in the official DeepSeek comparison table.