DeepSeek

Tier A · Strong with clear operating rules.

DeepSeek V4.1 Flash

An A-tier open-weight agentic-coding option, based on strong provider-reported results, exceptional API efficiency, and a now-published Superbash visual suite.

API pricing

USD per million tokens · DeepSeek API peak rate
Input
$0.3
Output
$1.2
Cached input
$0.006
Context window
1M
Pricing details and specifications

Try the active visual tests

Open the generated scene to test it, or compare that prompt with another model.

10 runs

City Scroll Journey

Run date: 2026-09-22

Specialist test · evaluated separately

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Low-Poly Tower Defense

Run date: 2026-09-22

Diorama tower defense testing economy balance, wave design, placement rules, and combat readability.

Mechanical Watch Simulator

Run date: 2026-09-22

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Run date: 2026-09-22

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Yingzao Fashi Assembly

Run date: 2026-09-22

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

Editorial verdict

Where this model fits

Back to all benchmarks →

DeepSeek V4.1 Flash is the current DeepSeek API model: a 552B-parameter multimodal MoE with a one-million-token context window. DeepSeek reports results ahead of its retiring V4 Pro on several agentic evaluations, including Terminal-Bench 2.1 and DeepSWE v1.1. A tier reflects that published capability and its unusually low peak API rate. All 16 same-prompt Superbash visual runs are published and stay separate from those first-party scores.

Best for

  • Long-context agentic coding
  • Tool-using and multimodal workflows
  • Cost-sensitive implementation with disciplined review

Watch out

  • Published scores are DeepSeek-reported and depend on the stated harness and maximum reasoning effort
  • Visual-run quality is prompt-specific; open the live artifacts
  • Twelve of the sixteen published runs resumed after a mid-run infrastructure restart, which the suite methodology records
  • Review product polish and regressions before shipping

Why it is ranked here

  1. DeepSeek reports 90.6% on Terminal-Bench 2.1 and 74.2% on DeepSWE v1.1 at maximum reasoning effort, with a one-million-token context window and the documented harnesses.
  2. The release replaces the older V4 Flash API route and temporarily serves legacy V4 Pro requests, so the current roster needs a separate V4.1 entry rather than treating the old V4 records as current evidence.
  3. The 16 same-prompt visual runs are published for direct inspection and are not folded into the provider scores. Thirteen of them vendor their own dependencies or generate assets procedurally, so they open locally without a CDN.

Evidence

2026-09-23

Superbash editorial model ranking

Takeaway: DeepSeek V4.1 Flash is currently placed in Tier A.

The September 2026 editorial roster places DeepSeek V4.1 Flash at rank 7.

Open source →
2026-09-10

DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient

Takeaway: V4.1 Flash is DeepSeek’s current API model, with native image understanding, lower API pricing than the retired V4 Flash route, and a now-published 16-run Superbash visual suite.

The official release confirms the live model ID, compatibility routing, native multimodal support, and the V4 Pro transition. The provider benchmark table is presented separately from the Superbash visual runs, which are a same-prompt evidence layer.

Open source →

How DeepSeek V4.1 Flash compares

2 peers with published comparison scores.

  • DeepSeek V4.1 Flash This model
  • DeepSeek V4 Flash
  • DeepSeek V4 Pro

6 benchmarks · 3 models

GPQA Diamond

STEM reasoning · Higher is better

Source
DeepSeek V4.1 Flash90.9%
DeepSeek V4 Flash89.9%
DeepSeek V4 Pro92.4%

Codeforces

Competitive programming · Higher is better

Source
DeepSeek V4.1 Flash3471
DeepSeek V4 Flash3289
DeepSeek V4 Pro3348

Terminal-Bench 2.1

Agentic coding · Higher is better

Source
DeepSeek V4.1 Flash90.6%
DeepSeek V4 Flash82.7%
DeepSeek V4 Pro87.9%

DeepSWE v1.1

Agentic coding · Higher is better

Source
DeepSeek V4.1 Flash74.2%
DeepSeek V4 Flash54.4%
DeepSeek V4 Pro62.7%

CyberGym

Cybersecurity agent · Higher is better

Source
DeepSeek V4.1 Flash88.1%
DeepSeek V4 Flash76.7%
DeepSeek V4 Pro83.3%

SEC-Bench Pro

Security coding · Higher is better

Source
DeepSeek V4.1 Flash62.8%
DeepSeek V4 Flash30.9%
DeepSeek V4 Pro56.4%

Official benchmark profile

Published scores outside our visual tests

DeepSeek sourceSeptember 2026Source report →
STEM reasoning

GPQA Diamond

90.9%
Competitive programming

Codeforces

3471
Agentic coding

Terminal-Bench 2.1

90.6%
Full official benchmark table10 scores and peer comparisons
BenchmarkAreaScoreComparison
GPQA DiamondSTEM reasoning90.9%
DeepSeek V4.1 Flash90.9%
DeepSeek V4 Flash89.9%
DeepSeek V4 Pro92.4%
CodeforcesCompetitive programming3471
Terminal-Bench 2.1Agentic coding90.6%
DeepSeek V4.1 Flash90.6%
DeepSeek V4 Pro87.9%
DeepSeek V4 Flash82.7%
DeepSWE v1.1Agentic coding74.2%
DeepSeek V4.1 Flash74.2%
DeepSeek V4 Pro62.7%
DeepSeek V4 Flash54.4%
CyberGymCybersecurity agent88.1%
DeepSeek V4.1 Flash88.1%
DeepSeek V4 Pro83.3%
DeepSeek V4 Flash76.7%
SEC-Bench ProSecurity coding62.8%
DeepSeek V4.1 Flash62.8%
DeepSeek V4 Pro56.4%
DeepSeek V4 Flash30.9%
HLE with toolsTool-using reasoning63.9%
DeepSeek V4.1 Flash63.9%
DeepSeek V4 Pro60.0%
DeepSeek V4 Flash51.5%
AutomationBenchAgentic automation54.8%
DeepSeek V4.1 Flash54.8%
DeepSeek V4 Pro43.2%
DeepSeek V4 Flash37.7%
Agent's Last ExamAgentic reasoning31.8%
DeepSeek V4.1 Flash31.8%
DeepSeek V4 Pro25.7%
DeepSeek V4 Flash25.2%
ZeroBench-main with toolsVisual agent49.0%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-21 ↗
Status
Current
API model ID
deepseek-flash
Context
1M
Max output
384K
License
MIT
Access route
DeepSeek API; the live API model ID is deepseek-flash. DeepSeek states that legacy V4 Flash, V4 Flash Vision Exp, and V4 Pro IDs temporarily route here.
Modalities
text input, image input, text output
Hardware
Not publicly verified
DeepSeek API peak rate · USD / 1M tokens
$0.3 input · $0.006 cached input · $1.2 output

Peak hours are Monday-Friday, 01:00-04:00 and 06:00-10:00 UTC. Off-peak rates are half: $0.15 input, $0.003 cached input, and $0.60 output per 1M tokens.

Benchmark sources

DeepSeek V4.1 Flash is a 552B-parameter multimodal MoE model with a 1M-token context window. The provider reports 8B active parameters during prefill and 16B during decoding. Scores below are the provider's maximum-reasoning results; agent benchmarks use different harnesses and are not a universal leaderboard.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

These are DeepSeek-reported results, not Superbash visual benchmark runs. The provider reports maximum reasoning effort (reasoning_effort=100), its named harnesses, and a 1M-token context limit for the code-agent evaluations.

Evaluation settings

  • Pass@1 at maximum reasoning effort; DeepSeek-reported.
  • Rating at maximum reasoning effort; DeepSeek-reported.
  • Pass@1 using DeepSeek Harness Minimal and a 1M-token context; DeepSeek-reported.
  • Resolved rate using mini-SWE and a 1M-token context; DeepSeek-reported.
  • Pass@1 using the Claude Code harness; DeepSeek-reported.
  • Pass@1 using the official scaffold; DeepSeek-reported.
  • Pass@5 using the Claude Code harness and a 512K-token context; DeepSeek-reported.

Local visual-run methodology and validation records