DeepSeek V4.1-Flash

The value pick: near its bigger sibling's quality for a fraction of the memory, under the most permissive licence in the catalogue, and a native 1M-token window with a cache cheap enough to actually use it. It activates 8B parameters per token while reading the prompt and 16B while writing the answer — the prefill figure is the one that matters for the long agentic work it is bought for. It reads documents itself: DeepSeek's own vision encoder, measured on its own table at 95.6 on DocVQA and 86.0 on RefCOCO. DeepSeek's figures for it are on the page below, each with the card it came from.

The licence, in one paragraph

MIT — commercial use, modification and redistribution with no revenue threshold, no territory clause and no naming requirement. The least encumbered licence of any model we stock.

What it needs

One configuration, the one we ship and quote: the Blackwell-optimised NVFP4 build, a Q8 cache at this model's native 1M window, plus the small runtime buffer.

Weights (NVFP4, 4-bit)503.7 GB
Cache at 1M (Q8)1.8 GB
Runtime buffer+0.75 GB
Total506.3 GB

Runs on: 4U4G-TURIN/HPR · NVIDIA DGX Station (GB300) · 4UXGM-TURIN2 DIRECT · 8U16X-TURIN2 B300.

Not this one: NVIDIA RTX 5090 · NVIDIA DGX Spark (GB10) · NVIDIA RTX PRO 6000 Blackwell · 4U4G-TURIN/HPR — smaller than this model at NVFP4 needs.

What the model is

Released 10 September 2026 — 244,457 downloads in the last month, 88 files, 510.3 GB of weights. There is no DeepSeek V4.1-Pro: no repository, no card, no endpoint, no benchmark. DeepSeek states that V4-Pro traffic routes to this model "until V4.1-Pro launches".

The architecture, in the vendor's own terms

Causal Encoder-Decoder (CED)
A 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder. The decoder's global KV cache is projected from the final encoder hidden states instead of being derived per decoder layer — which is why the two directions activate different amounts.
Activated per token
8B during prefill (reading the prompt), 16B during decode (writing the answer). The asymmetry is the point: input-heavy agentic work pays the 8B figure, not the 16B one.
SWA Bounded Replay
Missing sliding-window-attention KV states are reconstructed by replaying only the most recent n_win tokens, so SWA KV never has to be persisted to SSD. Persistent KV footprint is about 1/8 of DeepSeek-V4-Flash.
Compressed Sparse Attention 2 (CSA2)
Each attention layer takes one of three static modes — Full, Reindex, Reuse — sharing main KV and indexer K across layers and reusing Top-K sparse-attention indices. A Hierarchical Sparse Indexer bounds the deeper indexer cost independently of context length.
KV cache format
FP4 main KV caching — E2M1, with one E4M3 scale per 16 channels.
Global KV cache
890 bytes per token at the FP4 cache format it caches in — about 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1, and 1,780 bytes per token (0.00178 GB per 1,000) at the Q8 cache we ship. This is the vendor-stated reason the model is cheap to serve at long context.
Engram conditional memory
196B parameters, sparsely accessed by token-based lookup.
DSpark speculative decoding
Semi-autoregressive draft generation with confidence-scheduled verification.
Mixture of experts
1 shared expert plus 384 routed experts per MoE layer, with 6 routed experts active per token.
Vision
A DeepSeek-ViT encoder trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling, plus a two-layer MLP projector. Images and text are processed jointly from the start of pre-training.
Training
From scratch on a 45T-token multimodal corpus. Sparse attention was trained at 64K sequence length; the context was extended to 1M at the 34T-token mark.
Reasoning effort
Continuously controllable from 1 to 100 — a single integer that trades inference cost for accuracy. Every instruct figure on this page is at effort 100.
Input and output
Multimodal input (image and text), text output, contexts up to one million tokens.
Checkpoint
763.2B parameters on disk — 88 files, 510.3 GB.

Source: DeepSeek-V4.1-Flash model card — Hugging Face — retrieved 21 September 2026. MIT. The repository is public and not gated.

The parameter count, in five parts

"How big is it" has five defensible answers for this one artifact, and they are all correct about something different. The page used to give one of them — 763B — as the sum of two others, which does not add up.

Which numberFigureWhat it counts
Activated per token8B prefill / 16B decodeWhat moves per token. Two figures, not one — the CED split is why.
Backbone552BThe language model proper, and the figure DeepSeek headlines.
Engram conditional memory196B (196.6B)Sparse token-lookup tables, resident but sparsely accessed.
The vendor's own two figures, added748B552 + 196. This is the sum the site used to print as "763B", which is arithmetic that does not close.
Actually on disk763.2B88 files, 510.3 GB. The remaining ~15B is the vision encoder and auxiliary tensors, which DeepSeek does not itemise.

One artifact, five defensible answers, and a page that quotes one of them as "the" size is telling you which number it picked rather than what the model is. An independent reconstruction from the checkpoint's config.json reaches the same conclusion.

Source: DeepSeek-V4.1-Flash model card — Hugging Face — retrieved 21 September 2026.

Source: Parameter-count reconstruction from config.json.

Benchmarks — the base model

The pretrained checkpoint, no post-training, all three DeepSeek generations run the same way. This is the cleanest read on the architecture itself — and it is not what anyone runs in production.

Base model — the three DeepSeek generations on one table, run the same way

Benchmark · metric · shotsV4-Flash-BaseV4-Pro-BaseV4.1-Flash-Base
Backbone params 284B1.6T552B
Activated params 13B49B8B prefill / 16B decode
AGIEval · EM · 3–5 shots 83.984.4 (best)83.4
MMLU-Pro · EM · 5 shots 68.373.574.1 (best)
C-Eval · EM · 5 shots 92.193.1 (best)92.1
MultiLoKo · LLM-Judge · 5 shots 42.650.9 (best)45.5
SimpleQA-Verified · EM · 25 shots 30.155.2 (best)42.3
SuperGPQA · EM · 5 shots 46.553.9 (best)53.1
BBH · EM · 3 shots 86.987.5 (best)86.1
BBEH · EM · 1 shots 25.429.8 (best)27.2
DROP · F1 · 1 shots 88.688.7 (best)87.9
HellaSwag · EM · 0 shots 85.788.0 (best)87.2
BigCodeBench · Pass@1 · 3 shots 56.859.260.6 (best)
HumanEval · Pass@1 · 0 shots 69.576.879.4 (best)
GSM8K · EM · 8 shots 90.892.693.0 (best)
MATH · EM · 4 shots 57.464.5 (best)61.1
MGSM · EM · 8 shots 85.7 (best)84.480.2
LongBench-V2 · EM · 1 shots 44.751.5 (best)45.2
MMMU-Pro · EM · 4 shots ——56.5 (best)
CVBench · EM · 4 shots ——77.9 (best)
DocVQA · LLM-Judge · 4 shots ——95.6 (best)
RefCOCO-avg · Acc@0.5 · 0 shots ——86.0 (best)

V4-Pro-Base is the base checkpoint of DeepSeek-V4-Pro-0813 — the last Pro DeepSeek has released. There is no V4.1-Pro base, because there is no V4.1-Pro.

The four vision rows are published for V4.1-Flash only; the two earlier models are not scored on them.

Read the base table for what it is: a pretrained checkpoint with no post-training. It is the honest measure of the architecture, and it is not what a buyer runs.

Source: DeepSeek-V4.1-Flash model card — Hugging Face — retrieved 21 September 2026.

Benchmarks — instruct, against the frontier

The model we ship, at its maximum reasoning effort, against the closed frontier and the strongest open models. Reproduced whole, including the rows it loses.

Instruct model — against the frontier, every model at its maximum reasoning effort

Benchmark · metric · shotsOpus-5.0GPT-5.6 SolK3GLM-5.3DS-V4-ProDS-V4-FlashDS-V4.1-Flash
GPQA Diamond · Pass@1 93.494.1 (best)92.988.192.489.990.9
HLE · Pass@1 56.3 (best)44.543.542.0†42.7†37.8†36.8 (39.1†)
Codeforces · Rating ————334832893471 (best)
MathArena Apex · Pass@1 ——65.6 (best)—65.358.665.6
Terminal-Bench 2.1 · Pass@1 89.188.888.388.287.982.790.6 (best)
Terminal-Bench 3.0 · Pass@1 43.3 (best)34.417.728.311.87.630.0
Terminal-Bench 4.0 · Pass@1 51.8 (best)39.912.637.912.47.031.2
DeepSWE v1.1 · Resolved 74.073.067.566.962.754.474.2 (best)
ProgramBench · Almost@1 37.0 (best)23.017.519.015.5—20.3
NL2Repo-Bench · Score 75.3 (best)56.858.058.061.554.264.0
CyberGym · Pass@1 —84.580.084.583.376.788.1 (best)
SEC-Bench Pro · Pass@1 —74.3 (best)——56.430.962.8
ExploitGym · Pass@1 22.133.7 (best)—15.05.41.815.3
HLE with tools · Pass@1 63.6—59.862.560.051.563.9 (best)
AutomationBench · Pass@1 50.345.846.748.843.237.754.8 (best)
Agent's Last Exam · Pass@1 28.626.727.628.525.725.231.8 (best)
Chartography with tools · Pass@1 84.0 (best)79.968.1———78.9
BabyVision with tools · Pass@1 94.1 (best)88.985.7———89.6
ZeroBench-main with tools · Pass@5 52.053.0 (best)41.0———49.0

† Text-only subset of HLE.

DS-V4-Pro is DeepSeek-V4-Pro-0813. It is not DeepSeek V4.1-Pro — no such model is released: no repository, no card, no endpoint, no benchmark, and DeepSeek routes V4-Pro traffic to V4.1-Flash until it launches.

Settings: temperature 1.0, top_p 0.95. The code-agent benchmarks run in DeepSeek Harness Minimal mode at a 1M-token context window; DeepSWE v1.1 uses mini-SWE, SEC-Bench Pro uses Claude Code, and the visual agent benchmarks use Claude Code at 512k.

MathArena Apex is a tie — Kimi K3 and DeepSeek V4.1-Flash both measure 65.6.

These are the vendor's own runs on the vendor's own harness. They are not the independent measurements plotted on Benchmarks, and the two do not agree. The harness table below is where that gap starts.

Source: DeepSeek-V4.1-Flash model card — Hugging Face — retrieved 21 September 2026.

Where it leads

  • Codeforces rating — 3471, ahead of V4-Pro-0813's 3348 and V4-Flash's 3289.
  • Terminal-Bench 2.1 — 90.6, ahead of Opus-5.0's 89.1.
  • DeepSWE v1.1 — 74.2, ahead of Opus-5.0's 74.0.
  • CyberGym — 88.1, the highest on the table.
  • HLE with tools — 63.9.
  • AutomationBench — 54.8.
  • Agent's Last Exam — 31.8.

Where it trails

  • HLE without tools — 36.8 against Opus-5.0's 56.3. Give it tools or do not ask it this.
  • Terminal-Bench 3.0 — 30.0 against Opus-5.0's 43.3.
  • Terminal-Bench 4.0 — 31.2 against Opus-5.0's 51.8.
  • ProgramBench — 20.3 against 37.0.
  • ExploitGym — 15.3 against 33.7.
  • NL2Repo-Bench — 64.0 against 75.3.

Harness sensitivity — the same model, eight scaffolds

One checkpoint, eight agent harnesses. The harness moves the score by more than the model generation does, which is why a vendor number and an independent number for "the same model on the same test" can be far apart without either being wrong.

The same checkpoint under eight different agent harnesses

Benchmark · metric · shotsClaude CodeCodexOpenCodePimini-SWEDSH MinimalDSH StandardDSH PTC
DeepSWE v1.1 · Resolved 69.865.665.566.274.2 (best)72.670.567.6
Terminal-Bench 2.1 · Pass@1 88.084.185.086.190.390.6 (best)85.885.8

One checkpoint, eight scaffolds — these are not eight models, and DSH is DeepSeek's own harness. N=8 on DeepSWE v1.1, N=3 on Terminal-Bench 2.1, Linux containers, 1M context, max_steps 500, and Terminal-Bench 2.1 run without network access.

The harness moves the score by more than the model generation does: 65.5 to 74.2 on DeepSWE v1.1, from the harness alone. A score is only comparable to another score taken the same way.

Source: DeepSeek-V4.1-Flash model card — Hugging Face — retrieved 21 September 2026.

Independent measurement — Artificial Analysis

Artificial Analysis runs its own evaluations on its own harness, which is a different question from the vendor's tables above — and it is the same index, read on the same day, that the Benchmarks chart now carries.

Intelligence Index — index v4.3.2, retrieved 22 September 2026

Benchmark · metric · shotsIndex v4.3.2 (current, retrieved 22 September 2026)Superseded August index (v4.1.1)
Claude Opus 5.5 (max) — current index leader 58—
Claude Fable 5.1 (max) 53—
GPT-6 Astra (max) 53—
Claude Opus 5 (max) 5163
Muse Spark 1.3 (max) 48—
GPT-5.6 Sol (max) 4761
GLM-5.3 (max) 4560
Kimi K3 (max) 4460
DeepSeek V4.1 Flash (Reasoning, Max Effort) 39—
DeepSeek V4 Pro 0813 (max) 3653
DeepSeek V4 Flash 0731 3552

Source: Artificial Analysis — Intelligence Index v4.3.2, retrieved 22 September 2026, and the model pages.

  • The index changed under this model, and the Benchmarks chart has moved with it. Artificial Analysis announced Intelligence Index v4.3.2 on 7 September 2026 and every score on it moved; the chart now carries v4.3.2 as retrieved from AA's leaderboard on 22 September 2026, and the current column above is read from that same index on that same day. Nothing here is left standing on a superseded index.
  • V4.1-Flash was released on 10 September 2026, so it cannot appear in an August snapshot — and it does not. The superseded column above is AA's own v4.1.1 index of 28 August 2026, kept only for comparison: the 52 in it belongs to V4-Flash-0731 and the 53 to V4-Pro-0813, which reads 36 today. The Benchmarks chart no longer carries either figure — it was re-pulled whole onto the current index rather than half-migrated.
  • The sources disagree on the current figure for this model, and the disagreement is reported rather than resolved: Artificial Analysis publishes 39 on its leaderboard row and on this model's page, and OpenRouter's summary of the same evaluation says 39.5. Half a point apart on one evaluation, so both are stated here — neither is preferred and no average is taken.

No arena rating

No LMArena Elo exists for this model. The text-arena snapshot of 13 September 2026 — 402 models, 8.1M votes — lists nothing newer than deepseek-v4-pro-high-20260813 at 1463±7. There is no arena rating to quote for V4.1-Flash, so none is given here.

Source: LMArena text leaderboard.

Where the vendor and the independent runs disagree

Five places where the published figures did not line up about this model. Four of them still do not, and none is resolved by picking one: both are stated, and where the difference has a cause, the cause is named. The fifth was our own arithmetic, and it is corrected rather than reported.

Terminal-Bench 2.1

Vendor: 90.6 for V4.1-Flash, 87.9 for V4-Pro-0813 — DeepSeek Harness Minimal, 1M context

Independent: 54.68 for V4-Pro-0813 — Vals AI, Terminus 2 harness

Different harness, different number, and the harness is named on both sides. The site plots the independent run on Benchmarks and reproduces the vendor's own here; neither is presented as "the" score.

NL2Repo-Bench

Vendor: 64.0 on the model card, 65.4 in the API changelog

Independent: not independently run

DeepSeek's own two publications disagree by 1.4 points. The table above reproduces the model card figure, because the model card is the source this page is transcribed from.

Parameter count

Vendor: five defensible answers — 8B/16B, 552B, 196B, 748B, 763.2B

Independent: an independent reconstruction from the checkpoint's config.json reaches the same conclusion

Say which number is which rather than quoting one as the size. The table above does exactly that.

Intelligence Index

Vendor: the vendor publishes no Artificial Analysis score

Independent: 39 on index v4.3.2 (Artificial Analysis — its leaderboard row and this model's page), 39.5 on the same evaluation (OpenRouter)

Two summaries of one evaluation, half a point apart. Reported, not averaged.

KV cache per token

Vendor: 890 bytes per token at the FP4 cache it caches in — about 1/4 of V4-Flash and 1/437 of V1

Independent: 0.00178 GB per 1,000 tokens — the same 890 bytes, doubled for the 8-bit cache we ship

Not a disagreement, and it was one until 21 September 2026. The figure this page carried was 18 KB per token, about twenty times the architecture, and it was ours rather than DeepSeek's: an allocated cache pool divided by a scheduler's logical-token capacity (45 GB ÷ 2.5M tokens), not a cache measurement. Corrected to the vendor's published figure, the two agree by construction — 890 bytes at FP4 is 1,780 at Q8, which is what the machine tables' concurrency is derived from.

Where every figure on this page comes from

The rule on this site is that a number without a named source is not a number. These are the sources for everything above, all read on 21 September 2026.

The cache figure the sizing table above prints is the configuration we ship, not the vendor's global one: 1.8 GB at a 1M-token window, a Q8 cache. That is the vendor's published 890 bytes per token at its FP4 cache, doubled for the 8-bit cache we run — 0.00089 GB per 1,000 tokens at FP4, 0.00178 at Q8. Until 21 September 2026 this page carried 18.4 GB and about 18 KB per token, roughly twenty times the architecture; that figure was a cache-pool allocation divided by a scheduler's logical-token capacity rather than a cache measurement, and it was ours, not DeepSeek's.

What we change, and why

Updates we have tracked

No tracked updates yet. Every update to this model is published in the news feed first and listed here in the same pass — one update, one news item, one line here.

Register

Nothing is on sale yet. Register for a role and you are told the day it opens — what people register for is what gets built first.