DeepSeek V4.1-Flash

The value pick: near its bigger sibling's quality for a fraction of the memory, under the most permissive licence in the catalogue, and a native 1M-token window with a cache cheap enough to actually use it. Text and code only — where a job needs to read a document, it runs beside a seeing model rather than instead of one.

The licence, in one paragraph

MIT — commercial use, modification and redistribution with no revenue threshold, no territory clause and no naming requirement. The least encumbered licence of any model we stock.

What it needs

One configuration, the one we ship and quote: the Blackwell-optimised NVFP4 build, a Q8 cache at this model's native 1M window, plus the small runtime buffer.

Weights (NVFP4, 4-bit)187.4 GB
Cache at 1M (Q8)8.1 GB
Runtime buffer+0.75 GB
Total196.3 GB

Runs on: 4U4G-TURIN/HPR · 4UXGM-TURIN2 DIRECT · 8U16X-TURIN2 B300.

Not this one: NVIDIA RTX 5090 · NVIDIA DGX Spark (GB10) · NVIDIA RTX PRO 6000 Blackwell — smaller than this model at NVFP4 needs.

Not offered with it: NVIDIA DGX Station (GB300) — those machines are sold with the other model.

What we change, and why

Updates we have tracked

No tracked updates yet. Every update to this model is published in the news feed first and listed here in the same pass — one update, one news item, one line here.

Register

Nothing is on sale yet. Register for a role and you are told the day it opens — what people register for is what gets built first.