DeepSeek V4-Flash

The value pick: near its bigger sibling's quality for a fraction of the memory, under the most permissive licence in the catalogue, with a 1M-token window and a cache cheap enough to actually use it. Text and code only — where a job needs to read a document, it runs beside a seeing model rather than instead of one.

The licence, in one paragraph

MIT — commercial use, modification and redistribution with no revenue threshold, no territory clause and no naming requirement. The least encumbered licence of any model we stock.

What it needs

One configuration, the one we ship and quote: weights at Q4 (4-bit), a Q8 cache at 128k context — the setup most roles actually run — plus the small runtime buffer.

Weights (Q4, 4-bit)187.4 GB
Cache at 128k (Q8)1.0 GB
Runtime buffer+0.75 GB
Total189.2 GB

Runs on: NVIDIA DGX Station (GB300) · ASRock Rack 8U16X-TURIN2 B300 · ASRock Rack 8U8M-TURIN2 MI355X.

Not this one: Intel Arc Pro B70 · ASRock Radeon AI PRO R9700 · AMD Ryzen AI Max+ 395 station · NVIDIA RTX 5090 · NVIDIA DGX Spark (GB10) · NVIDIA RTX PRO 6000 Blackwell — smaller than this model at Q4 needs.

What we change, and why

Updates we have tracked

No tracked updates yet. Every update to this model is published in the news feed first and listed here in the same pass — one update, one news item, one line here.

Register

Nothing is on sale yet. Register for a role and you are told the day it opens — what people register for is what gets built first.