The smartest model that fits in a desk-side box, and what the long-memory roles run. The model card gives 125B of language with 6B active per token, plus a 51B n-gram embedding table and a 4B MTP head: 176B of weights, 180B packaged, which is the size Hugging Face reports for the checkpoint. It moves like a small model while knowing what a large one knows — 43.0 tok/s on a DGX Spark generating text, flat from 327 to 258,790 prompt tokens, though the same box measures 88.5 tok/s reproducing a file and 27.8 on free-form prose, which is why the task shape is named with every figure here — sees text, images and video, and its cache costs about 13 KB per token in the Q8 configuration we ship, 24 KB at full precision, DERIVED FROM THE ARCHITECTURE with no vendor figure published, which is what makes its full 262k window practical on one box.
The licence, in one paragraph
Qwen Community Licence 1.0 — free to run commercially, on two conditions that stand separately. (1) Naming: above 100M monthly users or $20M monthly revenue you must name the model publicly. (2) A separate licence from Qwen is required to run a Model as a Service or AI Work Assistant business — that condition carries NO user and NO revenue threshold, so it binds long before the naming one. Internal use that does not expose the model or its outputs to a third party is expressly exempt from both, which is where a business running it for its own work sits.
What it needs
One configuration, the one we ship and quote: the Blackwell-optimised NVFP4 build, a Q8 cache at this model's native 262k window, plus the small runtime buffer.
Weights (NVFP4, 4-bit)
118.8 GB
Cache at 262k (Q8) — derived
3.4 GB
Runtime buffer
+0.75 GB
Total
123.0 GB
Derived from the architecture; no vendor figure published. The cache line above is not a vendor publication: Alibaba (Qwen) publishes no bytes-per-token figure for this model at all, so it is computed from the architecture's own configuration — its KV heads, its head dimension, and the number of layers that keep a growing cache. The cache, the total, and every machine table on this site are derived from it in turn.
Runs on: NVIDIA DGX Spark (GB10) · 4U4G-TURIN/HPR · NVIDIA DGX Station (GB300) · 4UXGM-TURIN2 DIRECT.
Not this one: NVIDIA RTX 5090 · NVIDIA RTX PRO 6000 Blackwell — smaller than this model at NVFP4 needs.
Not offered with it: 8U16X-TURIN2 B300 — those machines are sold with the other model.
What we change, and why
We run the Blackwell-optimised NVFP4 build on NVIDIA hardware. NVFP4 is 4-bit in the tensor cores' native format on 5th-generation cards — the model keeps its quality and gains speed and memory, and it is the build our machines ship with.
Quantisation is sized from measured model files, never from nominal bits-per-weight arithmetic — real builds run 10–17% larger than their names suggest, and a card bought on a nominal figure is a card the model does not fit.
The KV cache runs at Q8_0 rather than F16: half the cache for well under 1% quality cost, measured in llama.cpp (768 MiB → 408 MiB for the same cache).
We run it at its native 262k window. Longer windows need RoPE/YaRN scaling — offered, marked, and never presented as equivalent.
Prompts and tool configurations are tuned per role, including the Danish document handling the roles are bought for. What each blueprint ships is the setup we run ourselves.
Vision stays on, because the roles that run this model are the ones that read documents.
Updates we have tracked
· Hardware
The four-bit builds arrived: NVIDIA published NVFP4 versions of the open models this site runs
Through September NVIDIA published official NVFP4 builds of Qwen3.8-Flash-Next and GLM-5.3-Flash (2 September), Qwen3.8-27B (4 September), GLM-5.3 (14 September) and DeepSeek-V4.1-Flash (16 September). A four-bit build is about a third of the memory the full-precision weights need — the difference between a datacentre card and the one in the machine you already own — and it is the configuration this site quotes. Take the vendor's own build over a community conversion: the calibration, the recipe and the licence travel with it.
Qwen3.8-Flash-Next opens its weights, and the licence has a clause aimed at businesses like ours
125B on the label, 176B once the separate n-gram embedding table is counted, and 180B packaged with the 4B MTP head — the model card's own numbers, with only 6B working per token — so it is far faster than its size suggests, and it reads text, images and video. Native context is 262k; the 1M figure needs manual configuration. Read the licence before planning on it: the Qwen Community Licence 1.0 requires a separate agreement from Qwen for anyone running a "Model as a Service or AI Work Assistant business", with no revenue or user threshold. Internal use is explicitly exempt. Our everyday models are Apache 2.0 and unaffected.