Qwen3.8-Flash-Next

The smartest model that fits in a desk-side box, and what the long-memory roles run. The model card gives 125B of language with 6B active per token, plus a 51B n-gram embedding table and a 4B MTP head: 176B of weights, 180B packaged, which is the size Hugging Face reports for the checkpoint. It moves like a small model while knowing what a large one knows — 43.0 tok/s on a DGX Spark generating text, flat from 327 to 258,790 prompt tokens, though the same box measures 88.5 tok/s reproducing a file and 27.8 on free-form prose, which is why the task shape is named with every figure here — sees text, images and video, and its cache costs about 13 KB per token in the Q8 configuration we ship, 24 KB at full precision, DERIVED FROM THE ARCHITECTURE with no vendor figure published, which is what makes its full 262k window practical on one box.

The licence, in one paragraph

Qwen Community Licence 1.0 — free to run commercially, on two conditions that stand separately. (1) Naming: above 100M monthly users or $20M monthly revenue you must name the model publicly. (2) A separate licence from Qwen is required to run a Model as a Service or AI Work Assistant business — that condition carries NO user and NO revenue threshold, so it binds long before the naming one. Internal use that does not expose the model or its outputs to a third party is expressly exempt from both, which is where a business running it for its own work sits.

What it needs

One configuration, the one we ship and quote: the Blackwell-optimised NVFP4 build, a Q8 cache at this model's native 262k window, plus the small runtime buffer.

Weights (NVFP4, 4-bit)118.8 GB
Cache at 262k (Q8) — derived3.4 GB
Runtime buffer+0.75 GB
Total123.0 GB

Derived from the architecture; no vendor figure published. The cache line above is not a vendor publication: Alibaba (Qwen) publishes no bytes-per-token figure for this model at all, so it is computed from the architecture's own configuration — its KV heads, its head dimension, and the number of layers that keep a growing cache. The cache, the total, and every machine table on this site are derived from it in turn.

Runs on: NVIDIA DGX Spark (GB10) · 4U4G-TURIN/HPR · NVIDIA DGX Station (GB300) · 4UXGM-TURIN2 DIRECT.

Not this one: NVIDIA RTX 5090 · NVIDIA RTX PRO 6000 Blackwell — smaller than this model at NVFP4 needs.

Not offered with it: 8U16X-TURIN2 B300 — those machines are sold with the other model.

What we change, and why

Updates we have tracked

· Hardware

The four-bit builds arrived: NVIDIA published NVFP4 versions of the open models this site runs

Through September NVIDIA published official NVFP4 builds of Qwen3.8-Flash-Next and GLM-5.3-Flash (2 September), Qwen3.8-27B (4 September), GLM-5.3 (14 September) and DeepSeek-V4.1-Flash (16 September). A four-bit build is about a third of the memory the full-precision weights need — the difference between a datacentre card and the one in the machine you already own — and it is the configuration this site quotes. Take the vendor's own build over a community conversion: the calibration, the recipe and the licence travel with it.

nvidia/DeepSeek-V4.1-Flash-NVFP4 — and the NVFP4 series on NVIDIA's model org

· Multimodal

Qwen3.8-Flash-Next opens its weights, and the licence has a clause aimed at businesses like ours

125B on the label, 176B once the separate n-gram embedding table is counted, and 180B packaged with the 4B MTP head — the model card's own numbers, with only 6B working per token — so it is far faster than its size suggests, and it reads text, images and video. Native context is 262k; the 1M figure needs manual configuration. Read the licence before planning on it: the Qwen Community Licence 1.0 requires a separate agreement from Qwen for anyone running a "Model as a Service or AI Work Assistant business", with no revenue or user threshold. Internal use is explicitly exempt. Our everyday models are Apache 2.0 and unaffected.

Qwen Community Licence 1.0, on the model repo

Register

Nothing is on sale yet. Register for a role and you are told the day it opens — what people register for is what gets built first.