Machines we designed, built for you
Five machines, NVIDIA only, one build: the Blackwell-optimised NVFP4 configuration we test and mod. Each section is one machine — what it runs, tokens a second for one person, how many people can work at once, what it draws, and the price ex-VAT.
M means measured on this machine. N means measured on one like it. E means estimated from a measurement. Every source is named at the foot of the page.
It prepares the work. You approve it, you sign it, you send it to the authorities, and you set the price. At every tier.
Some limits come from your other software, not our machine. If your booking system has no way in, the biggest machine we sell cannot open it. We check yours before you buy.
Before you compare the machines, see which work should stay off a cloud API
How it scales
- Start: one DGX Station. It runs the model and your AI team — in your office, under your desk.
- Need more? A second DGX Station. A second machine, linked to the first — the same team, more at once.
- Really a lot? The 8U16X-TURIN2. The top of the range — for volume that has outgrown a desk.
Companion workers. The DGX Station and the 8U16X run the background work — the automations, the reading, the drafting, the checking — while your people decide and sign. Instant, back-and-forth chat is not what these two machines are for.
Which one fits is a sizing question, not a sales one — the five machines below are the full line-up, and we will tell you when a smaller one is enough.
Tier 1 — NVIDIA DGX Spark (GB10)
The entry machine: the whole roster on one desk-side box, at the lowest price we sell.
Runs: Qwen3.8-Flash-Next.
- Memory: 128 GB · 273 GB/s
- Power draw: ≈ 240 W
- Price: €3.607–€6.334 ex-VAT
Qwen3.8-Flash-Next — The 99 GB NVFP4 build (RadixArk, token-pattern table memory-mapped from NVMe) with the 17.7 GB cache pool the published run was configured with — about 122 GB of the 128 GB unified pool. EVERY RATE HERE IS FOR GENERATING TEXT: the same box measures 88.5 tok/s reproducing a file against 27.8 on free-form prose, a 3.2× spread by task shape, so the task is named with every figure on this card. · max context 1M, scaled
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k M | 43 tok/s | 14 by throughput | 20.3 tok/s | 285 tok/s |
| 128k E | 43 tok/s | 10 by cache pool | 24.5 tok/s | 245 tok/s |
| 262k E (recommended) | 43 tok/s | 5 by cache pool | 32.9 tok/s | 165 tok/s |
| 512k* E | 43 tok/s | 2 by cache pool | 41.5 tok/s | 83 tok/s |
| 1M* E | 43 tok/s | 1 by cache pool | 45.4 tok/s | 45 tok/s |
| Prefill | 817 tok/s M — a cold, uncached 128k prompt in 160 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 17.7 GB cache pool would hold 42 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 14 conversations. At 128k (10) and 262k (5) and 512k (2) and 1M (1) the cache pool is the tighter limit, and those rows are the pool's number.
Measured on one DGX Spark (GB10, 128 GB unified) — the 99 GB NVFP4 build with its token-pattern table memory-mapped from NVMe, hybrid mode, MTP 3, thinking off. Two published runs on the same box: a context ladder and a concurrency sweep. THE TASK SHAPE IS NAMED WITH EVERY FIGURE, because on this box it moves the number more than the machine does.
One DGX Spark (GB10), Qwen3.8-Flash-Next as shipped — concurrency sweep, generating code, 2,048-token targets
| Streams | 1 | 2 | 4 | 6 | 8 |
|---|---|---|---|---|---|
| Aggregate tok/s | 45.4 | 82.2 | 131.9 | 171.6 | 218.3 |
| Each stream | 45.4 | 41.2 | 33.3 | 28.8 | 27.6 |
| Prose, aggregate | 31.8 | 56.4 | 67.7 | 92.2 | 110.0 |
Prompt length barely moves it. The context ladder generates at 37.4, 41.6, 41.1, 43.0, 42.1, 42.3, 42.6, 41.4, 40.6 and 44.2 tok/s across ten prompts from 327 to 258,790 tokens — flat, and that flatness is the point of the architecture's hybrid stack. Cold prefill on the same ladder runs 132 tok/s at the smallest prompt and ~1,600–1,900 tok/s once the n-gram table's page cache is warm; warm time-to-first-token stays under 3 s even at 258,790 prompt tokens. Two independent engines measure lower on the same class of box: llama.cpp with the model fully resident at 27.90 ± 0.13 tok/s decode against 817.11 ± 0.62 tok/s prefill, and SGLang at 27.5 tok/s single-stream and 71.7 at 8 concurrent. None of that is a disagreement about the box; it is a disagreement about the stack. AND THE SPREAD IS ALSO THE WORKLOAD: the same machine measures 88.5 tok/s reproducing a file, 46.1 on a targeted bug fix, 32.2 adding a function and 27.8 on free-form prose — a 3.2× spread between the easiest and hardest task shapes on identical hardware. A single headline rate without the task named is not a meaningful number, which is why this card names it.
Sources: pangoleen/qwen3.8-flash-next-dgx-spark — context ladder and sparkDash concurrency sweep · 0xBakeer/qwen38-flash-next-spark — decode by task shape (file reproduction 88.5, prose 27.8) · ncmalan/Qwen3.8-Flash-Next-Single-DGX-Spark — llama.cpp, fully resident, no offload · SGLang cookbook — Qwen3.8-Flash-Next
Measured on this box: 45.4 tok/s single-stream and 218.3 tok/s aggregate at eight streams (27.6 per stream) GENERATING CODE, and 14.0 per stream on free-form prose at the same eight — a 3.2× spread by task shape, so the task is named with every figure. One Spark, NVFP4 weights, the token-pattern table memory-mapped from NVMe, MTP 3, thinking off.
Holds Qwen3.8-Flash-Next at its native window with room to spare. Two can be linked over ConnectX for more throughput — the link, not the memory, is the limit.
Tier 2 — 4U4G-TURIN/HPR
Four RTX PRO 6000 cards on a Threadripper — compute and cache stay on the cards; DeepSeek's Engram tier is cached in host RAM with exact NVMe behind it.
Runs: Qwen3.8-Flash-Next · DeepSeek V4.1-Flash.
- Memory: 384 GB (4 cards pooled) · 7.2 TB/s
- Power draw: ≈ 3.3 kW
- Price: €57.143–€79.832 ex-VAT
Qwen3.8-Flash-Next — The 99 GB NVFP4 build, with the published run's KV cache budget of 2,662,752 tokens — 34.7 GB at the 13 KB-per-token Q8 cache we ship, out of the 384 GB the four cards pool. THAT BUDGET IS WHAT BINDS: 32 streams × (128k + 8k output) asks for 4.46M tokens, so the 32-stream/128k cell was omitted for capacity, not for measurement failure, and the table above inherits the limit. · max context 1M, scaled
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k M | 135 tok/s | 51 by throughput | 20.2 tok/s | 1,032 tok/s |
| 128k E | 129 tok/s | 20 by cache pool | 40.4 tok/s | 808 tok/s |
| 262k E (recommended) | 129 tok/s | 10 by cache pool | 62.7 tok/s | 627 tok/s |
| 512k* E | 129 tok/s | 5 by cache pool | 86.6 tok/s | 433 tok/s |
| 1M* E | 129 tok/s | 2 by cache pool | 112.2 tok/s | 224 tok/s |
| Prefill | 5,450 tok/s M — a cold, uncached 128k prompt in 24 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 34.68 GB cache pool would hold 83 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 51 conversations. At 128k (20) and 262k (10) and 512k (5) and 1M (2) the cache pool is the tighter limit, and those rows are the pool's number.
Measured on this box — four RTX PRO 6000 Blackwell cards, 384 GB pooled — on vLLM, tensor parallel 4, MTP 3, at the model's full 262,144-token window. Decode was measured at zero context and again at 16k, 32k, 64k and 128k; the single-stream rate sheds about 2% across that whole range.
4U4G-TURIN/HPR, Qwen3.8-Flash-Next FP8 — aggregate decode tok/s by concurrent streams, zero context
| Streams | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| Aggregate tok/s | 131.0 | 223.3 | 345.4 | 500.9 | 733.3 | 954.1 |
| Each stream | 131.0 | 111.7 | 86.3 | 62.6 | 45.8 | 29.8 |
Prefill is 5,447 / 5,459 / 5,468 / 5,421 / 5,279 tok/s at 8k / 16k / 32k / 64k / 128k — near 5,450, a 3% decline, effectively flat, which is what a 32k document being a 6-second wait and a 128k one a 24-second wait actually means. The limit on this card is the cache budget, not the cards: vLLM reports 2,662,752 KV tokens on startup, which is ten complete 262k sequences and nineteen 128k requests once an 8,192-token output allowance is reserved for each. That is why the matrix omits the 128k / 32-stream cell — 32 × (131,072 + 8,192) is 4.46M tokens against a 2.66M budget, so it was omitted for capacity, not because the run failed — and why the table above holds 20 conversations at 128k rather than 32. Summed GPU power across the whole matrix was 763 W average and 865 W peak against a 1,200 W card limit.
Sources: cstech.dev — Qwen3.8-Flash-Next FP8 on 4× RTX PRO 6000 Blackwell
DeepSeek V4.1-Flash — ~292 GB of non-Engram weights in VRAM plus ~73 GB of cache — ~365 GB of the 384 GB. The 183 GB Engram tier is cached in host RAM (64 GB) with exact NVMe backing.
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k M | 212 tok/s | 51 by throughput | 20.2 tok/s | 1,030 tok/s |
| 128k E | 208 tok/s | 50 by throughput | 20.2 tok/s | 1,009 tok/s |
| 262k E | 205 tok/s | 49 by throughput | 20.3 tok/s | 993 tok/s |
| 512k E | 202 tok/s | 48 by throughput | 20.3 tok/s | 976 tok/s |
| 1M E (recommended) | 195 tok/s | 40 by cache pool | 23.2 tok/s | 926 tok/s |
| Prefill | 7,065 tok/s M — a cold, uncached 128k prompt in 19 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 73 GB cache pool would hold 1,281 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 48–51 conversations. At 1M (40) the cache pool is the tighter limit, and those rows are the pool's number.
Measured on this box: Qwen 954.1 tok/s at 32 concurrent, 131.0 single-stream at zero context, and 5,450 tok/s prefill flat within 3% from 8k to 128k; DeepSeek 748 at 8. The Qwen run's 2,662,752-token cache budget is what binds at 32 streams × 128k, which is why that cell is absent rather than zero. DeepSeek measured with the cards capped at 275 W; the shipped cards run to 600 W.
The first machine that runs DeepSeek V4.1-Flash, and the cheapest that runs both models side by side. Qwen3.8-Flash-Next runs here too — it is the same four cards.
Tier 3 — NVIDIA DGX Station (GB300)
A desk-side superchip: the fastest single box we sell, for work that reads long documents all day.
Runs: Qwen3.8-Flash-Next · DeepSeek V4.1-Flash.
- Memory: 748 GB · 7.1 TB/s · 252 GB HBM at 7.1 TB/s + 496 GB Grace at 396 GB/s
- Power draw: ≈ 1.5 kW
- Price: €79.832–€105.042 ex-VAT
Qwen3.8-Flash-Next — The 99 GB NVFP4 build (local-inference-lab) in HBM, leaving ~619 GB of the 748 GB coherent pool for cache — 252 GB HBM at 7.1 TB/s and 496 GB Grace at 396 GB/s. The rates on this card are one engine on one Station on SGLang, with MTP3 and ReplaySSM on, which is the configuration we ship. · max context 1M, scaled
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k M | 355 tok/s | 157 by throughput | 20.0 tok/s | 3,142 tok/s |
| 128k E | 355 tok/s | 157 by throughput | 20.0 tok/s | 3,142 tok/s |
| 262k E (recommended) | 355 tok/s | 157 by throughput | 20.0 tok/s | 3,142 tok/s |
| 512k* E | 355 tok/s | 92 by cache pool | 33.0 tok/s | 3,034 tok/s |
| 1M* E | 355 tok/s | 46 by cache pool | 60.9 tok/s | 2,802 tok/s |
| Prefill | 33,020 tok/s M — a cold, uncached 128k prompt in 4.0 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 619 GB cache pool would hold 1,485 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 157 conversations. At 512k (92) and 1M (46) the cache pool is the tighter limit, and those rows are the pool's number.
Measured on one DGX Station (GB300) — the box this card is for — with the local-inference-lab NVFP4 checkpoint on SGLang, one serving engine per Station. The source publishes a full C1–C64 concurrency ladder twice: plain autoregressive, and with MTP3 speculation plus ReplaySSM, which is the configuration this site ships and the ladder the rows above are fitted to.
One DGX Station (GB300), Qwen3.8-Flash-Next — aggregate decode tok/s by concurrent requests
| Configuration | C1 | C16 | C64 |
|---|---|---|---|
| TP1/MTP3 + ReplaySSM — what we ship | 354.6 | 1,733.2 | 2,927.8 |
| TP1/AR — no speculation | 202.1 | 1,883.9 | 4,090.4 |
Speculation is what the shipped configuration does, and it buys per-stream speed: 354.6 tok/s against 202.1 at one request, and 1,733.2 against 1,883.9 at sixteen. At sixty-four it costs aggregate throughput instead — 2,927.8 against 4,090.4 — which is the trade the source measured rather than argued. The Qwen rows on this card are fitted to the speculated ladder, so a reader who turns speculation off is looking at the second row. Cold prefill on the same box runs 33,020–39,338 tok/s from an 8k to a 128k prompt, with 64K measured at 37,884 tok/s (MTP3) and 38,653 (AR); the card quotes the low end of that range. Note what this replaces: until 21 September 2026 this card carried a derived 1,086 tok/s ceiling for Qwen, which the measurement contradicts in both configurations — it was never a measurement of anything.
Sources: catid/dgx_station_benchmarks — Qwen3.8-Flash-Next on one DGX Station GB300
DeepSeek V4.1-Flash — The 476 GiB source checkpoint (NVIDIA's NVFP4 conversion is ~492 GiB): ~225 GB of weights in HBM, the Engram tier in Grace, leaving ~493 GB of the 748 GB coherent pool for cache — 252 GB HBM at 7.1 TB/s and 496 GB Grace at 396 GB/s. The live GB300 campaign allocated a smaller 2.5M-token pool and held 2.2 concurrent full-1M windows.
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k N | 181 tok/s | 62 by throughput | 20.2 tok/s | 1,253 tok/s |
| 128k E | 179 tok/s | 62 by throughput | 20.0 tok/s | 1,242 tok/s |
| 262k E | 173 tok/s | 59 by throughput | 20.2 tok/s | 1,194 tok/s |
| 512k E | 168 tok/s | 57 by throughput | 20.3 tok/s | 1,156 tok/s |
| 1M E (recommended) | 158 tok/s | 54 by throughput | 20.0 tok/s | 1,081 tok/s |
| Prefill | 18,000 tok/s N — a cold, uncached 128k prompt in 7.3 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 493 GB cache pool would hold 8,655 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 54–62 conversations.
Measured on one DGX Station (GB300) — the box this card is for — on vLLM with UVA expert offload and DSpark k=5 speculative decoding, the model as shipped (routed experts MXFP4, Engram and attention FP8), at a 1M context. The machine table above is fitted to v20, the newest sweep on that box; v11 is the anchor the later revisions are read against.
One DGX Station (GB300), DeepSeek V4.1-Flash as shipped — aggregate tok/s by concurrent requests
| Revision | 1 | 2 | 4 | 8 | 12 | 16 |
|---|---|---|---|---|---|---|
| v11 — anchor sweep, 10–11 September 2026 | 82.1 | 119.2 | 164.5 | 237.0 | 251.6 | 286.5 |
| v12 — 12 September 2026 | 89.6 | — | — | — | — | — |
| v18 — 18 September 2026 | 172.0 | — | — | 655 | — | 945 |
| v20 — later revision, same box, 21 September 2026 | 180.6 | — | — | 670 | — | 979 |
Prefill, same campaign: 972,435 prompt tokens in 85.0 s (11,400 tok/s), a 207k prompt in 11.3 s, a 6.5k prompt at 17,900 tok/s. One stream, by content class: shell 150.0, code 145.6, tool-JSON 131, structured 115.8, prose 92.8 — prose is the slowest of the five, and on identical hardware shell and code run 57–62% faster than prose, tool-JSON 41% faster and structured 25% faster. The caveat that governs every figure here: the 510 GB checkpoint does not fit this box's ~250 GiB of HBM, so routed experts and the 189 GiB Engram table are offloaded to Grace LPDDR5X over NVLink-C2C, and each number depends on how many GiB of each are offloaded — that offload split is why the same box reads 82.1 tok/s in one revision and 180.6 eleven days later, and why prose, the slowest class, runs 38% below shell and code and 29% below tool-JSON on identical hardware.
Two machines over 400GbE — not this box
A separate team (catid/dgx_station_benchmarks) measured two DGX Stations joined over 400GbE on this same model: 252.9 tok/s per user at 1 concurrent, 3,401.6 tok/s aggregate at 64 concurrent, and 55,992 prompt tok/s on a single 128k request. Those are two boxes working together, and that team did not attempt this model on one Station — so none of it is a single-box figure, and nothing in the table above is a two-box figure. It is not the older, smaller DeepSeek-V4-Flash either: that checkpoint fits in HBM, so its DGX Station figures run higher and do not apply to this model.
Sources: J-M-Recipes — recipe dgx-station-gb300/deepseek-v4.1-flash-vllm-uva-dspark · al-engr.com — the v11, v12, v18 and v20 write-ups · catid/dgx_station_benchmarks — two DGX Stations over 400GbE
Measured on this class of box: Qwen3.8-Flash-Next on one DGX Station (GB300) — a C1–C64 ladder, 202.1 → 4,090.4 tok/s plain autoregressive and 354.6 → 2,927.8 with MTP3 + ReplaySSM, which is the configuration we ship and the ladder the Qwen rows are fitted to, plus cold prefill of 33,020–39,338 tok/s from an 8k to a 128k prompt. The DeepSeek V4.1-Flash rows are anchored to the sweeps measured on one DGX Station (GB300) — v11 through v20, UVA expert offload and DSpark on, quoted in full under that table. No row for either model is scaled from another box.
Holds Qwen3.8-Flash-Next out to a 1M window, and DeepSeek V4.1-Flash beside it — the only desk-side box that runs both. No rack, no server room.
Tier 4 — 4UXGM-TURIN2 DIRECT
Eight RTX PRO 6000 cards with a basic host — the box exists to feed the GPUs.
Runs: Qwen3.8-Flash-Next · DeepSeek V4.1-Flash.
- Memory: 768 GB (8 cards pooled) · 14.3 TB/s
- Power draw: ≈ 5.3 kW
- Price: €107.563–€147.899 ex-VAT
Qwen3.8-Flash-Next — Two independent copies of the 99 GB NVFP4 build with their own caches — 198 GB of the 768 GB, leaving ~530 GB for cache, which is what lets 64 requests be in flight at once. The published run measures AGGREGATE throughput at 64 concurrent requests and publishes no single-stream run at all, so the one-user row here is that measurement read through this box's own saturation curve. · max context 1M, scaled
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k E | 138 tok/s | 193 by throughput | 20.0 tok/s | 3,864 tok/s |
| 128k E | 134 tok/s | 187 by throughput | 20.0 tok/s | 3,743 tok/s |
| 262k E (recommended) | 130 tok/s | 155 by cache pool | 23.0 tok/s | 3,572 tok/s |
| 512k* E | 124 tok/s | 79 by cache pool | 39.5 tok/s | 3,122 tok/s |
| 1M* E | 118 tok/s | 39 by cache pool | 64.8 tok/s | 2,529 tok/s |
| Prefill | 8,000 tok/s E — a cold, uncached 128k prompt in 16 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 530 GB cache pool would hold 1,271 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 187–193 conversations. At 262k (155) and 512k (79) and 1M (39) the cache pool is the tighter limit, and those rows are the pool's number.
Measured on this box — eight RTX PRO 6000 Blackwell Server Edition cards, 768 GB pooled — on vLLM, run as two independent copies of the model, each spread across four GPUs, with a load balancer in front. The published figure is the box's AGGREGATE throughput at 64 concurrent requests, with a 251k-token prompt in the mix and MTP speculation on. No single-stream run was published.
4UXGM-TURIN2 DIRECT, Qwen3.8-Flash-Next — the published measurement
| Requests in flight | Aggregate tok/s |
|---|---|
| Two TP4 engines, 64 concurrent | 3,340.5 |
What the source actually varies is the draft head, not the width: predicting three tokens ahead measured +72.1% at 1 request, +38.0% at 8 and +17.3% at 16, and −4.6% at 32 — so speculation is a low-load win and a high-load cost on this box, which is why the rows above are fitted to the measured aggregate rather than to a speculated single-stream rate. Reusing sparse-attention work between prediction steps bought +3.7% at 8 requests. The one-user row on this card is not a measurement: it is 3,340.5 tok/s read back through the same saturation curve this site fits to every pair, marked E for that reason. The published run is capped at 64 concurrent requests by the test, not by the machine.
Sources: helix.ml — Qwen3.8-Flash-Next on eight GPUs: what actually helped
DeepSeek V4.1-Flash — The full 476 GiB source checkpoint resident — about 64 GB of weights per card — leaving ~247 GB of the 768 GB for cache. No host memory and no NVMe in the serving path.
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k E | 205 tok/s | 106 by throughput | 20.2 tok/s | 2,137 tok/s |
| 128k E | 200 tok/s | 104 by throughput | 20.0 tok/s | 2,084 tok/s |
| 262k E | 195 tok/s | 101 by throughput | 20.1 tok/s | 2,029 tok/s |
| 512k E | 188 tok/s | 97 by throughput | 20.1 tok/s | 1,953 tok/s |
| 1M E (recommended) | 180 tok/s | 93 by throughput | 20.1 tok/s | 1,866 tok/s |
| Prefill | 8,000 tok/s E — a cold, uncached 128k prompt in 16 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 247 GB cache pool would hold 4,336 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 93–106 conversations.
Measured on this box: 3,340.5 tok/s at 64 concurrent, 251k-token prompt in the mix. Qwen measured with two model copies across the eight cards, 64 requests in flight, a 251k-token prompt among them, MTP speculation on — and no single-stream run published, which is why its one-user rows are marked E. DeepSeek has no decode sweep at this width: its rows are scaled from the measured four-card box, which ran DSpark block 5 with the cards capped at 275 W.
DeepSeek V4.1-Flash with the whole checkpoint resident, and Qwen3.8-Flash-Next run as two independent four-card copies — measured at 3,340.5 tok/s across the box with 64 requests in flight.
Tier 5 — 8U16X-TURIN2 B300
The top of the range: an 8× B300 node for volume that has outgrown a desk.
Runs: DeepSeek V4.1-Flash.
- Memory: 2.10 TB (8 cards pooled) · 64 TB/s · 2.30 TB installed
- Power draw: ≈ 12.0 kW
- Price: €596.639–€963.866 ex-VAT
DeepSeek V4.1-Flash — 361 GB of weights and 1,039 GB of cache (~1.93 TB of the 2.30 TB installed, 2.1 TB usable), plus the 189 GB Engram tier pinned in host RAM.
| Context | One user | At once | Each, all running | Machine total |
|---|---|---|---|---|
| 32k E | 110 tok/s | 2,629 by throughput | 20.0 tok/s | 52,583 tok/s |
| 128k E | 108 tok/s | 2,558 by throughput | 20.0 tok/s | 51,168 tok/s |
| 262k E | 105 tok/s | 2,227 by cache pool | 21.3 tok/s | 47,422 tok/s |
| 512k E | 102 tok/s | 1,140 by cache pool | 30.1 tok/s | 34,268 tok/s |
| 1M E (recommended) | 98 tok/s | 570 by cache pool | 37.9 tok/s | 21,594 tok/s |
| Prefill | 40,000 tok/s E — a cold, uncached 128k prompt in 3.3 s, server loaded and otherwise idle | |||
At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 1,039 GB cache pool would hold 18,240 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 2,558–2,629 conversations. At 262k (2,227) and 512k (1,140) and 1M (570) the cache pool is the tighter limit, and those rows are the pool's number. On this machine that floor is not a measurement and we do not present it as one: no published measurement of DeepSeek V4.1-Flash exists on an 8× B300 or HGX B300 node, and the nearest measured evidence on this exact node class — TheAI Cloud's 8× B300 SXM6 node on out-of-box vLLM — puts a DeepSeek Flash-line model at a 4,722 tok/s aggregate plateau across the whole node, with that team's ramp hard-capped at C=512. Read 2,629 conversations against that plateau and each conversation gets about 1.8 output tokens a second — below the floor this table quotes, and below any interactive threshold. The evidence behind both numbers, and this model's own measured frontier on B300, are set out under this table.
NO PUBLISHED MEASUREMENT OF THIS MODEL ON AN 8× B300 OR HGX B300 NODE EXISTS. Not from NVIDIA, not from DeepSeek, and not in the vLLM or SGLang recipe tables — those mark DeepSeek V4.1-Flash verified on H200, GB200, GB300 and MI355X, and do not mark B300 at all. So the table above is a model, not a measurement: a curve fitted to this node class and read as a floor, at the concurrency where each conversation still gets 20 tok/s. What IS measured on this node class is below, and it is a different, smaller DeepSeek model.
One 8× NVIDIA B300 SXM6 node, ~2.1 TiB, out-of-box vLLM 0.26.0 — A DIFFERENT MODEL, NOT DeepSeek V4.1-Flash — aggregate output tok/s, concurrency ramp doubling to a hard cap of C=512
| Model on the same node class | Active params | Weights | One user (tok/s) | Node plateau (tok/s) | Sustained (tok/s) |
|---|---|---|---|---|---|
| DeepSeek V4-Flash | 304B MoE · ~14B | 155 GB FP4 | 172 | 4,722 | 4,660 |
| DeepSeek V4-Pro | 1.6T MoE · 49B | 805 GB FP4 | 102 | 2,051 | 1,338 |
All output tokens a second, whole node, on the same node class and the same pooled memory as this card, with the concurrency ramp doubling to a hard cap of C=512 — the plateau lives at C=256–512. Across the seven frontier MoEs in that run the plateau band is 1,137–5,769 tok/s, and the team flags that its own single-process load client saturated near 4.7k tok/s, so 4,722 is what that harness could draw rather than a proven ceiling. What it does establish is the order of magnitude: the machine totals this card's curve implies sit well above every plateau measured on this node class, which is why we publish them as modelled. No context-resolved concurrency table for DeepSeek V4.1-Flash exists on any B300 either, so the 32k and 128k rows above can be neither checked nor replaced by a measurement — they are a modelled floor, and they are tagged E for exactly that reason.
InferenceX B300 — per chip, not a node, and not a conversation count
SemiAnalysis's InferenceX publishes the closest measurement of this exact model on B300 silicon: 2- and 4-chip slices of a node (TP2 and TP4), concurrency 1 to 128, on vLLM with prefix-cache reuse, run against AgentX agentic coding traces — input p50 92k tokens, p90 304k, p99 734k; output p50 435. Its unit is tokens per chip per second, and it counts TOTAL tokens, input plus output, much of it served from prefix cache — not output only. So it is a different measurement from ours: it is not a node figure, it cannot be converted into a number of conversations, and it is not a context-resolved concurrency table. Its highest published point is 155,834 tok/s per chip at 64 concurrent; its widest is 128 concurrent, at 125,107 tok/s per chip and 38.8 tok/s per user on a two-chip slice, and 103,823 per chip and 27.0 per user on a four-chip slice. Per-user rates across its published B300 rows run 27.0–292 tok/s, and the 20 tok/s floor this site quotes sits under the slowest of them — which is what a policy floor should do.
Sources: cloud.theai.com — seven open models on one 8× B300 SXM6 node, out-of-box vLLM · inferencex.semianalysis.com — DeepSeek V4.1-Flash on B300 (TP2/TP4, AgentX traces, per chip)
No published measurement of DeepSeek V4.1-Flash exists on an 8× B300 or HGX B300 node — not from NVIDIA, not from DeepSeek, and not in the vLLM or SGLang recipe tables, which mark this model verified on H200, GB200, GB300 and MI355X and do not mark B300 at all. Every DeepSeek figure on this card is modelled, and the measured block above names the nearest evidence that does exist on this node class. Qwen3.8-Flash-Next: NO MEASUREMENT OF THIS MODEL ON AN 8× B300 OR HGX B300 NODE EXISTS, so we publish no node figure for it and derive none — this card is not offered with it, and the model's own page says the same. The one-user row here is this node read back through its own curve to one conversation, not a single-stream measurement: no single-stream run of this model on this node class has been published, and the four-card box reads higher per user partly because its published run had DSpark on. On the ceiling: the largest concurrency ever documented on B300-class silicon is 8,192, on a 72-GPU GB300 NVL72 rack, published as an accuracy evaluation with a gsm8k score rather than a throughput sweep — so 2,048 is a limit of what has been run on a node, not the node's ceiling.
DeepSeek V4.1-Flash sized for many people at once. At this size power and cooling are a design job — see the note on the page. We do not offer Qwen3.8-Flash-Next on this node, and no measurement of that model on an 8× B300 or HGX B300 exists — so there is no node figure for it anywhere on this site, derived or otherwise.
How to read these numbers
- Every machine here is one we build and test. None of it is a demo.
- One user is a single conversation with nothing else running on the machine.
- At once is how many conversations we quote running at that window. A 128k conversation takes four times the memory of a 32k one, so each time the window doubles, half as many fit — unless a throughput limit is the tighter of the two, in which case the table says so.
- Each, all running is what one conversation gets when they are all going at once. Machine total is what the box does across all of them.
- M, N and E say where a number came from: M measured on this machine, N measured on one like it, E estimated from a measurement. Every source is named below.
- Prefill is the machine reading a prompt, cold and uncached. It slows down while other answers are being written.
- Power is the usual draw at the wall. It is a design figure, not a measurement.
- recommended marks the window we ship and quote. Past it the model needs scaling, which costs accuracy on long documents.
- * marks an extended window: it works, and it is the one thing here that makes answers less reliable.
- Prices are ex-VAT, and we quote every machine before you buy. The two Threadripper boxes are estimated from their parts.
Every system is quoted. That covers the machine, the build, the operating system and inference stack, delivery, and your first employee running before handover. Prices are ex-VAT: a VAT-registered business reclaims it, and a cross-border EU purchase is normally invoiced under reverse charge with no VAT at all.