Machines we designed, built for you

Five machines, NVIDIA only, one build: the Blackwell-optimised NVFP4 configuration we test and mod. Each section is one machine — what it runs, tokens a second for one person, how many people can work at once, what it draws, and the price ex-VAT.

M means measured on this machine. N means measured on one like it. E means estimated from a measurement. Every source is named at the foot of the page.

It prepares the work. You approve it, you sign it, you send it to the authorities, and you set the price. At every tier.

Some limits come from your other software, not our machine. If your booking system has no way in, the biggest machine we sell cannot open it. We check yours before you buy.

Before you compare the machines, see which work should stay off a cloud API

How it scales

Companion workers. The DGX Station and the 8U16X run the background work — the automations, the reading, the drafting, the checking — while your people decide and sign. Instant, back-and-forth chat is not what these two machines are for.

Which one fits is a sizing question, not a sales one — the five machines below are the full line-up, and we will tell you when a smaller one is enough.

Tier 1 — NVIDIA DGX Spark (GB10)

NVIDIA DGX Spark (GB10) — Desktop AI box · 128 GB

The entry machine: the whole roster on one desk-side box, at the lowest price we sell.

Runs: Qwen3.8-Flash-Next.

  • Memory: 128 GB · 273 GB/s
  • Power draw: ≈ 240 W
  • Price: €3.607–€6.334 ex-VAT

Qwen3.8-Flash-Next — The 99 GB NVFP4 build (RadixArk, token-pattern table memory-mapped from NVMe) with the 17.7 GB cache pool the published run was configured with — about 122 GB of the 128 GB unified pool. EVERY RATE HERE IS FOR GENERATING TEXT: the same box measures 88.5 tok/s reproducing a file against 27.8 on free-form prose, a 3.2× spread by task shape, so the task is named with every figure on this card. · max context 1M, scaled

ContextOne userAt onceEach, all runningMachine total
32k M43 tok/s14 by throughput20.3 tok/s285 tok/s
128k E43 tok/s10 by cache pool24.5 tok/s245 tok/s
262k E (recommended)43 tok/s5 by cache pool32.9 tok/s165 tok/s
512k* E43 tok/s2 by cache pool41.5 tok/s83 tok/s
1M* E43 tok/s1 by cache pool45.4 tok/s45 tok/s
Prefill817 tok/s M — a cold, uncached 128k prompt in 160 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 17.7 GB cache pool would hold 42 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 14 conversations. At 128k (10) and 262k (5) and 512k (2) and 1M (1) the cache pool is the tighter limit, and those rows are the pool's number.

Measured on one DGX Spark (GB10, 128 GB unified) — the 99 GB NVFP4 build with its token-pattern table memory-mapped from NVMe, hybrid mode, MTP 3, thinking off. Two published runs on the same box: a context ladder and a concurrency sweep. THE TASK SHAPE IS NAMED WITH EVERY FIGURE, because on this box it moves the number more than the machine does.

One DGX Spark (GB10), Qwen3.8-Flash-Next as shipped — concurrency sweep, generating code, 2,048-token targets

Streams12468
Aggregate tok/s45.482.2131.9171.6218.3
Each stream45.441.233.328.827.6
Prose, aggregate31.856.467.792.2110.0

Prompt length barely moves it. The context ladder generates at 37.4, 41.6, 41.1, 43.0, 42.1, 42.3, 42.6, 41.4, 40.6 and 44.2 tok/s across ten prompts from 327 to 258,790 tokens — flat, and that flatness is the point of the architecture's hybrid stack. Cold prefill on the same ladder runs 132 tok/s at the smallest prompt and ~1,600–1,900 tok/s once the n-gram table's page cache is warm; warm time-to-first-token stays under 3 s even at 258,790 prompt tokens. Two independent engines measure lower on the same class of box: llama.cpp with the model fully resident at 27.90 ± 0.13 tok/s decode against 817.11 ± 0.62 tok/s prefill, and SGLang at 27.5 tok/s single-stream and 71.7 at 8 concurrent. None of that is a disagreement about the box; it is a disagreement about the stack. AND THE SPREAD IS ALSO THE WORKLOAD: the same machine measures 88.5 tok/s reproducing a file, 46.1 on a targeted bug fix, 32.2 adding a function and 27.8 on free-form prose — a 3.2× spread between the easiest and hardest task shapes on identical hardware. A single headline rate without the task named is not a meaningful number, which is why this card names it.

Sources: pangoleen/qwen3.8-flash-next-dgx-spark — context ladder and sparkDash concurrency sweep · 0xBakeer/qwen38-flash-next-spark — decode by task shape (file reproduction 88.5, prose 27.8) · ncmalan/Qwen3.8-Flash-Next-Single-DGX-Spark — llama.cpp, fully resident, no offload · SGLang cookbook — Qwen3.8-Flash-Next

Measured on this box: 45.4 tok/s single-stream and 218.3 tok/s aggregate at eight streams (27.6 per stream) GENERATING CODE, and 14.0 per stream on free-form prose at the same eight — a 3.2× spread by task shape, so the task is named with every figure. One Spark, NVFP4 weights, the token-pattern table memory-mapped from NVMe, MTP 3, thinking off.

Holds Qwen3.8-Flash-Next at its native window with room to spare. Two can be linked over ConnectX for more throughput — the link, not the memory, is the limit.

Tier 2 — 4U4G-TURIN/HPR

4U4G-TURIN/HPR — 4U rack · 4× RTX PRO 6000 + display GPU

Four RTX PRO 6000 cards on a Threadripper — compute and cache stay on the cards; DeepSeek's Engram tier is cached in host RAM with exact NVMe behind it.

Runs: Qwen3.8-Flash-Next · DeepSeek V4.1-Flash.

  • Memory: 384 GB (4 cards pooled) · 7.2 TB/s
  • Power draw: ≈ 3.3 kW
  • Price: €57.143–€79.832 ex-VAT

Qwen3.8-Flash-Next — The 99 GB NVFP4 build, with the published run's KV cache budget of 2,662,752 tokens — 34.7 GB at the 13 KB-per-token Q8 cache we ship, out of the 384 GB the four cards pool. THAT BUDGET IS WHAT BINDS: 32 streams × (128k + 8k output) asks for 4.46M tokens, so the 32-stream/128k cell was omitted for capacity, not for measurement failure, and the table above inherits the limit. · max context 1M, scaled

ContextOne userAt onceEach, all runningMachine total
32k M135 tok/s51 by throughput20.2 tok/s1,032 tok/s
128k E129 tok/s20 by cache pool40.4 tok/s808 tok/s
262k E (recommended)129 tok/s10 by cache pool62.7 tok/s627 tok/s
512k* E129 tok/s5 by cache pool86.6 tok/s433 tok/s
1M* E129 tok/s2 by cache pool112.2 tok/s224 tok/s
Prefill5,450 tok/s M — a cold, uncached 128k prompt in 24 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 34.68 GB cache pool would hold 83 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 51 conversations. At 128k (20) and 262k (10) and 512k (5) and 1M (2) the cache pool is the tighter limit, and those rows are the pool's number.

Measured on this box — four RTX PRO 6000 Blackwell cards, 384 GB pooled — on vLLM, tensor parallel 4, MTP 3, at the model's full 262,144-token window. Decode was measured at zero context and again at 16k, 32k, 64k and 128k; the single-stream rate sheds about 2% across that whole range.

4U4G-TURIN/HPR, Qwen3.8-Flash-Next FP8 — aggregate decode tok/s by concurrent streams, zero context

Streams12481632
Aggregate tok/s131.0223.3345.4500.9733.3954.1
Each stream131.0111.786.362.645.829.8

Prefill is 5,447 / 5,459 / 5,468 / 5,421 / 5,279 tok/s at 8k / 16k / 32k / 64k / 128k — near 5,450, a 3% decline, effectively flat, which is what a 32k document being a 6-second wait and a 128k one a 24-second wait actually means. The limit on this card is the cache budget, not the cards: vLLM reports 2,662,752 KV tokens on startup, which is ten complete 262k sequences and nineteen 128k requests once an 8,192-token output allowance is reserved for each. That is why the matrix omits the 128k / 32-stream cell — 32 × (131,072 + 8,192) is 4.46M tokens against a 2.66M budget, so it was omitted for capacity, not because the run failed — and why the table above holds 20 conversations at 128k rather than 32. Summed GPU power across the whole matrix was 763 W average and 865 W peak against a 1,200 W card limit.

Sources: cstech.dev — Qwen3.8-Flash-Next FP8 on 4× RTX PRO 6000 Blackwell

DeepSeek V4.1-Flash — ~292 GB of non-Engram weights in VRAM plus ~73 GB of cache — ~365 GB of the 384 GB. The 183 GB Engram tier is cached in host RAM (64 GB) with exact NVMe backing.

ContextOne userAt onceEach, all runningMachine total
32k M212 tok/s51 by throughput20.2 tok/s1,030 tok/s
128k E208 tok/s50 by throughput20.2 tok/s1,009 tok/s
262k E205 tok/s49 by throughput20.3 tok/s993 tok/s
512k E202 tok/s48 by throughput20.3 tok/s976 tok/s
1M E (recommended)195 tok/s40 by cache pool23.2 tok/s926 tok/s
Prefill7,065 tok/s M — a cold, uncached 128k prompt in 19 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 73 GB cache pool would hold 1,281 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 48–51 conversations. At 1M (40) the cache pool is the tighter limit, and those rows are the pool's number.

Measured on this box: Qwen 954.1 tok/s at 32 concurrent, 131.0 single-stream at zero context, and 5,450 tok/s prefill flat within 3% from 8k to 128k; DeepSeek 748 at 8. The Qwen run's 2,662,752-token cache budget is what binds at 32 streams × 128k, which is why that cell is absent rather than zero. DeepSeek measured with the cards capped at 275 W; the shipped cards run to 600 W.

The first machine that runs DeepSeek V4.1-Flash, and the cheapest that runs both models side by side. Qwen3.8-Flash-Next runs here too — it is the same four cards.

Tier 3 — NVIDIA DGX Station (GB300)

NVIDIA DGX Station (GB300) — Desktop superchip · 748 GB coherent

A desk-side superchip: the fastest single box we sell, for work that reads long documents all day.

Runs: Qwen3.8-Flash-Next · DeepSeek V4.1-Flash.

  • Memory: 748 GB · 7.1 TB/s · 252 GB HBM at 7.1 TB/s + 496 GB Grace at 396 GB/s
  • Power draw: ≈ 1.5 kW
  • Price: €79.832–€105.042 ex-VAT

Qwen3.8-Flash-Next — The 99 GB NVFP4 build (local-inference-lab) in HBM, leaving ~619 GB of the 748 GB coherent pool for cache — 252 GB HBM at 7.1 TB/s and 496 GB Grace at 396 GB/s. The rates on this card are one engine on one Station on SGLang, with MTP3 and ReplaySSM on, which is the configuration we ship. · max context 1M, scaled

ContextOne userAt onceEach, all runningMachine total
32k M355 tok/s157 by throughput20.0 tok/s3,142 tok/s
128k E355 tok/s157 by throughput20.0 tok/s3,142 tok/s
262k E (recommended)355 tok/s157 by throughput20.0 tok/s3,142 tok/s
512k* E355 tok/s92 by cache pool33.0 tok/s3,034 tok/s
1M* E355 tok/s46 by cache pool60.9 tok/s2,802 tok/s
Prefill33,020 tok/s M — a cold, uncached 128k prompt in 4.0 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 619 GB cache pool would hold 1,485 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 157 conversations. At 512k (92) and 1M (46) the cache pool is the tighter limit, and those rows are the pool's number.

Measured on one DGX Station (GB300) — the box this card is for — with the local-inference-lab NVFP4 checkpoint on SGLang, one serving engine per Station. The source publishes a full C1–C64 concurrency ladder twice: plain autoregressive, and with MTP3 speculation plus ReplaySSM, which is the configuration this site ships and the ladder the rows above are fitted to.

One DGX Station (GB300), Qwen3.8-Flash-Next — aggregate decode tok/s by concurrent requests

ConfigurationC1C16C64
TP1/MTP3 + ReplaySSM — what we ship354.61,733.22,927.8
TP1/AR — no speculation202.11,883.94,090.4

Speculation is what the shipped configuration does, and it buys per-stream speed: 354.6 tok/s against 202.1 at one request, and 1,733.2 against 1,883.9 at sixteen. At sixty-four it costs aggregate throughput instead — 2,927.8 against 4,090.4 — which is the trade the source measured rather than argued. The Qwen rows on this card are fitted to the speculated ladder, so a reader who turns speculation off is looking at the second row. Cold prefill on the same box runs 33,020–39,338 tok/s from an 8k to a 128k prompt, with 64K measured at 37,884 tok/s (MTP3) and 38,653 (AR); the card quotes the low end of that range. Note what this replaces: until 21 September 2026 this card carried a derived 1,086 tok/s ceiling for Qwen, which the measurement contradicts in both configurations — it was never a measurement of anything.

Sources: catid/dgx_station_benchmarks — Qwen3.8-Flash-Next on one DGX Station GB300

DeepSeek V4.1-Flash — The 476 GiB source checkpoint (NVIDIA's NVFP4 conversion is ~492 GiB): ~225 GB of weights in HBM, the Engram tier in Grace, leaving ~493 GB of the 748 GB coherent pool for cache — 252 GB HBM at 7.1 TB/s and 496 GB Grace at 396 GB/s. The live GB300 campaign allocated a smaller 2.5M-token pool and held 2.2 concurrent full-1M windows.

ContextOne userAt onceEach, all runningMachine total
32k N181 tok/s62 by throughput20.2 tok/s1,253 tok/s
128k E179 tok/s62 by throughput20.0 tok/s1,242 tok/s
262k E173 tok/s59 by throughput20.2 tok/s1,194 tok/s
512k E168 tok/s57 by throughput20.3 tok/s1,156 tok/s
1M E (recommended)158 tok/s54 by throughput20.0 tok/s1,081 tok/s
Prefill18,000 tok/s N — a cold, uncached 128k prompt in 7.3 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 493 GB cache pool would hold 8,655 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 54–62 conversations.

Measured on one DGX Station (GB300) — the box this card is for — on vLLM with UVA expert offload and DSpark k=5 speculative decoding, the model as shipped (routed experts MXFP4, Engram and attention FP8), at a 1M context. The machine table above is fitted to v20, the newest sweep on that box; v11 is the anchor the later revisions are read against.

One DGX Station (GB300), DeepSeek V4.1-Flash as shipped — aggregate tok/s by concurrent requests

Revision12481216
v11 — anchor sweep, 10–11 September 202682.1119.2164.5237.0251.6286.5
v12 — 12 September 202689.6—————
v18 — 18 September 2026172.0——655—945
v20 — later revision, same box, 21 September 2026180.6——670—979

Prefill, same campaign: 972,435 prompt tokens in 85.0 s (11,400 tok/s), a 207k prompt in 11.3 s, a 6.5k prompt at 17,900 tok/s. One stream, by content class: shell 150.0, code 145.6, tool-JSON 131, structured 115.8, prose 92.8 — prose is the slowest of the five, and on identical hardware shell and code run 57–62% faster than prose, tool-JSON 41% faster and structured 25% faster. The caveat that governs every figure here: the 510 GB checkpoint does not fit this box's ~250 GiB of HBM, so routed experts and the 189 GiB Engram table are offloaded to Grace LPDDR5X over NVLink-C2C, and each number depends on how many GiB of each are offloaded — that offload split is why the same box reads 82.1 tok/s in one revision and 180.6 eleven days later, and why prose, the slowest class, runs 38% below shell and code and 29% below tool-JSON on identical hardware.

Two machines over 400GbE — not this box

A separate team (catid/dgx_station_benchmarks) measured two DGX Stations joined over 400GbE on this same model: 252.9 tok/s per user at 1 concurrent, 3,401.6 tok/s aggregate at 64 concurrent, and 55,992 prompt tok/s on a single 128k request. Those are two boxes working together, and that team did not attempt this model on one Station — so none of it is a single-box figure, and nothing in the table above is a two-box figure. It is not the older, smaller DeepSeek-V4-Flash either: that checkpoint fits in HBM, so its DGX Station figures run higher and do not apply to this model.

Sources: J-M-Recipes — recipe dgx-station-gb300/deepseek-v4.1-flash-vllm-uva-dspark · al-engr.com — the v11, v12, v18 and v20 write-ups · catid/dgx_station_benchmarks — two DGX Stations over 400GbE

Measured on this class of box: Qwen3.8-Flash-Next on one DGX Station (GB300) — a C1–C64 ladder, 202.1 → 4,090.4 tok/s plain autoregressive and 354.6 → 2,927.8 with MTP3 + ReplaySSM, which is the configuration we ship and the ladder the Qwen rows are fitted to, plus cold prefill of 33,020–39,338 tok/s from an 8k to a 128k prompt. The DeepSeek V4.1-Flash rows are anchored to the sweeps measured on one DGX Station (GB300) — v11 through v20, UVA expert offload and DSpark on, quoted in full under that table. No row for either model is scaled from another box.

Holds Qwen3.8-Flash-Next out to a 1M window, and DeepSeek V4.1-Flash beside it — the only desk-side box that runs both. No rack, no server room.

Tier 4 — 4UXGM-TURIN2 DIRECT

4UXGM-TURIN2 DIRECT — 4U rack · 8× RTX PRO 6000

Eight RTX PRO 6000 cards with a basic host — the box exists to feed the GPUs.

Runs: Qwen3.8-Flash-Next · DeepSeek V4.1-Flash.

  • Memory: 768 GB (8 cards pooled) · 14.3 TB/s
  • Power draw: ≈ 5.3 kW
  • Price: €107.563–€147.899 ex-VAT

Qwen3.8-Flash-Next — Two independent copies of the 99 GB NVFP4 build with their own caches — 198 GB of the 768 GB, leaving ~530 GB for cache, which is what lets 64 requests be in flight at once. The published run measures AGGREGATE throughput at 64 concurrent requests and publishes no single-stream run at all, so the one-user row here is that measurement read through this box's own saturation curve. · max context 1M, scaled

ContextOne userAt onceEach, all runningMachine total
32k E138 tok/s193 by throughput20.0 tok/s3,864 tok/s
128k E134 tok/s187 by throughput20.0 tok/s3,743 tok/s
262k E (recommended)130 tok/s155 by cache pool23.0 tok/s3,572 tok/s
512k* E124 tok/s79 by cache pool39.5 tok/s3,122 tok/s
1M* E118 tok/s39 by cache pool64.8 tok/s2,529 tok/s
Prefill8,000 tok/s E — a cold, uncached 128k prompt in 16 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. Qwen3.8-Flash-Next caches 13.03 KB per token in the Q8 configuration we ship, so this machine's 530 GB cache pool would hold 1,271 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 187–193 conversations. At 262k (155) and 512k (79) and 1M (39) the cache pool is the tighter limit, and those rows are the pool's number.

Measured on this box — eight RTX PRO 6000 Blackwell Server Edition cards, 768 GB pooled — on vLLM, run as two independent copies of the model, each spread across four GPUs, with a load balancer in front. The published figure is the box's AGGREGATE throughput at 64 concurrent requests, with a 251k-token prompt in the mix and MTP speculation on. No single-stream run was published.

4UXGM-TURIN2 DIRECT, Qwen3.8-Flash-Next — the published measurement

Requests in flightAggregate tok/s
Two TP4 engines, 64 concurrent3,340.5

What the source actually varies is the draft head, not the width: predicting three tokens ahead measured +72.1% at 1 request, +38.0% at 8 and +17.3% at 16, and −4.6% at 32 — so speculation is a low-load win and a high-load cost on this box, which is why the rows above are fitted to the measured aggregate rather than to a speculated single-stream rate. Reusing sparse-attention work between prediction steps bought +3.7% at 8 requests. The one-user row on this card is not a measurement: it is 3,340.5 tok/s read back through the same saturation curve this site fits to every pair, marked E for that reason. The published run is capped at 64 concurrent requests by the test, not by the machine.

Sources: helix.ml — Qwen3.8-Flash-Next on eight GPUs: what actually helped

DeepSeek V4.1-Flash — The full 476 GiB source checkpoint resident — about 64 GB of weights per card — leaving ~247 GB of the 768 GB for cache. No host memory and no NVMe in the serving path.

ContextOne userAt onceEach, all runningMachine total
32k E205 tok/s106 by throughput20.2 tok/s2,137 tok/s
128k E200 tok/s104 by throughput20.0 tok/s2,084 tok/s
262k E195 tok/s101 by throughput20.1 tok/s2,029 tok/s
512k E188 tok/s97 by throughput20.1 tok/s1,953 tok/s
1M E (recommended)180 tok/s93 by throughput20.1 tok/s1,866 tok/s
Prefill8,000 tok/s E — a cold, uncached 128k prompt in 16 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 247 GB cache pool would hold 4,336 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 93–106 conversations.

Measured on this box: 3,340.5 tok/s at 64 concurrent, 251k-token prompt in the mix. Qwen measured with two model copies across the eight cards, 64 requests in flight, a 251k-token prompt among them, MTP speculation on — and no single-stream run published, which is why its one-user rows are marked E. DeepSeek has no decode sweep at this width: its rows are scaled from the measured four-card box, which ran DSpark block 5 with the cards capped at 275 W.

DeepSeek V4.1-Flash with the whole checkpoint resident, and Qwen3.8-Flash-Next run as two independent four-card copies — measured at 3,340.5 tok/s across the box with 64 requests in flight.

Tier 5 — 8U16X-TURIN2 B300

8U16X-TURIN2 B300 — 8U rack · 8× B300 · IPMI

The top of the range: an 8× B300 node for volume that has outgrown a desk.

Runs: DeepSeek V4.1-Flash.

  • Memory: 2.10 TB (8 cards pooled) · 64 TB/s · 2.30 TB installed
  • Power draw: ≈ 12.0 kW
  • Price: €596.639–€963.866 ex-VAT

DeepSeek V4.1-Flash — 361 GB of weights and 1,039 GB of cache (~1.93 TB of the 2.30 TB installed, 2.1 TB usable), plus the 189 GB Engram tier pinned in host RAM.

ContextOne userAt onceEach, all runningMachine total
32k E110 tok/s2,629 by throughput20.0 tok/s52,583 tok/s
128k E108 tok/s2,558 by throughput20.0 tok/s51,168 tok/s
262k E105 tok/s2,227 by cache pool21.3 tok/s47,422 tok/s
512k E102 tok/s1,140 by cache pool30.1 tok/s34,268 tok/s
1M E (recommended)98 tok/s570 by cache pool37.9 tok/s21,594 tok/s
Prefill40,000 tok/s E — a cold, uncached 128k prompt in 3.3 s, server loaded and otherwise idle

At once is the smaller of two limits, and the row says which one showed. DeepSeek V4.1-Flash caches 1.78 KB per token in the Q8 configuration we ship, so this machine's 1,039 GB cache pool would hold 18,240 conversations at 32k on memory alone. We quote instead the largest number at which every conversation still gets 20 tok/s — a floor rather than a measurement, read off this machine's own fitted saturation curve. It binds at 2,558–2,629 conversations. At 262k (2,227) and 512k (1,140) and 1M (570) the cache pool is the tighter limit, and those rows are the pool's number. On this machine that floor is not a measurement and we do not present it as one: no published measurement of DeepSeek V4.1-Flash exists on an 8× B300 or HGX B300 node, and the nearest measured evidence on this exact node class — TheAI Cloud's 8× B300 SXM6 node on out-of-box vLLM — puts a DeepSeek Flash-line model at a 4,722 tok/s aggregate plateau across the whole node, with that team's ramp hard-capped at C=512. Read 2,629 conversations against that plateau and each conversation gets about 1.8 output tokens a second — below the floor this table quotes, and below any interactive threshold. The evidence behind both numbers, and this model's own measured frontier on B300, are set out under this table.

NO PUBLISHED MEASUREMENT OF THIS MODEL ON AN 8× B300 OR HGX B300 NODE EXISTS. Not from NVIDIA, not from DeepSeek, and not in the vLLM or SGLang recipe tables — those mark DeepSeek V4.1-Flash verified on H200, GB200, GB300 and MI355X, and do not mark B300 at all. So the table above is a model, not a measurement: a curve fitted to this node class and read as a floor, at the concurrency where each conversation still gets 20 tok/s. What IS measured on this node class is below, and it is a different, smaller DeepSeek model.

One 8× NVIDIA B300 SXM6 node, ~2.1 TiB, out-of-box vLLM 0.26.0 — A DIFFERENT MODEL, NOT DeepSeek V4.1-Flash — aggregate output tok/s, concurrency ramp doubling to a hard cap of C=512

Model on the same node classActive paramsWeightsOne user (tok/s)Node plateau (tok/s)Sustained (tok/s)
DeepSeek V4-Flash304B MoE · ~14B155 GB FP41724,7224,660
DeepSeek V4-Pro1.6T MoE · 49B805 GB FP41022,0511,338

All output tokens a second, whole node, on the same node class and the same pooled memory as this card, with the concurrency ramp doubling to a hard cap of C=512 — the plateau lives at C=256–512. Across the seven frontier MoEs in that run the plateau band is 1,137–5,769 tok/s, and the team flags that its own single-process load client saturated near 4.7k tok/s, so 4,722 is what that harness could draw rather than a proven ceiling. What it does establish is the order of magnitude: the machine totals this card's curve implies sit well above every plateau measured on this node class, which is why we publish them as modelled. No context-resolved concurrency table for DeepSeek V4.1-Flash exists on any B300 either, so the 32k and 128k rows above can be neither checked nor replaced by a measurement — they are a modelled floor, and they are tagged E for exactly that reason.

InferenceX B300 — per chip, not a node, and not a conversation count

SemiAnalysis's InferenceX publishes the closest measurement of this exact model on B300 silicon: 2- and 4-chip slices of a node (TP2 and TP4), concurrency 1 to 128, on vLLM with prefix-cache reuse, run against AgentX agentic coding traces — input p50 92k tokens, p90 304k, p99 734k; output p50 435. Its unit is tokens per chip per second, and it counts TOTAL tokens, input plus output, much of it served from prefix cache — not output only. So it is a different measurement from ours: it is not a node figure, it cannot be converted into a number of conversations, and it is not a context-resolved concurrency table. Its highest published point is 155,834 tok/s per chip at 64 concurrent; its widest is 128 concurrent, at 125,107 tok/s per chip and 38.8 tok/s per user on a two-chip slice, and 103,823 per chip and 27.0 per user on a four-chip slice. Per-user rates across its published B300 rows run 27.0–292 tok/s, and the 20 tok/s floor this site quotes sits under the slowest of them — which is what a policy floor should do.

Sources: cloud.theai.com — seven open models on one 8× B300 SXM6 node, out-of-box vLLM · inferencex.semianalysis.com — DeepSeek V4.1-Flash on B300 (TP2/TP4, AgentX traces, per chip)

No published measurement of DeepSeek V4.1-Flash exists on an 8× B300 or HGX B300 node — not from NVIDIA, not from DeepSeek, and not in the vLLM or SGLang recipe tables, which mark this model verified on H200, GB200, GB300 and MI355X and do not mark B300 at all. Every DeepSeek figure on this card is modelled, and the measured block above names the nearest evidence that does exist on this node class. Qwen3.8-Flash-Next: NO MEASUREMENT OF THIS MODEL ON AN 8× B300 OR HGX B300 NODE EXISTS, so we publish no node figure for it and derive none — this card is not offered with it, and the model's own page says the same. The one-user row here is this node read back through its own curve to one conversation, not a single-stream measurement: no single-stream run of this model on this node class has been published, and the four-card box reads higher per user partly because its published run had DSpark on. On the ceiling: the largest concurrency ever documented on B300-class silicon is 8,192, on a 72-GPU GB300 NVL72 rack, published as an accuracy evaluation with a gsm8k score rather than a throughput sweep — so 2,048 is a limit of what has been run on a node, not the node's ceiling.

DeepSeek V4.1-Flash sized for many people at once. At this size power and cooling are a design job — see the note on the page. We do not offer Qwen3.8-Flash-Next on this node, and no measurement of that model on an 8× B300 or HGX B300 exists — so there is no node figure for it anywhere on this site, derived or otherwise.

How to read these numbers

Every system is quoted. That covers the machine, the build, the operating system and inference stack, delivery, and your first employee running before handover. Prices are ex-VAT: a VAT-registered business reclaims it, and a cross-border EU purchase is normally invoiced under reverse charge with no VAT at all.