Two models, and why exactly these two
Every machine we build runs one of two open models, and we mod both: quantisation chosen from measured files, a Q8 cache by default, prompts tuned per role. Each page carries the licence, what the model costs in memory, what we change, and every update we have tracked.
The smartest model that fits in a desk-side box. Long-memory roles.
The smartest model that fits in a desk-side box, and what the long-memory roles run. 176B of weights with only 6B active per token, so it moves like a small model while knowing what a large one knows — 41 tok/s on a DGX Spark. It sees text, images and video, and its cache is the cheapest in the catalogue at 24 KB per token, which is what makes long context practical on one box.
Cheap, strong text and code at a 1M window.
The value pick: near its bigger sibling's quality for a fraction of the memory, under the most permissive licence in the catalogue, with a 1M-token window and a cache cheap enough to actually use it. Text and code only — where a job needs to read a document, it runs beside a seeing model rather than instead of one.