What moved this week

Qwen3.8-Flash-Next opens its weights, and the licence has a clause aimed at businesses like ours

· Multimodal

125B on the label, 176 GB in practice once the separate embedding table is counted, and only 6B working per token — so it is far faster than its size suggests, and it reads text, images and video. Native context is 262k; the 1M figure needs manual configuration. Read the licence before planning on it: the Qwen Community Licence 1.0 requires a separate agreement from Qwen for anyone running a "Model as a Service or AI Work Assistant business", with no revenue or user threshold. Internal use is explicitly exempt. Our everyday models are Apache 2.0 and unaffected.

Qwen Community Licence 1.0, on the model repo

GLM-5.3 Flash lands the same day

· Multimodal

Z.AI shipped GLM-5.3 Flash on the same date as Qwen's release. Two open releases in one day is now ordinary. The question worth asking of each is not which scores higher but which one fits a machine you can afford to keep running.

LLM release tracker

LTX-2.5: open-weight video with sound, 22B

· Video

A 22-billion-parameter open model that generates picture and audio together rather than separately, six to twenty seconds, up to 4K. Worth knowing about and worth being honest about: video is still the most expensive thing to run locally, it needs the larger machine, and for most small businesses it is not a job at all.

Open Source For You

llama.cpp is tagging builds several times a day

· Harness

The runtime behind most single-machine setups was in the b10600s by 28 August, released as sequential build tags rather than versions. Good news for support of new models. Also a reason to pin a build and test before moving, rather than tracking the tip on a machine somebody depends on.

llama.cpp releases

Memory prices are the biggest change to what a machine costs

· Hardware

A 32 GB DDR5-6000 kit is around $392, against $110–140 in late 2025. Reported 64 GB prices differ by source — roughly $850 to $1,118 — and we quote the range rather than pick one. The cause is memory makers moving wafer capacity to HBM for AI datacentres, where a gigabyte costs about three times the wafer of ordinary DDR5. This changes the price of every machine we quote, and it may last into 2027.

XenoSpectrum, on the DDR5 squeeze

HiDream opens an 8B image model under MIT

· Image

The best picture quality per gigabyte you can self-host, and MIT means no revenue threshold and no restriction on selling what it makes. Eight billion parameters, so it fits the everyday machine rather than needing the large one.

HiDream-O1 model card

Mistral opens Voxtral TTS, and it beats ElevenLabs in blind tests

· Voice

Apache 2.0, 4.1 billion parameters, about 3 GB of memory. Nine languages, and it clones a voice from three to five seconds of audio. Human listeners preferred it to ElevenLabs Flash v2.5 in 68 per cent of blind comparisons. A European lab, a permissive licence, and it runs on your own machine.

Mistral release coverage

Qwen3-Coder-Next: a coding model built to run on your own machine

· Coding

Apache 2.0, 80 billion parameters but only 3 billion working per token, so it runs far lighter than its size suggests. A 262,000-token window, which is enough to hold a real codebase. Around 70 per cent on SWE-bench — the practical choice if you want a coding agent that never sends your source anywhere.

Qwen3-Coder-Next model card

HeartMuLa opens a music model you can actually license

· Music

Apache 2.0 on the code and the weights, which is rare here. It writes clearer sung lyrics than anything else measured, including the paid services, and places second overall in listening tests behind Suno. The trade is speed: it generates at about real time, so a four-minute track takes about four minutes. Around 4.8 GB has to stay in memory once the audio codec is counted.

HeartMuLa-oss-3B model card

ACE-Step 1.5 XL: a full song in seconds, on a card you already own

· Music

Fast enough to change the job — a four-minute track with vocals in under ten seconds on a mid-range card, under 4 GB for the smaller version. We are not stocking it. The code is MIT but the weight files carry no stated licence at all, which is not the same as permissive, and several guides report the project as Apache 2.0 when it is not. Until that is fixed you cannot safely sell what it makes.

ACE-Step 1.5 repository

Second-hand cards are no longer the cheap way in

· Hardware

Used RTX 4090 cards are trading around $2,100, up from $1,200–1,400 earlier in the year: production ended in 2024 and demand keeps absorbing what is left. If you were planning to save money with a used card, price a new one before you decide.

Petronella, AI workstation build costs