Benchmarks
Independent results for the open models a small business can actually run. Five categories; in each, three open-weight models measured against the one closed model to beat. Every number names its source, and where no measurement exists the page says so rather than filling the gap.
Virtual workers — our own model list
Open weights, vision and thinking — the latest from each lab that we would deploy, and the licence is part of why each one qualifies: it has to let a Danish business ship what it makes. Our picks are DeepSeek V4.1-Flash and Qwen3.8-Flash-Next, the models this site runs and mods.
DeepSeek V4.1-Flash — DeepSeek — our pick
763.2B on disk / 8B prefill, 16B decode · 1M tokens native; reads images · runs on: The 4UXGM or the 8U16X node
The benchmark
- AA Intelligence Index 39, retrieved 22 September 2026
- Terminal-Bench 2.1: 74.53% at $0.10 per test — Vals AI's independent harness
- 890 bytes per token of cache at FP4 (1.78 KB at the Q8 we ship), so a 1M-token conversation holds in about 1.8 GB
Licence: MIT — no threshold, no territory clause, no naming requirement.
One of the two models this site runs and mods.
Qwen3.8-Flash-Next — Alibaba (Qwen) — our pick
180B packaged / 6B active per token · 262k native, 1M configured; reads text, images and video · runs on: The 4U4G, the DGX Station or the Spark (99 GB NVFP4 build)
The benchmark
- AA Intelligence Index 40, retrieved 22 September 2026
- AA's own Terminal-Bench 2.1 run: 86.14%
- Measured on one DGX Spark at 43.0 tok/s generating text, flat from 327 to 258,790 prompt tokens
Licence: Qwen Community Licence 1.0 — commercial use with naming above 100M users / $20M revenue; a separate licence is required to run it AS a service, which does not bind internal use.
One of the two models this site runs and mods.
MiMo-V2.6-Pro — Xiaomi
1.02T / 42B active per token · 1M tokens; reads text, images, video and audio · runs on: The 8U16X node (~673 GB at the shipped 4-bit build)
The benchmark
- AA Intelligence Index 46 — the highest open score on the index
- Terminal-Bench 2.1: vendor reports 89.9; no independent run exists
Licence: MIT — verified on the model card 22 September 2026.
Muse Glimmer-30B — Meta
29.6B dense · 128k, 256k extended; reads images · runs on: The Spark, the Station or the 4U4G
The benchmark
- AA Intelligence Index 17, retrieved 22 September 2026
- AA measured an 82% hallucination rate — read its scores with care
Licence: Apache 2.0 — no user cap, no territorial carve-out, no attribution clause.
Nemotron-3-Nano-Omni-30B-A3B — NVIDIA
31B / 3B active per token · 262k tokens; hears speech, sees images and video · runs on: The Spark, the Station or the 4U4G
The benchmark
- AA Intelligence Index 10 (AA's own estimate flag — a placement, not a measurement)
- No coding benchmark has been published, so we do not claim it codes
Licence: NVIDIA Open Model Licence — perpetual, worldwide, royalty-free, no territorial exclusion.
Ornith-1.5-35B-A3B — Ornith (DeepReinforce)
35B / 3B active per token · 256k, 1M extended; vision via a 903 MB projector · runs on: The 4U4G
The benchmark
- SWE-bench Pro 59.6; SWE-bench Verified 79.0 (catalogue figures)
- Not run on the AA index at all — no figure exists
Licence: MIT — and the vendor states it is globally accessible, free from regional limitations.
The cheapest cache on the roster (~20 KB/token), which is what makes its long context practical.
Mistral Large 3 — Mistral
675B / 41B active per token · 256k tokens; reads images · runs on: The 4UXGM or the 8U16X node
The benchmark
- AA Intelligence Index 9 — the lowest of the three Mistral general models
- Vision and thinking both present, which is why it qualifies here
Licence: Apache 2.0.
Office/CAD
The benchmark — the full field
| Model | score |
| Claude Opus 5.5 (Anthropic) (cloud API) | 67 |
| GPT-6 Astra (OpenAI) (cloud API) | 60 |
| Claude Opus 5 (Anthropic) (cloud API) | 59 |
| GPT-5.6 Sol (OpenAI) (cloud API) | 56 |
| GPT-6 Sol (OpenAI) (cloud API) | 55 |
| MiMo-V2.6-Pro | 52 |
| GLM-5.3 | 52 |
| DeepSeek V4.1-Flash — our pick | 49 |
| Kimi K3 | 48 |
| Qwen3.8-Flash-Next — our pick | 44 |
source: components: artificialanalysis.ai/leaderboards/models — weighting: dobkra · our retrieval date for the components — the weighting is ours
Our own weighting of the index components for office and CAD work: the science evals removed and the abstention component left out.
dobkra's office/CAD weighting of the Artificial Analysis Intelligence Index components · source: components: artificialanalysis.ai — weighting: dobkra · our retrieval date for the components
Agentic LLM
The benchmark — the full field
| Model | index |
| Claude Opus 5.5 (Anthropic) (cloud API) — office/CAD 67 | 58 |
| GPT-6 Astra (OpenAI) (cloud API) — office/CAD 60 | 53 |
| GPT-6 Sol (OpenAI) (cloud API) — office/CAD 55 | 48 |
| MiMo-V2.6-Pro — office/CAD 52 | 46 |
| GLM-5.3 — office/CAD 52 | 45 |
| Kimi K3 — office/CAD 48 | 44 |
| Claude Opus 4.7 (Anthropic) (cloud API) | 41 |
| Qwen3.8-Flash-Next — our pick — office/CAD 44 | 40 |
| DeepSeek V4.1-Flash — our pick — office/CAD 49 | 39 |
| GPT-5.5 (OpenAI) (cloud API) — office/CAD 47 | 38 |
source: artificialanalysis.ai/leaderboards/models · our retrieval date — Artificial Analysis publishes no snapshot date of its own; index v4.3.2 was announced 7 September 2026
Every figure here is on one index: Artificial Analysis Intelligence Index v4.3.2, retrieved from its leaderboard on 23 September 2026. AA publishes no snapshot date for its own leaderboard, so the date above is our retrieval date, not an AA "as of". Scores are rounded to whole index points; AA's own payload carries decimals. Scored at each model's highest reasoning effort. The chart plots the paid frontier (OpenAI and Anthropic, with the previous generation carried), the three best open-weight models you can run, and the two models this site runs and mods — marked as our picks. GPT-5.5, Claude Opus 4.6 and Claude Opus 4.7 are carried for reference: Artificial Analysis marks all three deprecated (replaced between February and April 2026), and their rows sit below the current open models. The office/CAD figure beside each score is this site's own weighting of the same index for the work this shop actually does: the two scientific-reasoning evals removed (HLE and CritPt — 20% of AA's weight, and not office or CAD work), the remaining components re-scaled to 100% at AA's own relative weights — GDPval-AA 22%, Terminal-Bench 4.0 22%, SciCode 22%, AA-Omniscience accuracy 22%, AA-LCR 11%. The non-hallucination component is deliberately left out: it measures whether a model declines to answer when it does not know, which is not what office or CAD work asks of it — a drafting or drawing task is checked against the document, not against the model's willingness to abstain. The index itself still carries that component; only this figure leaves it out. The figure is derived here, not published by AA, and it covers half the index's weight: three office-relevant components — AA-Briefcase, AutomationBench and GDP.pdf — are not published per model and are therefore absent. Where AA has not run a component, the row carries no figure rather than a guess.
- Not run on this index: Ornith-1.5-35B-A3B — Artificial Analysis has never benchmarked it, so no figure exists to plot. Solar-Open2-250B's licence is the Upstage Solar License — the Apache License 2.0 in full, including commercial use, with one added condition: a derivative model must carry the Solar brand. Commercially usable, so it plots as open.
Reasoning, tool use and coding — the model that does the work. Ranked on the independent intelligence index, with the terminal-agent benchmark and what a task costs to run under it.
Artificial Analysis Intelligence Index v4.3.2 (retrieved 23 September 2026), with Terminal-Bench 2.1 — Vals AI, Terminus 2 harness (board re-read 23 September 2026) · source: artificialanalysis.ai/leaderboards/models · vals.ai/benchmarks/terminal-bench-2-1 · our retrieval date — neither source publishes a snapshot date of its own
The three open models, against the one to beat
Claude Opus 5.5 — Anthropic — closed, the one to beat
Not published · 1M tokens · released 17 September 2026
The benchmark
- AA Intelligence Index 58 — the highest score on the index, and the leader among 148 reasoning models in AA's own words, retrieved 22 September 2026
- Terminal-Bench 2.1: 87.64% at $0.45 per test — Vals AI's independent Terminus 2 harness, board updated 21 September 2026
- AA's published cost per Intelligence-Index task: $5.98
Licence and price
Licence: Cloud only — no weights. The yardstick, not a candidate.
Price: USD 4 in / 20 out per 1M tokens — Artificial Analysis's price basis, retrieved 22 September 2026
Cost to do a job: $0.45 per Terminal-Bench 2.1 task (Vals AI) and $5.98 per Intelligence-Index task (AA) — one model, two harnesses, two prices for a job.
Took the top of the index on 17 September 2026, above the model the site carried as leader before it (Claude Fable 5.1, 53). The DeepSeek model page has been corrected to match.
MiMo-V2.6-Pro — Xiaomi — open weights
1.02T total / 42B active per token (vendor's figures) · 1M tokens; reads text, images, video and audio · released 21 September 2026
The benchmark
- AA Intelligence Index 46 — the highest score of any open-weight model on the index, retrieved 22 September 2026
- AA's published cost per Intelligence-Index task: $0.13 — the cheapest on the index
- Terminal-Bench 2.1: no independent run exists. Neither AA nor Vals AI has measured this model on any TB2.1 harness.
The vendor's own benchmarks
- Terminal Bench 2.1 · 89.9 — against Claude Opus 5 89.1, GPT-5.6 Sol 88.8, Claude Fable 5 84.3
- Terminal Bench 4.0 · 34.9 — against Opus 5 49.0, Fable 5 42.4, Sol 39.9
- DeepSWE v1.1 · 71.9 — against Opus 5 74.0, Sol 73.0, Fable 5 70.0
- AutomationBench v1.0.6 · 53.1 — ahead of Opus 5 50.3
- Toolathlon-Verified · 76.9 — against Opus 5 80.6
- GDPval-AA 2.1 (Elo) · 1673 — against Opus 5 1708, Sol 1588
- OSWorld-Verified · 82.0 — against Opus 5 83.4
- Agents' Last Exam · 31.6 — level with Opus 5 31.6
Source: MiMo-V2.6-Pro-RL model card — the vendor's own runs, on the vendor's own harness — retrieved 22 September 2026.
Licence and price
Licence: MIT — commercial use, no revenue threshold, no territory clause, verified on the model card 22 September 2026.
public, ungated; 985 downloads, last modified 22 September 2026 (HF API, read 22 September 2026)
Price: USD 0.435 in / 0.87 out per 1M tokens — OpenRouter, retrieved 22 September 2026; AA's record carries the same figures
Cost to do a job: $0.13 per Intelligence-Index task (AA, measured) — about a forty-fifth of the closed leader's $5.98 for the same job.
Runs on: The 8U16X node: 1.02T at the shipped 4-bit build is about 673 GB.
Where the two disagree
The vendor's table puts it level with Claude Opus 5 on terminal work (89.9 against 89.1); AA's independent index puts it twelve points below the current closed leader (46 against 58). No independent TB2.1 run exists to settle it, so both figures are shown and neither is averaged.
Released 21 September 2026 under MIT; the news feed carries the release, and this is the first chart pull that includes it.
GLM-5.3 — Z.ai — open weights
753B total / 40B active per token · 1M tokens; text only · released 4 September 2026
The benchmark
- AA Intelligence Index 45, retrieved 22 September 2026
- AA's published cost per Intelligence-Index task: $2.01
- Terminal-Bench 2.1: 71.54% at $0.31 per test — Vals AI's independent Terminus 2 harness, board updated 21 September 2026
The vendor's own benchmarks
- Terminal Bench 2.1 · 88.2 — against Kimi K3 88.3, Sol 88.8, DeepSeek-V4-Pro-0813 87.9
- Terminal Bench 3.0 · 28.3 — against Fable 5 33.7, Sol 34.6
- DeepSWE v1.1 · 66.9 — against Sol 72.7, Kimi K3 67.5
- CyberGym · 84.5 — the highest on its own table
- AutomationBench v1.0.6 · 48.2 — ahead of Sol 45.8
- HLE with tools · 62.5 — against Sol 64.5
- GDPval-AA v2 (Elo) · 1769 — ahead of Qwen3.8-Max 1739 and Sol 1730
- NL2Repo · 58.0
Source: GLM-5.3 model card — the vendor's own runs, on the vendor's own harness — retrieved 22 September 2026.
Licence and price
Licence: Free for commercial use. A security review is required only if you run it AS a service and your group turns over more than $10bn in any 12 months — no territory is excluded (GLM licence, on the model card).
public, ungated; 1,039,477 downloads, last modified 4 September 2026 (HF API, read 22 September 2026)
Price: USD 0.6538 in / 2.0548 out per 1M tokens — OpenRouter, retrieved 22 September 2026 (AA's record carries $1.4 / $4.4 for the same model).
Cost to do a job: $0.31 per Terminal-Bench 2.1 task (Vals AI) and $2.01 per Intelligence-Index task (AA).
Runs on: The 4UXGM or the 8U16X node: 753B at the shipped 4-bit build is about 497 GB.
Where the two disagree
Terminal-Bench 2.1: the vendor's own harness reads 88.2, Vals AI's independent Terminus 2 harness reads 71.54 on the same benchmark — 16.7 points apart, the harness named on both sides. The price basis also disagrees: OpenRouter lists $0.6538 in / $2.0548 out per 1M tokens where AA's record carries $1.4 / $4.4 — different routes, both stated.
Kimi K3 — Moonshot — open weights
2.8T total / 104B active per token; MoonViT-V2 vision encoder · 1M tokens; reads text and images · released 2 September 2026
The benchmark
- AA Intelligence Index 44, retrieved 22 September 2026
- AA's published cost per Intelligence-Index task: $2.00
- Terminal-Bench 2.1: 80.90% at $0.34 per test — Vals AI's independent Terminus 2 harness, board updated 21 September 2026
The vendor's own benchmarks
- Terminal-Bench 2.1 · 88.3 — against Sol 88.8, Fable 5 88.0
- DeepSWE · 67.5 — against Sol 73.0
- GPQA Diamond · 93.5 — against Sol 94.1
- HLE-Full (no tools / with tools) · 43.5 / 56.0 — against Fable 5 53.3 / 63.0
- OSWorld-Verified · 84.8 — ahead of Sol 83.0
- MMMU-Pro (CoT / direct) · 81.6 / 83.4
- AutomationBench · 30.8 — ahead of Sol 29.7
- GDPval-AA v2 (Elo) · 1686 — against Fable 5 1747, Sol 1736
Source: Kimi K3 model card — the vendor's own runs, on the vendor's own harness — retrieved 22 September 2026.
Licence and price
Licence: Custom Kimi licence — free for commercial use; above 100M monthly users or $20M monthly revenue the model must be named in your product.
public, ungated; 1,900,376 downloads, last modified 2 September 2026 (HF API, read 22 September 2026)
Price: USD 3 in / 15 out per 1M tokens — OpenRouter, retrieved 22 September 2026
Cost to do a job: $0.34 per Terminal-Bench 2.1 task (Vals AI) and $2.00 per Intelligence-Index task (AA).
Runs on: No machine we sell — named here as a benchmark leader, not a candidate for owned hardware.
Where the two disagree
Terminal-Bench 2.1: vendor 88.3 against Vals AI's independent 80.90 — 7.4 points apart. It is also the open leader on the index that no machine we sell can hold: 2.8T at the shipped 4-bit build is about 1.85 TB.
Named, not measured
- Terminal-Bench 2.1: MiMo-V2.6-Pro has no independent run — neither AA nor Vals AI has measured it on any TB2.1 harness, so the vendor's own 89.9 stands alone.
- No intelligence-index figure at all: Ornith-1.5-35B-A3B — Artificial Analysis has never benchmarked it.
- SWE-bench Pro: no figure published for Grok 4.6, Kimi K3, Gemma 4 31B, Nemotron-3-Nano-Omni, Mistral Small 4, Mistral Medium 3.5 or Mistral Large 3.
The category's ceiling moved on 17 September 2026 and the page moves with it — the closed reference is the current top performer, not the one the site last carried. The chart plots the comparison: the one to beat and the three best open-weight models you can run. The wider open field stays in the record behind this page, and the refresh log tracks it.
Image
The benchmark — the full field
| Model | Elo |
| GPT Image 2.5 Sunburst (OpenAI) (cloud API) | 1197 |
| Nano Banana Pro (Google) (cloud API) | 1100 |
| Qwen-Image-2.1 (estimated — no arena measurement exists. Placed between the two Qwen rows this arena does measure that bracket it in the vendor's own comparison — Qwen-Image-3.0-Pro at 1,089 and Qwen Image 2.0 Pro at 1,028 — using the vendor's Qwen-Image-Bench scores (2.1 at 60.28, between 3 Pro's 62.36 and 2.0 Pro's 57.84), so the placement is proportional: 1,061, band 1,041 to 1,081. Inputs: Qwen's release blog, 20 September 2026, and this arena's rows, retrieved 23 September 2026. One disagreement is stated rather than averaged: the vendor's own chart ranks 2.1 above Google's Nano Banana Pro, which this arena measures at 1,100 — the two scales are not comparable, so the placement stays anchored to the family's own measured rows.) | 1061 |
| Qwen-Image-2512 | 998 |
| Ming-Image-0.1-Design (InclusionAI) | 995 |
source: artificialanalysis.ai/image/leaderboard/text-to-image · our retrieval date — the arena publishes no snapshot date of its own
Read the licence column, not just the ranking. Several models marketed as open are API-only, or carry a revenue cap or a territorial exclusion that rules out a Danish business — and the two highest-placed open rows on this board, Ideogram 4.0 and FLUX.2 [dev], are non-commercial by default — BFL sells a commercial licence for FLUX.2 [dev] — so they are named rather than plotted as free-to-use. The chart plots the paid leaders, the open-weight models you can run, and our pick — which is modelled and says so on its row.
- No arena Elo, so not ranked: Mage-Flow (Microsoft, MIT).
- Our pick for image generation: Qwen-Image-2.1 — released 20 September 2026 and not on this arena yet, so its chart position is modelled, not measured: the method, its inputs and its band are on the row itself. Its licence covers serving the model to others, not generating with it.
- Non-commercial open models that would outrank these three are named, not plotted: Ideogram 4.0 (1,011) and FLUX.2 [dev] (1,000).
Text-to-image quality by blind pairwise human votes — with the licence read as carefully as the ranking.
Artificial Analysis Text-to-Image Arena — Elo with 95% CI and appearance counts · source: artificialanalysis.ai/image/leaderboard/text-to-image · our retrieval date — the arena publishes no snapshot date of its own
The three open models, against the one to beat
GPT Image 2.5 Sunburst (max) — OpenAI — closed, the one to beat
Not published · Image generation; API only · released 8 September 2026
The benchmark
- Arena Elo 1,197 with 13,446 appearances — the arena's top row, retrieved 23 September 2026
- The arena's listed API price: $210.72 per 1,000 images
Licence and price
Licence: Cloud only — the arena's #1 at pull time.
Price: $210.72 per 1,000 images — the arena's own price column, retrieved 23 September 2026
Cost to do a job: About 21 cents per image at the arena's listed rate.
Its sibling GPT Image 2.5 Flare (max) sits second at 1,190. Google's Nano Banana Pro — the pro tier, 1,100 — is plotted beside it so the closed field reads at a glance.
Qwen-Image-2512 — Alibaba (Qwen) — open weights
Not stated on the card · Image generation; self-hosted · released 31 December 2025 (HF last modified)
The benchmark
- Arena Elo 998 with 2,186 appearances, retrieved 23 September 2026 — the arena's row reads "Qwen Image Max 2512" and links Qwen's 2512 demo space; the checkpoint repo is Qwen/Qwen-Image-2512
- The arena's listed API price: $20 per 1,000 images
Licence and price
Licence: Apache 2.0 — covers the outputs as well as the weights (HF model card, verified 22 September 2026).
public, ungated; 52,288 downloads (HF API, read 22 September 2026)
Price: $20 per 1,000 images (the arena's price column) — self-hosted, the marginal cost is your own electricity.
The site's catalogue carries it as the safest licence on the shelf: weights and outputs both covered.
Ming-Image-0.1-Design — InclusionAI — open weights
Not stated on the card · Image generation; self-hosted · released 22 September 2026 — the day of this pull
The benchmark
- Arena Elo 995 with 21,297 appearances — placed by the arena on its release day, retrieved 23 September 2026
- The arena's listed API price: $30 per 1,000 images
Licence and price
Licence: MIT — verified on the model card, 22 September 2026.
public, ungated; newly published, 0 downloads at pull time (HF API, read 22 September 2026)
Price: $30 per 1,000 images (the arena's price column) — self-hosted, the marginal cost is your own electricity.
Cosmos3-Super-Text2Image (agentic) — NVIDIA — open weights
64B · Image generation; self-hosted, needs a 96 GB card · released 31 May 2026
The benchmark
- Arena Elo 994 with 11,060 appearances — the "agentic" variant; the base variant reads 983 and the 4-step 970, retrieved 23 September 2026
- No API price listed on the arena — it is a self-hosted model
Licence and price
Licence: OpenMDW-1.1 (NVIDIA Open Model Development Agreement) — commercially usable, no territory exclusion (catalogue check, 1 September 2026).
public, ungated; 1,997 downloads, last modified 16 September 2026 (HF API, read 22 September 2026)
Price: Self-hosted — no published per-image price.
Runs on: A 96 GB card (the 4U4G class and up).
The two open models, side by side
| Qwen-Image-2512 | Ming-Image-0.1-Design |
| Arena Elo | 998 on 2,186 appearances | 995 on 21,297 appearances |
| Licence | Apache 2.0 — weights and outputs | MIT |
| Released | 30 December 2025 | 17 September 2026 |
| Weights, BF16 | 57.7 GB — 40.9 transformer + 16.6 text encoder + VAE | 52.9 GB — 34.0 MLLM + 12.3 transformer + 6.2 connector + VAE |
| Sampling | 50 steps, CFG 4.0 (the card's example) | 12 steps, CFG 1.0 |
| Resolution | aspect buckets; no number stated on the card | 2048² recommended, 1024² faster |
| Hardware | no requirement stated; 80 GB class for 58 GB of BF16 weights | one 80 GB CUDA GPU, validated |
| Serving | diffusers | vLLM-Omni |
| Built for | general and photoreal generation, editing; 2512 improved human realism and text rendering | text-rich design — UI, infographics, posters; RGBA transparency; editable-PPT skill |
| Design board | Qwen-Image (the original) 946; 2512 not listed | 1,082 — first, above Ideogram 4.0 Quality (1,052) |
| Adoption | 51,443 downloads, 964 likes | brand-new — 221 likes, downloads counter still 0 |
| Measured speed | no independent measurement exists | no independent measurement exists |
Sources, retrieved 23 September 2026: both model cards and repos, the arena, and the design board as shown on Ming's card — that board could not be reached directly, so it is vendor-cited rather than independent. Weights are read from the repositories' own file listings.
Named, not measured
- The two highest-placed open rows are not free to use: Ideogram 4.0 (1,011) is non-commercial, and FLUX.2 [dev] (1,000) is non-commercial by default — Black Forest Labs sells commercial weights licences that include it (their Platform, Professional and Enterprise tiers, via bfl.ai/licensing, retrieved 23 September 2026). Neither is plotted as free-to-use; only the second has a published path to commercial use.
- HiDream-O1-Image (979, MIT) is the next eligible open model below the three.
The arena's scale moved again since the previous pull (top row 1,370 then, 1,197 now, retrieved 23 September 2026): the whole board is re-pulled in one pass, never patched row by row. The chart carries the paid leaders, the open rows a business can run, and our pick — modelled, and marked as such on its row.
Video
The benchmark — the full field
| Model | Elo |
| Gemini Omni Flash (Google) (cloud API) | 1513 |
| Dreamina Seedance 2.5 720p (ByteDance) (cloud API) | 1474 |
| MiniMax H3 — EU excluded by default — obtainable | 1460 |
| Sora 2 Pro (OpenAI) (cloud API) | 1368 |
| Kandinsky 5.0 T2V Pro | 1175 |
| LTX-2 19B | 1154 |
| Wan 2.2 A14B — our pick | 1132 |
source: arena.ai/leaderboard/text-to-video · our retrieval date — the arena publishes no snapshot date of its own
Read the vote counts before the ranking. The newest models sit near the top on a few thousand votes, which is exactly when a leaderboard is least reliable and most quoted — the arena marks those rows Preliminary, and anything under 10,000 votes is treated as provisional here. The chart plots the paid leader and the open-weight models you can run, including the one this site stocks (Wan 2.2, marked); our catalogue's LTX-2.5 is not on this board, so the newest LTX row here is LTX-2 19B.
- MiniMax H3 (1,460) and HunyuanVideo 1.5 (1,169) are open weights with territory-limited licences. MiniMax's default terms exclude the EU, the UK, the Republic of Korea and the USA, and EU coverage is obtainable through MiniMax's licence process (licence text read 22 September 2026); HunyuanVideo's Tencent Community Licence expressly does not apply in the EU. H3 is shown in full above; HunyuanVideo is named, not plotted.
- Wan 3.0, 2.7 and 2.6 outscore the open field and have released no weights: API only, and the arena lists them as Proprietary. The same board shows HappyHorse-1.0 as Proprietary and carries no LTX-2.3 or LTX-2.5 row at all — the newest LTX row it measures is LTX-2 19B, which is what this chart plots.
Text-to-video quality by blind pairwise votes — and a field where the open weights sit a long way below the services.
LMArena / arena.ai Text-to-Video Arena — Elo with 95% CI and vote counts · source: arena.ai/leaderboard/text-to-video · our retrieval date — the arena publishes no snapshot date of its own; its rows load client-side, so this pull is a browser read
The three open models, against the one to beat
Gemini Omni Flash — Google — closed, the one to beat
Not published · Text-to-video; API only · released 2026
The benchmark
- Arena Elo 1,513 ±9 on 26,576 votes, retrieved 22 September 2026
- The arena's rank-1 row is Gemini Omni 1.1 Flash at 1,516 ±15 on 1,784 votes — under 10,000 votes, so provisional by the arena's own mark
Licence and price
Licence: Cloud only — the top row with a settled sample; the row above it is provisional.
Price: API pricing is per second of video on Google's own schedule; no arena price column exists for video.
The owner confirmed the media references as each arena's #1 at pull time; on this board the #1 row is provisional, so the yardstick is the top settled row and the provisional leader is named beside it.
MiniMax H3 — MiniMax — open weights
33B dense (H3-Omni-Transformer); about 13B of that sits in AdaLN branches whose outputs cache, so inference-only deployment loads the remainder · Text-to-video and image-to-video; 4–15 second clips at 768p and 24 fps, 2K by a second regeneration pass · released 2 August 2026 (licence date); repository last modified 13 August 2026
The benchmark
- Arena Elo 1,460 ±9 on 11,417 votes — the highest-placed open-weight row on this board, retrieved 23 September 2026
Licence and price
Licence: MiniMax H3 Community Licence — its default grant is limited to an Applicable Territory that excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America (licence text read 22 September 2026). EU coverage is obtainable: MiniMax grants it through a licence process, and the default terms are what they are subject to.
public, ungated; 3,664,216 downloads, last modified 13 August 2026 (HF API, read 23 September 2026)
Price: Self-hosted — no published per-clip price; MiniMax's hosted service is a different product at its own price.
Runs on: A 32 GB card and up: 33B dense at the shipped 4-bit build is about 22 GB before working buffers.
The top open-weight video model this arena has measured, and it moves Kandinsky 5.0 T2V Pro (1,175) down to the second open slot. Its default terms exclude the EU and EU coverage is obtainable through MiniMax's licence process — the entry states both halves rather than leaving either to a footnote.
Kandinsky 5.0 T2V Pro — Kandinsky (Sber) — open weights
Not stated on the arena row · Text-to-video; self-hosted · released 2026
The benchmark
- Arena Elo 1,175 ±21 on 2,019 votes — the highest-placed open-weight model on the board, retrieved 22 September 2026
Licence and price
Licence: MIT — as listed on the arena's licence column, and the family's Hugging Face cards carry the same licence.
Price: Self-hosted — no published per-clip price.
The highest-placed open row a Danish business can run as it stands — 338 points below the closed yardstick, and below MiniMax H3, whose EU coverage takes a licence process.
LTX-2 19B — Lightricks — open weights
19B · Text-to-video; self-hosted · released 4 August 2026 (HF last modified)
The benchmark
- Arena Elo 1,154 ±8 on 75,596 votes — the fattest sample of any open row on the board, retrieved 22 September 2026
Licence and price
Licence: LTX-2 community licence — free below $10M annual revenue; above that Lightricks asks you to license it. One binding condition: do not strip its provenance watermarking.
public; 304,795 downloads (HF API, read 22 September 2026)
Price: Self-hosted — no published per-clip price.
Runs on: The 4UXGM class and up (19B plus working buffers).
Wan 2.2 A14B — Alibaba — open weights
14B · Text-to-video; self-hosted · released 7 August 2025 (HF last modified)
The benchmark
- Arena Elo 1,132 ±15 on 10,458 votes, retrieved 22 September 2026
Licence and price
Licence: Apache 2.0 — no revenue threshold, no territorial exclusion, no watermark condition.
public, ungated; 3,146 downloads (HF API, read 22 September 2026)
Price: Self-hosted — no published per-clip price.
The last Wan released with weights; everything since is API-only.
Named, not measured
- HunyuanVideo 1.5 (1,169): its Tencent Community Licence expressly does not apply in the EU — named, not plotted. MiniMax H3's default terms exclude the EU as well, and EU coverage is obtainable through MiniMax's licence process; the entry states both halves.
- Wan 3.0 (1,476), 2.7 and 2.6 outscore the open field and have released no weights: API only.
Music
The benchmark — the full field
| Model | Elo |
| Mureka V9 (cloud API) | 1177.2 |
| Suno V5.5 (cloud API) | 1171.24 |
| HeartMuLa-oss-3B — our pick — PREDICTED, not measured: placed one measured generation-step below Suno V4.5, whose row it is compared against in its own vendor's listening test (69.93 MOS against 76.08) — the method, inputs and band are under this chart | 1010 |
| ACE-Step v1 3.5B — PREDICTED, not measured: placed one vendor-test rank below HeartMuLa — 66.66 MOS against its 69.93 in the same listening test — which is a second generation-step from the measured anchor, Suno V4.5 at 1,065.36; the method and band are under this chart | 955 |
| MusicGen (Meta) — non-commercial licence | 845.51 |
source: artificialanalysis.ai/music/leaderboard/instrumental · our retrieval date — the arena publishes no snapshot date of its own
The open side of this chart is MODELLED, not measured — no independent benchmark of an open music model exists, and every open row says so on its face. HeartMuLa's 1,010 is a prediction, and this is how it was derived. Its vendor's blind listening test (HeartMuLa paper, arXiv 2601.10547, Table 13) ranks it second of five — 69.93 MOS against Suno-v4.5's 76.08, both with 95% confidence intervals — and ahead of every open model in that test. Suno V4.5 is measured on this arena at 1,065.36, and the arena's own Suno generations sit 35 to 71 Elo apart (V4.5 to V5 to V5.5), so the placement is one measured generation-step below its reference: a band of 990 to 1,030, taken at its midpoint, 1,010. ACE-Step is placed the same way, one vendor-test rank lower (66.66 MOS): a second generation-step from the measured anchor, band 920 to 990, taken at 955 — a wider band, because the uncertainty compounds the further a placement sits from a measurement. The inputs are named: the vendor's MOS table (the paper, above) and this arena's measured Suno V4.5 and MusicGen rows (retrieved 22 September 2026). MusicGen is measured, but its weights are non-commercial, which is what the ✕ on its row says. Nothing here blends a prediction into a measured row.
- DiffRhythm2 and YuE sit in the same vendor table (58.33 and 57.93 MOS) and could be placed the same way, but neither has a confirmed commercial licence — named, not plotted. ACE-Step's model card declares Apache 2.0, and the owner accepted that as the statement of record on 23 September 2026, so it plots as open — with the note that the repository ships no separate licence file with the weights.
- Self-reported, not plotted: HeartMuLa 69.93 in its own listening test, against Suno-v4.5 at 76.08 — the vendor's figures, not this arena's.
- Stable Audio 2.0 carries open weights under a non-commercial licence, so it is named and never a candidate for a business.
Instrumental music quality by blind pairwise listening votes — the open side is modelled, because no independent measurement of an open music model exists.
Artificial Analysis Instrumental Music Arena — Elo with confidence bounds and appearance counts · source: artificialanalysis.ai/music/leaderboard/instrumental · our retrieval date — the arena publishes no snapshot date of its own
The three open models, against the one to beat
Mureka V9 — Mureka — closed, the one to beat
Not published · Instrumental music generation; API only · released 27 March 2026
The benchmark
- Arena Elo 1,177.2 on 2,237 appearances, retrieved 22 September 2026
Licence and price
Licence: Cloud only — the arena's #1 at pull time.
Price: Subscription/API pricing on the vendor's own schedule; no arena price column exists for music.
Suno V5.5 sits second at 1,171.24.
No open-weight model in this category has an independent measurement on this board, so the three slots are named below rather than filled with a placement.
Named, not measured
- No open-weight music model has an independent measurement on this arena — the open rows are modelled and say so on their face. The method, its inputs and its band are stated under the chart; nothing here is blended silently with a measured row.
- HeartMuLa-oss-3B (Apache 2.0) is the open model this site runs, and ACE-Step v1 3.5B plots beside it as open — its card's Apache 2.0 accepted as the statement of record (owner, 23 September 2026). Both are predicted rows: the method, inputs and bands are under the chart. DiffRhythm2 and YuE sit in the same vendor table but have no confirmed licence — named, not plotted.
- Nothing open released since 2023 has been measured on this board at all — that gap is the finding; the predicted row is where the open side stands.
The open side of this category is modelled by the owner's ruling of 2026-09-23: no independent measurement of an open music model exists, so a predicted row is placed instead — with its method, inputs and uncertainty stated under the chart, and marked as predicted on the row itself. A slot is still never filled with an unlabelled number.
Voice
The benchmark — the full field
| Model | Elo |
| Cartesia Sonic 3.6 (cloud API) | 1273 |
| ElevenLabs v3 Conversational (cloud API) | 1196 |
| OpenAI TTS-1 HD (cloud API) | 1099 |
| Magpie-Multilingual 357M (NVIDIA) | 1063 |
| Kokoro 82M v1.0 | 1061 |
| Chatterbox (Resemble AI) | 1021 |
source: artificialanalysis.ai/text-to-speech/leaderboard · our retrieval date — the arena publishes no snapshot date of its own
The widest open/closed gap on the site: the arena's open leaders are licence-blocked — Breeze TTS 2 is research and non-commercial, Fish Audio S2 Pro is research only, and Step Audio EditX ships weights with no licence of their own — so the best open voice a business can actually build on is Magpie-Multilingual at 1,063, with Kokoro at 1,061. The arena's own top row is Cartesia Sonic 3.6 at 1,273. The closed leaders this chart now carries sit far above them: ElevenLabs' v3 Conversational at 1,196 and OpenAI's TTS-1 HD at 1,099 — cloud-only, and plotted so the gap is visible rather than described. Whisper and other speech-to-text models are a different job and stay off this board.
- Not ranked on this arena, so not stocked: MisoTTS 8B, Orpheus TTS 3B, Higgs Audio v2. Our own voice-overs run on VoxCPM2 — local, cloned from the presenter voice, and not on this arena, so it has no Elo here.
- The rest of each family stays in the data beneath the chart, not dropped: ElevenLabs Eleven v3 1,167, Turbo v2.5 1,095, Multilingual v2 1,092, Flash v2.5 1,074; OpenAI TTS-1 1,086, GPT-Realtime-2 1,072; Google Gemini 3.8 Flash-Lite TTS 1,235 and Gemini 3.1 Flash TTS 1,199.
Text-to-speech quality by blind pairwise listening votes — the widest open/closed gap on the site, and mostly a licence story.
Artificial Analysis Speech Arena — Elo with 95% CI and sample counts · source: artificialanalysis.ai/text-to-speech/leaderboard · our retrieval date — the arena publishes no snapshot date of its own; its rows load client-side, so this pull is a browser read
The three open models, against the one to beat
Cartesia Sonic 3.6 — Cartesia — closed, the one to beat
Not published · Text-to-speech; API only · released August 2026
The benchmark
- Arena Elo 1,273 on 1,757 samples, retrieved 23 September 2026
- The arena's listed API price: $49.0 per 1M characters
Licence and price
Licence: Cloud only — the arena's #1 at pull time.
Price: $49.0 per 1M characters — the arena's own price column, retrieved 23 September 2026
Cost to do a job: $49.0 per 1M characters at the arena's listed rate.
The Inworld row the site carried last (Realtime TTS 1.5 Max, 1,209.6) has moved down the board: Realtime TTS-2 sits fourth at 1,245 and Realtime TTS-2 Flash eighth at 1,210. The two closed rows this chart gained on 23 September — ElevenLabs' v3 Conversational and OpenAI's TTS-1 HD — are the best each lab fields on this arena.
Magpie-Multilingual 357M — NVIDIA — open weights
357M · Text-to-speech; self-hosted, runs on almost anything · released February 2026
The benchmark
- Arena Elo 1,063 on 1,971 samples — the highest-placed open model a business can build on, retrieved 23 September 2026
Licence and price
Licence: NVIDIA Open Model Licence — perpetual, worldwide, royalty-free, no territorial exclusion (model card, verified 22 September 2026).
public, ungated; 6,956 downloads, last modified 9 September 2026 (HF API, read 22 September 2026)
Price: Self-hosted — no published per-character price.
Kokoro 82M v1.0 — hexgrad — open weights
82M — runs on a CPU · Text-to-speech; self-hosted · released January 2025
The benchmark
- Arena Elo 1,061 on 5,232 samples, retrieved 23 September 2026
- The arena's listed API price: $0.7 per 1M characters
Licence and price
Licence: Apache 2.0 (model card, verified 22 September 2026).
public, ungated; 11,722,492 downloads (HF API, read 22 September 2026)
Price: $0.7 per 1M characters at the arena's listed rate — or nothing but electricity, self-hosted.
Chatterbox — Resemble AI — open weights
0.5B · Text-to-speech with voice cloning; self-hosted · released May 2025
The benchmark
- Arena Elo 1,021 on 4,600 samples, retrieved 23 September 2026
- The arena's listed API price: $25.0 per 1M characters
Licence and price
Licence: MIT — what it says is yours to use (model card, verified 22 September 2026).
public, ungated; 1,794,542 downloads, last modified 10 June 2026 (HF API, read 22 September 2026)
Price: $25.0 per 1M characters at the arena's listed rate — or nothing but electricity, self-hosted.
Named, not measured
- The open leaders are licence-blocked and named, never candidates: Breeze TTS 2 (1,204 — research and non-commercial licence), Fish Audio S2 Pro (1,120 — research only), Step Audio EditX (1,094 — the code is Apache 2.0 but the weights carry no licence of their own).
- Voxtral TTS (1,076) is held off the candidate list: the licence on its model card and the licence in this site's own catalogue disagree, and that is being settled before it can be recommended.
- Whisper large-v3 and other speech-to-text models are a different job and stay off this board.
The category's closed reference moved with the arena: the Inworld row named when this structure was approved (Realtime TTS 1.5 Max) is no longer the top row, so the yardstick is the current #1.
Sources this page stands on
Refresh protocol: a source whose version moves is re-pulled whole — never patched row by row — and every chart prints its own retrieval date, so a stale snapshot is visible rather than assumed current. The Office/CAD figure is computed the same way for any model: the five published components from Artificial Analysis's model payload at the weights named on its chart, one row per model at its highest reasoning effort — so a new model gets a figure the day AA publishes its components.
The five machines and what each one runs · the two models we run and mod
A number here you want checked?
Every score on this page names its source, and the thin ones say so. If one of them is carrying a decision you are making, write to us and we will show the measurement behind it.
Email us · Register for a role