Benchmarks

Independent results for the open models a small business can actually run. Five categories; in each, three open-weight models measured against the one closed model to beat. Every number names its source, and where no measurement exists the page says so rather than filling the gap.

Virtual workers — our own model list

Open weights, vision and thinking — the latest from each lab that we would deploy, and the licence is part of why each one qualifies: it has to let a Danish business ship what it makes. Our picks are DeepSeek V4.1-Flash and Qwen3.8-Flash-Next, the models this site runs and mods.

DeepSeek V4.1-Flash — DeepSeek — our pick

763.2B on disk / 8B prefill, 16B decode · 1M tokens native; reads images · runs on: The 4UXGM or the 8U16X node

The benchmark

  • AA Intelligence Index 39, retrieved 22 September 2026
  • Terminal-Bench 2.1: 74.53% at $0.10 per test — Vals AI's independent harness
  • 890 bytes per token of cache at FP4 (1.78 KB at the Q8 we ship), so a 1M-token conversation holds in about 1.8 GB

Licence: MIT — no threshold, no territory clause, no naming requirement.

One of the two models this site runs and mods.

Qwen3.8-Flash-Next — Alibaba (Qwen) — our pick

180B packaged / 6B active per token · 262k native, 1M configured; reads text, images and video · runs on: The 4U4G, the DGX Station or the Spark (99 GB NVFP4 build)

The benchmark

  • AA Intelligence Index 40, retrieved 22 September 2026
  • AA's own Terminal-Bench 2.1 run: 86.14%
  • Measured on one DGX Spark at 43.0 tok/s generating text, flat from 327 to 258,790 prompt tokens

Licence: Qwen Community Licence 1.0 — commercial use with naming above 100M users / $20M revenue; a separate licence is required to run it AS a service, which does not bind internal use.

One of the two models this site runs and mods.

MiMo-V2.6-Pro — Xiaomi

1.02T / 42B active per token · 1M tokens; reads text, images, video and audio · runs on: The 8U16X node (~673 GB at the shipped 4-bit build)

The benchmark

  • AA Intelligence Index 46 — the highest open score on the index
  • Terminal-Bench 2.1: vendor reports 89.9; no independent run exists

Licence: MIT — verified on the model card 22 September 2026.

Muse Glimmer-30B — Meta

29.6B dense · 128k, 256k extended; reads images · runs on: The Spark, the Station or the 4U4G

The benchmark

  • AA Intelligence Index 17, retrieved 22 September 2026
  • AA measured an 82% hallucination rate — read its scores with care

Licence: Apache 2.0 — no user cap, no territorial carve-out, no attribution clause.

Nemotron-3-Nano-Omni-30B-A3B — NVIDIA

31B / 3B active per token · 262k tokens; hears speech, sees images and video · runs on: The Spark, the Station or the 4U4G

The benchmark

  • AA Intelligence Index 10 (AA's own estimate flag — a placement, not a measurement)
  • No coding benchmark has been published, so we do not claim it codes

Licence: NVIDIA Open Model Licence — perpetual, worldwide, royalty-free, no territorial exclusion.

Ornith-1.5-35B-A3B — Ornith (DeepReinforce)

35B / 3B active per token · 256k, 1M extended; vision via a 903 MB projector · runs on: The 4U4G

The benchmark

  • SWE-bench Pro 59.6; SWE-bench Verified 79.0 (catalogue figures)
  • Not run on the AA index at all — no figure exists

Licence: MIT — and the vendor states it is globally accessible, free from regional limitations.

The cheapest cache on the roster (~20 KB/token), which is what makes its long context practical.

Mistral Large 3 — Mistral

675B / 41B active per token · 256k tokens; reads images · runs on: The 4UXGM or the 8U16X node

The benchmark

  • AA Intelligence Index 9 — the lowest of the three Mistral general models
  • Vision and thinking both present, which is why it qualifies here

Licence: Apache 2.0.

Office/CAD

The benchmark — the full field

Modelscore
Claude Opus 5.5 (Anthropic) (cloud API)67
GPT-6 Astra (OpenAI) (cloud API)60
Claude Opus 5 (Anthropic) (cloud API)59
GPT-5.6 Sol (OpenAI) (cloud API)56
GPT-6 Sol (OpenAI) (cloud API)55
MiMo-V2.6-Pro52
GLM-5.352
DeepSeek V4.1-Flash — our pick49
Kimi K348
Qwen3.8-Flash-Next — our pick44

source: components: artificialanalysis.ai/leaderboards/models — weighting: dobkra · our retrieval date for the components — the weighting is ours

Our own weighting of the index components for office and CAD work: the science evals removed and the abstention component left out.

dobkra's office/CAD weighting of the Artificial Analysis Intelligence Index components · source: components: artificialanalysis.ai — weighting: dobkra · our retrieval date for the components

Agentic LLM

The benchmark — the full field

Modelindex
Claude Opus 5.5 (Anthropic) (cloud API) — office/CAD 6758
GPT-6 Astra (OpenAI) (cloud API) — office/CAD 6053
GPT-6 Sol (OpenAI) (cloud API) — office/CAD 5548
MiMo-V2.6-Pro — office/CAD 5246
GLM-5.3 — office/CAD 5245
Kimi K3 — office/CAD 4844
Claude Opus 4.7 (Anthropic) (cloud API)41
Qwen3.8-Flash-Next — our pick — office/CAD 4440
DeepSeek V4.1-Flash — our pick — office/CAD 4939
GPT-5.5 (OpenAI) (cloud API) — office/CAD 4738

source: artificialanalysis.ai/leaderboards/models · our retrieval date — Artificial Analysis publishes no snapshot date of its own; index v4.3.2 was announced 7 September 2026

Every figure here is on one index: Artificial Analysis Intelligence Index v4.3.2, retrieved from its leaderboard on 23 September 2026. AA publishes no snapshot date for its own leaderboard, so the date above is our retrieval date, not an AA "as of". Scores are rounded to whole index points; AA's own payload carries decimals. Scored at each model's highest reasoning effort. The chart plots the paid frontier (OpenAI and Anthropic, with the previous generation carried), the three best open-weight models you can run, and the two models this site runs and mods — marked as our picks. GPT-5.5, Claude Opus 4.6 and Claude Opus 4.7 are carried for reference: Artificial Analysis marks all three deprecated (replaced between February and April 2026), and their rows sit below the current open models. The office/CAD figure beside each score is this site's own weighting of the same index for the work this shop actually does: the two scientific-reasoning evals removed (HLE and CritPt — 20% of AA's weight, and not office or CAD work), the remaining components re-scaled to 100% at AA's own relative weights — GDPval-AA 22%, Terminal-Bench 4.0 22%, SciCode 22%, AA-Omniscience accuracy 22%, AA-LCR 11%. The non-hallucination component is deliberately left out: it measures whether a model declines to answer when it does not know, which is not what office or CAD work asks of it — a drafting or drawing task is checked against the document, not against the model's willingness to abstain. The index itself still carries that component; only this figure leaves it out. The figure is derived here, not published by AA, and it covers half the index's weight: three office-relevant components — AA-Briefcase, AutomationBench and GDP.pdf — are not published per model and are therefore absent. Where AA has not run a component, the row carries no figure rather than a guess.

  • Not run on this index: Ornith-1.5-35B-A3B — Artificial Analysis has never benchmarked it, so no figure exists to plot. Solar-Open2-250B's licence is the Upstage Solar License — the Apache License 2.0 in full, including commercial use, with one added condition: a derivative model must carry the Solar brand. Commercially usable, so it plots as open.

Reasoning, tool use and coding — the model that does the work. Ranked on the independent intelligence index, with the terminal-agent benchmark and what a task costs to run under it.

Artificial Analysis Intelligence Index v4.3.2 (retrieved 23 September 2026), with Terminal-Bench 2.1 — Vals AI, Terminus 2 harness (board re-read 23 September 2026) · source: artificialanalysis.ai/leaderboards/models · vals.ai/benchmarks/terminal-bench-2-1 · our retrieval date — neither source publishes a snapshot date of its own

The three open models, against the one to beat

Claude Opus 5.5 — Anthropic — closed, the one to beat

Not published · 1M tokens · released 17 September 2026

The benchmark

  • AA Intelligence Index 58 — the highest score on the index, and the leader among 148 reasoning models in AA's own words, retrieved 22 September 2026
  • Terminal-Bench 2.1: 87.64% at $0.45 per test — Vals AI's independent Terminus 2 harness, board updated 21 September 2026
  • AA's published cost per Intelligence-Index task: $5.98

Licence and price

Licence: Cloud only — no weights. The yardstick, not a candidate.

Price: USD 4 in / 20 out per 1M tokens — Artificial Analysis's price basis, retrieved 22 September 2026

Cost to do a job: $0.45 per Terminal-Bench 2.1 task (Vals AI) and $5.98 per Intelligence-Index task (AA) — one model, two harnesses, two prices for a job.

Took the top of the index on 17 September 2026, above the model the site carried as leader before it (Claude Fable 5.1, 53). The DeepSeek model page has been corrected to match.

MiMo-V2.6-Pro — Xiaomi — open weights

1.02T total / 42B active per token (vendor's figures) · 1M tokens; reads text, images, video and audio · released 21 September 2026

The benchmark

  • AA Intelligence Index 46 — the highest score of any open-weight model on the index, retrieved 22 September 2026
  • AA's published cost per Intelligence-Index task: $0.13 — the cheapest on the index
  • Terminal-Bench 2.1: no independent run exists. Neither AA nor Vals AI has measured this model on any TB2.1 harness.

The vendor's own benchmarks

  • Terminal Bench 2.1 · 89.9 — against Claude Opus 5 89.1, GPT-5.6 Sol 88.8, Claude Fable 5 84.3
  • Terminal Bench 4.0 · 34.9 — against Opus 5 49.0, Fable 5 42.4, Sol 39.9
  • DeepSWE v1.1 · 71.9 — against Opus 5 74.0, Sol 73.0, Fable 5 70.0
  • AutomationBench v1.0.6 · 53.1 — ahead of Opus 5 50.3
  • Toolathlon-Verified · 76.9 — against Opus 5 80.6
  • GDPval-AA 2.1 (Elo) · 1673 — against Opus 5 1708, Sol 1588
  • OSWorld-Verified · 82.0 — against Opus 5 83.4
  • Agents' Last Exam · 31.6 — level with Opus 5 31.6

Source: MiMo-V2.6-Pro-RL model card — the vendor's own runs, on the vendor's own harness — retrieved 22 September 2026.

Licence and price

Licence: MIT — commercial use, no revenue threshold, no territory clause, verified on the model card 22 September 2026.

public, ungated; 985 downloads, last modified 22 September 2026 (HF API, read 22 September 2026)

Price: USD 0.435 in / 0.87 out per 1M tokens — OpenRouter, retrieved 22 September 2026; AA's record carries the same figures

Cost to do a job: $0.13 per Intelligence-Index task (AA, measured) — about a forty-fifth of the closed leader's $5.98 for the same job.

Runs on: The 8U16X node: 1.02T at the shipped 4-bit build is about 673 GB.

Where the two disagree

The vendor's table puts it level with Claude Opus 5 on terminal work (89.9 against 89.1); AA's independent index puts it twelve points below the current closed leader (46 against 58). No independent TB2.1 run exists to settle it, so both figures are shown and neither is averaged.

Released 21 September 2026 under MIT; the news feed carries the release, and this is the first chart pull that includes it.

GLM-5.3 — Z.ai — open weights

753B total / 40B active per token · 1M tokens; text only · released 4 September 2026

The benchmark

  • AA Intelligence Index 45, retrieved 22 September 2026
  • AA's published cost per Intelligence-Index task: $2.01
  • Terminal-Bench 2.1: 71.54% at $0.31 per test — Vals AI's independent Terminus 2 harness, board updated 21 September 2026

The vendor's own benchmarks

  • Terminal Bench 2.1 · 88.2 — against Kimi K3 88.3, Sol 88.8, DeepSeek-V4-Pro-0813 87.9
  • Terminal Bench 3.0 · 28.3 — against Fable 5 33.7, Sol 34.6
  • DeepSWE v1.1 · 66.9 — against Sol 72.7, Kimi K3 67.5
  • CyberGym · 84.5 — the highest on its own table
  • AutomationBench v1.0.6 · 48.2 — ahead of Sol 45.8
  • HLE with tools · 62.5 — against Sol 64.5
  • GDPval-AA v2 (Elo) · 1769 — ahead of Qwen3.8-Max 1739 and Sol 1730
  • NL2Repo · 58.0

Source: GLM-5.3 model card — the vendor's own runs, on the vendor's own harness — retrieved 22 September 2026.

Licence and price

Licence: Free for commercial use. A security review is required only if you run it AS a service and your group turns over more than $10bn in any 12 months — no territory is excluded (GLM licence, on the model card).

public, ungated; 1,039,477 downloads, last modified 4 September 2026 (HF API, read 22 September 2026)

Price: USD 0.6538 in / 2.0548 out per 1M tokens — OpenRouter, retrieved 22 September 2026 (AA's record carries $1.4 / $4.4 for the same model).

Cost to do a job: $0.31 per Terminal-Bench 2.1 task (Vals AI) and $2.01 per Intelligence-Index task (AA).

Runs on: The 4UXGM or the 8U16X node: 753B at the shipped 4-bit build is about 497 GB.

Where the two disagree

Terminal-Bench 2.1: the vendor's own harness reads 88.2, Vals AI's independent Terminus 2 harness reads 71.54 on the same benchmark — 16.7 points apart, the harness named on both sides. The price basis also disagrees: OpenRouter lists $0.6538 in / $2.0548 out per 1M tokens where AA's record carries $1.4 / $4.4 — different routes, both stated.

Kimi K3 — Moonshot — open weights

2.8T total / 104B active per token; MoonViT-V2 vision encoder · 1M tokens; reads text and images · released 2 September 2026

The benchmark

  • AA Intelligence Index 44, retrieved 22 September 2026
  • AA's published cost per Intelligence-Index task: $2.00
  • Terminal-Bench 2.1: 80.90% at $0.34 per test — Vals AI's independent Terminus 2 harness, board updated 21 September 2026

The vendor's own benchmarks

  • Terminal-Bench 2.1 · 88.3 — against Sol 88.8, Fable 5 88.0
  • DeepSWE · 67.5 — against Sol 73.0
  • GPQA Diamond · 93.5 — against Sol 94.1
  • HLE-Full (no tools / with tools) · 43.5 / 56.0 — against Fable 5 53.3 / 63.0
  • OSWorld-Verified · 84.8 — ahead of Sol 83.0
  • MMMU-Pro (CoT / direct) · 81.6 / 83.4
  • AutomationBench · 30.8 — ahead of Sol 29.7
  • GDPval-AA v2 (Elo) · 1686 — against Fable 5 1747, Sol 1736

Source: Kimi K3 model card — the vendor's own runs, on the vendor's own harness — retrieved 22 September 2026.

Licence and price

Licence: Custom Kimi licence — free for commercial use; above 100M monthly users or $20M monthly revenue the model must be named in your product.

public, ungated; 1,900,376 downloads, last modified 2 September 2026 (HF API, read 22 September 2026)

Price: USD 3 in / 15 out per 1M tokens — OpenRouter, retrieved 22 September 2026

Cost to do a job: $0.34 per Terminal-Bench 2.1 task (Vals AI) and $2.00 per Intelligence-Index task (AA).

Runs on: No machine we sell — named here as a benchmark leader, not a candidate for owned hardware.

Where the two disagree

Terminal-Bench 2.1: vendor 88.3 against Vals AI's independent 80.90 — 7.4 points apart. It is also the open leader on the index that no machine we sell can hold: 2.8T at the shipped 4-bit build is about 1.85 TB.

Named, not measured

  • Terminal-Bench 2.1: MiMo-V2.6-Pro has no independent run — neither AA nor Vals AI has measured it on any TB2.1 harness, so the vendor's own 89.9 stands alone.
  • No intelligence-index figure at all: Ornith-1.5-35B-A3B — Artificial Analysis has never benchmarked it.
  • SWE-bench Pro: no figure published for Grok 4.6, Kimi K3, Gemma 4 31B, Nemotron-3-Nano-Omni, Mistral Small 4, Mistral Medium 3.5 or Mistral Large 3.

The category's ceiling moved on 17 September 2026 and the page moves with it — the closed reference is the current top performer, not the one the site last carried. The chart plots the comparison: the one to beat and the three best open-weight models you can run. The wider open field stays in the record behind this page, and the refresh log tracks it.

Image

The benchmark — the full field

ModelElo
GPT Image 2.5 Sunburst (OpenAI) (cloud API)1197
Nano Banana Pro (Google) (cloud API)1100
Qwen-Image-2.1 (estimated — no arena measurement exists. Placed between the two Qwen rows this arena does measure that bracket it in the vendor's own comparison — Qwen-Image-3.0-Pro at 1,089 and Qwen Image 2.0 Pro at 1,028 — using the vendor's Qwen-Image-Bench scores (2.1 at 60.28, between 3 Pro's 62.36 and 2.0 Pro's 57.84), so the placement is proportional: 1,061, band 1,041 to 1,081. Inputs: Qwen's release blog, 20 September 2026, and this arena's rows, retrieved 23 September 2026. One disagreement is stated rather than averaged: the vendor's own chart ranks 2.1 above Google's Nano Banana Pro, which this arena measures at 1,100 — the two scales are not comparable, so the placement stays anchored to the family's own measured rows.)1061
Qwen-Image-2512998
Ming-Image-0.1-Design (InclusionAI)995

source: artificialanalysis.ai/image/leaderboard/text-to-image · our retrieval date — the arena publishes no snapshot date of its own

Read the licence column, not just the ranking. Several models marketed as open are API-only, or carry a revenue cap or a territorial exclusion that rules out a Danish business — and the two highest-placed open rows on this board, Ideogram 4.0 and FLUX.2 [dev], are non-commercial by default — BFL sells a commercial licence for FLUX.2 [dev] — so they are named rather than plotted as free-to-use. The chart plots the paid leaders, the open-weight models you can run, and our pick — which is modelled and says so on its row.

  • No arena Elo, so not ranked: Mage-Flow (Microsoft, MIT).
  • Our pick for image generation: Qwen-Image-2.1 — released 20 September 2026 and not on this arena yet, so its chart position is modelled, not measured: the method, its inputs and its band are on the row itself. Its licence covers serving the model to others, not generating with it.
  • Non-commercial open models that would outrank these three are named, not plotted: Ideogram 4.0 (1,011) and FLUX.2 [dev] (1,000).

Text-to-image quality by blind pairwise human votes — with the licence read as carefully as the ranking.

Artificial Analysis Text-to-Image Arena — Elo with 95% CI and appearance counts · source: artificialanalysis.ai/image/leaderboard/text-to-image · our retrieval date — the arena publishes no snapshot date of its own

The three open models, against the one to beat

GPT Image 2.5 Sunburst (max) — OpenAI — closed, the one to beat

Not published · Image generation; API only · released 8 September 2026

The benchmark

  • Arena Elo 1,197 with 13,446 appearances — the arena's top row, retrieved 23 September 2026
  • The arena's listed API price: $210.72 per 1,000 images

Licence and price

Licence: Cloud only — the arena's #1 at pull time.

Price: $210.72 per 1,000 images — the arena's own price column, retrieved 23 September 2026

Cost to do a job: About 21 cents per image at the arena's listed rate.

Its sibling GPT Image 2.5 Flare (max) sits second at 1,190. Google's Nano Banana Pro — the pro tier, 1,100 — is plotted beside it so the closed field reads at a glance.

Qwen-Image-2512 — Alibaba (Qwen) — open weights

Not stated on the card · Image generation; self-hosted · released 31 December 2025 (HF last modified)

The benchmark

  • Arena Elo 998 with 2,186 appearances, retrieved 23 September 2026 — the arena's row reads "Qwen Image Max 2512" and links Qwen's 2512 demo space; the checkpoint repo is Qwen/Qwen-Image-2512
  • The arena's listed API price: $20 per 1,000 images

Licence and price

Licence: Apache 2.0 — covers the outputs as well as the weights (HF model card, verified 22 September 2026).

public, ungated; 52,288 downloads (HF API, read 22 September 2026)

Price: $20 per 1,000 images (the arena's price column) — self-hosted, the marginal cost is your own electricity.

The site's catalogue carries it as the safest licence on the shelf: weights and outputs both covered.

Ming-Image-0.1-Design — InclusionAI — open weights

Not stated on the card · Image generation; self-hosted · released 22 September 2026 — the day of this pull

The benchmark

  • Arena Elo 995 with 21,297 appearances — placed by the arena on its release day, retrieved 23 September 2026
  • The arena's listed API price: $30 per 1,000 images

Licence and price

Licence: MIT — verified on the model card, 22 September 2026.

public, ungated; newly published, 0 downloads at pull time (HF API, read 22 September 2026)

Price: $30 per 1,000 images (the arena's price column) — self-hosted, the marginal cost is your own electricity.

Cosmos3-Super-Text2Image (agentic) — NVIDIA — open weights

64B · Image generation; self-hosted, needs a 96 GB card · released 31 May 2026

The benchmark

  • Arena Elo 994 with 11,060 appearances — the "agentic" variant; the base variant reads 983 and the 4-step 970, retrieved 23 September 2026
  • No API price listed on the arena — it is a self-hosted model

Licence and price

Licence: OpenMDW-1.1 (NVIDIA Open Model Development Agreement) — commercially usable, no territory exclusion (catalogue check, 1 September 2026).

public, ungated; 1,997 downloads, last modified 16 September 2026 (HF API, read 22 September 2026)

Price: Self-hosted — no published per-image price.

Runs on: A 96 GB card (the 4U4G class and up).

The two open models, side by side

Qwen-Image-2512Ming-Image-0.1-Design
Arena Elo998 on 2,186 appearances995 on 21,297 appearances
LicenceApache 2.0 — weights and outputsMIT
Released30 December 202517 September 2026
Weights, BF1657.7 GB — 40.9 transformer + 16.6 text encoder + VAE52.9 GB — 34.0 MLLM + 12.3 transformer + 6.2 connector + VAE
Sampling50 steps, CFG 4.0 (the card's example)12 steps, CFG 1.0
Resolutionaspect buckets; no number stated on the card2048² recommended, 1024² faster
Hardwareno requirement stated; 80 GB class for 58 GB of BF16 weightsone 80 GB CUDA GPU, validated
ServingdiffusersvLLM-Omni
Built forgeneral and photoreal generation, editing; 2512 improved human realism and text renderingtext-rich design — UI, infographics, posters; RGBA transparency; editable-PPT skill
Design boardQwen-Image (the original) 946; 2512 not listed1,082 — first, above Ideogram 4.0 Quality (1,052)
Adoption51,443 downloads, 964 likesbrand-new — 221 likes, downloads counter still 0
Measured speedno independent measurement existsno independent measurement exists

Sources, retrieved 23 September 2026: both model cards and repos, the arena, and the design board as shown on Ming's card — that board could not be reached directly, so it is vendor-cited rather than independent. Weights are read from the repositories' own file listings.

Named, not measured

  • The two highest-placed open rows are not free to use: Ideogram 4.0 (1,011) is non-commercial, and FLUX.2 [dev] (1,000) is non-commercial by default — Black Forest Labs sells commercial weights licences that include it (their Platform, Professional and Enterprise tiers, via bfl.ai/licensing, retrieved 23 September 2026). Neither is plotted as free-to-use; only the second has a published path to commercial use.
  • HiDream-O1-Image (979, MIT) is the next eligible open model below the three.

The arena's scale moved again since the previous pull (top row 1,370 then, 1,197 now, retrieved 23 September 2026): the whole board is re-pulled in one pass, never patched row by row. The chart carries the paid leaders, the open rows a business can run, and our pick — modelled, and marked as such on its row.

Video

The benchmark — the full field

ModelElo
Gemini Omni Flash (Google) (cloud API)1513
Dreamina Seedance 2.5 720p (ByteDance) (cloud API)1474
MiniMax H3 — EU excluded by default — obtainable1460
Sora 2 Pro (OpenAI) (cloud API)1368
Kandinsky 5.0 T2V Pro1175
LTX-2 19B1154
Wan 2.2 A14B — our pick1132

source: arena.ai/leaderboard/text-to-video · our retrieval date — the arena publishes no snapshot date of its own

Read the vote counts before the ranking. The newest models sit near the top on a few thousand votes, which is exactly when a leaderboard is least reliable and most quoted — the arena marks those rows Preliminary, and anything under 10,000 votes is treated as provisional here. The chart plots the paid leader and the open-weight models you can run, including the one this site stocks (Wan 2.2, marked); our catalogue's LTX-2.5 is not on this board, so the newest LTX row here is LTX-2 19B.

  • MiniMax H3 (1,460) and HunyuanVideo 1.5 (1,169) are open weights with territory-limited licences. MiniMax's default terms exclude the EU, the UK, the Republic of Korea and the USA, and EU coverage is obtainable through MiniMax's licence process (licence text read 22 September 2026); HunyuanVideo's Tencent Community Licence expressly does not apply in the EU. H3 is shown in full above; HunyuanVideo is named, not plotted.
  • Wan 3.0, 2.7 and 2.6 outscore the open field and have released no weights: API only, and the arena lists them as Proprietary. The same board shows HappyHorse-1.0 as Proprietary and carries no LTX-2.3 or LTX-2.5 row at all — the newest LTX row it measures is LTX-2 19B, which is what this chart plots.

Text-to-video quality by blind pairwise votes — and a field where the open weights sit a long way below the services.

LMArena / arena.ai Text-to-Video Arena — Elo with 95% CI and vote counts · source: arena.ai/leaderboard/text-to-video · our retrieval date — the arena publishes no snapshot date of its own; its rows load client-side, so this pull is a browser read

The three open models, against the one to beat

Gemini Omni Flash — Google — closed, the one to beat

Not published · Text-to-video; API only · released 2026

The benchmark

  • Arena Elo 1,513 ±9 on 26,576 votes, retrieved 22 September 2026
  • The arena's rank-1 row is Gemini Omni 1.1 Flash at 1,516 ±15 on 1,784 votes — under 10,000 votes, so provisional by the arena's own mark

Licence and price

Licence: Cloud only — the top row with a settled sample; the row above it is provisional.

Price: API pricing is per second of video on Google's own schedule; no arena price column exists for video.

The owner confirmed the media references as each arena's #1 at pull time; on this board the #1 row is provisional, so the yardstick is the top settled row and the provisional leader is named beside it.

MiniMax H3 — MiniMax — open weights

33B dense (H3-Omni-Transformer); about 13B of that sits in AdaLN branches whose outputs cache, so inference-only deployment loads the remainder · Text-to-video and image-to-video; 4–15 second clips at 768p and 24 fps, 2K by a second regeneration pass · released 2 August 2026 (licence date); repository last modified 13 August 2026

The benchmark

  • Arena Elo 1,460 ±9 on 11,417 votes — the highest-placed open-weight row on this board, retrieved 23 September 2026

Licence and price

Licence: MiniMax H3 Community Licence — its default grant is limited to an Applicable Territory that excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America (licence text read 22 September 2026). EU coverage is obtainable: MiniMax grants it through a licence process, and the default terms are what they are subject to.

public, ungated; 3,664,216 downloads, last modified 13 August 2026 (HF API, read 23 September 2026)

Price: Self-hosted — no published per-clip price; MiniMax's hosted service is a different product at its own price.

Runs on: A 32 GB card and up: 33B dense at the shipped 4-bit build is about 22 GB before working buffers.

The top open-weight video model this arena has measured, and it moves Kandinsky 5.0 T2V Pro (1,175) down to the second open slot. Its default terms exclude the EU and EU coverage is obtainable through MiniMax's licence process — the entry states both halves rather than leaving either to a footnote.

Kandinsky 5.0 T2V Pro — Kandinsky (Sber) — open weights

Not stated on the arena row · Text-to-video; self-hosted · released 2026

The benchmark

  • Arena Elo 1,175 ±21 on 2,019 votes — the highest-placed open-weight model on the board, retrieved 22 September 2026

Licence and price

Licence: MIT — as listed on the arena's licence column, and the family's Hugging Face cards carry the same licence.

Price: Self-hosted — no published per-clip price.

The highest-placed open row a Danish business can run as it stands — 338 points below the closed yardstick, and below MiniMax H3, whose EU coverage takes a licence process.

LTX-2 19B — Lightricks — open weights

19B · Text-to-video; self-hosted · released 4 August 2026 (HF last modified)

The benchmark

  • Arena Elo 1,154 ±8 on 75,596 votes — the fattest sample of any open row on the board, retrieved 22 September 2026

Licence and price

Licence: LTX-2 community licence — free below $10M annual revenue; above that Lightricks asks you to license it. One binding condition: do not strip its provenance watermarking.

public; 304,795 downloads (HF API, read 22 September 2026)

Price: Self-hosted — no published per-clip price.

Runs on: The 4UXGM class and up (19B plus working buffers).

Wan 2.2 A14B — Alibaba — open weights

14B · Text-to-video; self-hosted · released 7 August 2025 (HF last modified)

The benchmark

  • Arena Elo 1,132 ±15 on 10,458 votes, retrieved 22 September 2026

Licence and price

Licence: Apache 2.0 — no revenue threshold, no territorial exclusion, no watermark condition.

public, ungated; 3,146 downloads (HF API, read 22 September 2026)

Price: Self-hosted — no published per-clip price.

The last Wan released with weights; everything since is API-only.

Named, not measured

  • HunyuanVideo 1.5 (1,169): its Tencent Community Licence expressly does not apply in the EU — named, not plotted. MiniMax H3's default terms exclude the EU as well, and EU coverage is obtainable through MiniMax's licence process; the entry states both halves.
  • Wan 3.0 (1,476), 2.7 and 2.6 outscore the open field and have released no weights: API only.

Music

The benchmark — the full field

ModelElo
Mureka V9 (cloud API)1177.2
Suno V5.5 (cloud API)1171.24
HeartMuLa-oss-3B — our pick — PREDICTED, not measured: placed one measured generation-step below Suno V4.5, whose row it is compared against in its own vendor's listening test (69.93 MOS against 76.08) — the method, inputs and band are under this chart1010
ACE-Step v1 3.5B — PREDICTED, not measured: placed one vendor-test rank below HeartMuLa — 66.66 MOS against its 69.93 in the same listening test — which is a second generation-step from the measured anchor, Suno V4.5 at 1,065.36; the method and band are under this chart955
MusicGen (Meta) — non-commercial licence845.51

source: artificialanalysis.ai/music/leaderboard/instrumental · our retrieval date — the arena publishes no snapshot date of its own

The open side of this chart is MODELLED, not measured — no independent benchmark of an open music model exists, and every open row says so on its face. HeartMuLa's 1,010 is a prediction, and this is how it was derived. Its vendor's blind listening test (HeartMuLa paper, arXiv 2601.10547, Table 13) ranks it second of five — 69.93 MOS against Suno-v4.5's 76.08, both with 95% confidence intervals — and ahead of every open model in that test. Suno V4.5 is measured on this arena at 1,065.36, and the arena's own Suno generations sit 35 to 71 Elo apart (V4.5 to V5 to V5.5), so the placement is one measured generation-step below its reference: a band of 990 to 1,030, taken at its midpoint, 1,010. ACE-Step is placed the same way, one vendor-test rank lower (66.66 MOS): a second generation-step from the measured anchor, band 920 to 990, taken at 955 — a wider band, because the uncertainty compounds the further a placement sits from a measurement. The inputs are named: the vendor's MOS table (the paper, above) and this arena's measured Suno V4.5 and MusicGen rows (retrieved 22 September 2026). MusicGen is measured, but its weights are non-commercial, which is what the ✕ on its row says. Nothing here blends a prediction into a measured row.

  • DiffRhythm2 and YuE sit in the same vendor table (58.33 and 57.93 MOS) and could be placed the same way, but neither has a confirmed commercial licence — named, not plotted. ACE-Step's model card declares Apache 2.0, and the owner accepted that as the statement of record on 23 September 2026, so it plots as open — with the note that the repository ships no separate licence file with the weights.
  • Self-reported, not plotted: HeartMuLa 69.93 in its own listening test, against Suno-v4.5 at 76.08 — the vendor's figures, not this arena's.
  • Stable Audio 2.0 carries open weights under a non-commercial licence, so it is named and never a candidate for a business.

Instrumental music quality by blind pairwise listening votes — the open side is modelled, because no independent measurement of an open music model exists.

Artificial Analysis Instrumental Music Arena — Elo with confidence bounds and appearance counts · source: artificialanalysis.ai/music/leaderboard/instrumental · our retrieval date — the arena publishes no snapshot date of its own

The three open models, against the one to beat

Mureka V9 — Mureka — closed, the one to beat

Not published · Instrumental music generation; API only · released 27 March 2026

The benchmark

  • Arena Elo 1,177.2 on 2,237 appearances, retrieved 22 September 2026

Licence and price

Licence: Cloud only — the arena's #1 at pull time.

Price: Subscription/API pricing on the vendor's own schedule; no arena price column exists for music.

Suno V5.5 sits second at 1,171.24.

No open-weight model in this category has an independent measurement on this board, so the three slots are named below rather than filled with a placement.

Named, not measured

  • No open-weight music model has an independent measurement on this arena — the open rows are modelled and say so on their face. The method, its inputs and its band are stated under the chart; nothing here is blended silently with a measured row.
  • HeartMuLa-oss-3B (Apache 2.0) is the open model this site runs, and ACE-Step v1 3.5B plots beside it as open — its card's Apache 2.0 accepted as the statement of record (owner, 23 September 2026). Both are predicted rows: the method, inputs and bands are under the chart. DiffRhythm2 and YuE sit in the same vendor table but have no confirmed licence — named, not plotted.
  • Nothing open released since 2023 has been measured on this board at all — that gap is the finding; the predicted row is where the open side stands.

The open side of this category is modelled by the owner's ruling of 2026-09-23: no independent measurement of an open music model exists, so a predicted row is placed instead — with its method, inputs and uncertainty stated under the chart, and marked as predicted on the row itself. A slot is still never filled with an unlabelled number.

Voice

The benchmark — the full field

ModelElo
Cartesia Sonic 3.6 (cloud API)1273
ElevenLabs v3 Conversational (cloud API)1196
OpenAI TTS-1 HD (cloud API)1099
Magpie-Multilingual 357M (NVIDIA)1063
Kokoro 82M v1.01061
Chatterbox (Resemble AI)1021

source: artificialanalysis.ai/text-to-speech/leaderboard · our retrieval date — the arena publishes no snapshot date of its own

The widest open/closed gap on the site: the arena's open leaders are licence-blocked — Breeze TTS 2 is research and non-commercial, Fish Audio S2 Pro is research only, and Step Audio EditX ships weights with no licence of their own — so the best open voice a business can actually build on is Magpie-Multilingual at 1,063, with Kokoro at 1,061. The arena's own top row is Cartesia Sonic 3.6 at 1,273. The closed leaders this chart now carries sit far above them: ElevenLabs' v3 Conversational at 1,196 and OpenAI's TTS-1 HD at 1,099 — cloud-only, and plotted so the gap is visible rather than described. Whisper and other speech-to-text models are a different job and stay off this board.

  • Not ranked on this arena, so not stocked: MisoTTS 8B, Orpheus TTS 3B, Higgs Audio v2. Our own voice-overs run on VoxCPM2 — local, cloned from the presenter voice, and not on this arena, so it has no Elo here.
  • The rest of each family stays in the data beneath the chart, not dropped: ElevenLabs Eleven v3 1,167, Turbo v2.5 1,095, Multilingual v2 1,092, Flash v2.5 1,074; OpenAI TTS-1 1,086, GPT-Realtime-2 1,072; Google Gemini 3.8 Flash-Lite TTS 1,235 and Gemini 3.1 Flash TTS 1,199.

Text-to-speech quality by blind pairwise listening votes — the widest open/closed gap on the site, and mostly a licence story.

Artificial Analysis Speech Arena — Elo with 95% CI and sample counts · source: artificialanalysis.ai/text-to-speech/leaderboard · our retrieval date — the arena publishes no snapshot date of its own; its rows load client-side, so this pull is a browser read

The three open models, against the one to beat

Cartesia Sonic 3.6 — Cartesia — closed, the one to beat

Not published · Text-to-speech; API only · released August 2026

The benchmark

  • Arena Elo 1,273 on 1,757 samples, retrieved 23 September 2026
  • The arena's listed API price: $49.0 per 1M characters

Licence and price

Licence: Cloud only — the arena's #1 at pull time.

Price: $49.0 per 1M characters — the arena's own price column, retrieved 23 September 2026

Cost to do a job: $49.0 per 1M characters at the arena's listed rate.

The Inworld row the site carried last (Realtime TTS 1.5 Max, 1,209.6) has moved down the board: Realtime TTS-2 sits fourth at 1,245 and Realtime TTS-2 Flash eighth at 1,210. The two closed rows this chart gained on 23 September — ElevenLabs' v3 Conversational and OpenAI's TTS-1 HD — are the best each lab fields on this arena.

Magpie-Multilingual 357M — NVIDIA — open weights

357M · Text-to-speech; self-hosted, runs on almost anything · released February 2026

The benchmark

  • Arena Elo 1,063 on 1,971 samples — the highest-placed open model a business can build on, retrieved 23 September 2026

Licence and price

Licence: NVIDIA Open Model Licence — perpetual, worldwide, royalty-free, no territorial exclusion (model card, verified 22 September 2026).

public, ungated; 6,956 downloads, last modified 9 September 2026 (HF API, read 22 September 2026)

Price: Self-hosted — no published per-character price.

Kokoro 82M v1.0 — hexgrad — open weights

82M — runs on a CPU · Text-to-speech; self-hosted · released January 2025

The benchmark

  • Arena Elo 1,061 on 5,232 samples, retrieved 23 September 2026
  • The arena's listed API price: $0.7 per 1M characters

Licence and price

Licence: Apache 2.0 (model card, verified 22 September 2026).

public, ungated; 11,722,492 downloads (HF API, read 22 September 2026)

Price: $0.7 per 1M characters at the arena's listed rate — or nothing but electricity, self-hosted.

Chatterbox — Resemble AI — open weights

0.5B · Text-to-speech with voice cloning; self-hosted · released May 2025

The benchmark

  • Arena Elo 1,021 on 4,600 samples, retrieved 23 September 2026
  • The arena's listed API price: $25.0 per 1M characters

Licence and price

Licence: MIT — what it says is yours to use (model card, verified 22 September 2026).

public, ungated; 1,794,542 downloads, last modified 10 June 2026 (HF API, read 22 September 2026)

Price: $25.0 per 1M characters at the arena's listed rate — or nothing but electricity, self-hosted.

Named, not measured

  • The open leaders are licence-blocked and named, never candidates: Breeze TTS 2 (1,204 — research and non-commercial licence), Fish Audio S2 Pro (1,120 — research only), Step Audio EditX (1,094 — the code is Apache 2.0 but the weights carry no licence of their own).
  • Voxtral TTS (1,076) is held off the candidate list: the licence on its model card and the licence in this site's own catalogue disagree, and that is being settled before it can be recommended.
  • Whisper large-v3 and other speech-to-text models are a different job and stay off this board.

The category's closed reference moved with the arena: the Inworld row named when this structure was approved (Realtime TTS 1.5 Max) is no longer the top row, so the yardstick is the current #1.

Sources this page stands on

Refresh protocol: a source whose version moves is re-pulled whole — never patched row by row — and every chart prints its own retrieval date, so a stale snapshot is visible rather than assumed current. The Office/CAD figure is computed the same way for any model: the five published components from Artificial Analysis's model payload at the weights named on its chart, one row per model at its highest reasoning effort — so a new model gets a figure the day AA publishes its components.

The five machines and what each one runs · the two models we run and mod

A number here you want checked?

Every score on this page names its source, and the thin ones say so. If one of them is carrying a decision you are making, write to us and we will show the measurement behind it.

Email us · Register for a role