The Open-Weight League
The weekly form guide to open-weight models worth running on your own infrastructure.
INT Intelligence · COST $ per task · OPEN Openness / 18 · SPD Speed t/s · PTS Blend 50·20·20·10
Season 2026/27 · Matchday 04| # | Club | Nation | INTIntelligence Index · reasoning, knowledge & agentic ability · AA Index v4.1.1.1 | COSTCost per task · USD · scored on a log scale, cheapest = 100 | OPENOpenness Index · weights & transparency, scored out of 18 | SPDOutput Speed · provider-median throughput · a proxy for how heavy the model is to serve | PTSLeague Points · blended score · INT 50 · COST 20 · OPEN 20 · SPD 10 |
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0731DeepSeek | CHN | 52 | 0.03 | 8 | 106 | NEW70 |
| 2 | DeepSeek V4 ProDeepSeek | CHN | 45 | 0.05 | 9 | 63 | ▲262 |
| 3 | MiMo V2.5Xiaomi | CHN | 38 | 0.01 | 7 | 93 | ▲1062 |
| 4 | MiMo V2.5 ProXiaomi | CHN | 43 | 0.03 | 7 | 70 | ▲460 |
| 5 | GLM-5.2Z.AI · Zhipu AI | CHN | 53 | 0.31 | 8 | 109 | ▼360 |
| 6 | HYHy3Tencent | CHN | 42 | 0.04 | 8 | 67 | NEW59 |
| 7 | Kimi K3Moonshot AI | CHN | 60 | 0.84 | 7 | 39 | ▼659 |
| 8 | NXNex-N2-ProNex AGI | CHN | 42 | — | 7 | 129 | ▲257 |
| 9 | InklingThinking Machines | USA | 42 | — | 7 | 82 | ▼256 |
| 10 | Inkling SmallThinking Machines | USA | 41 | 0.07 | 6 | 143 | NEW56 |
| 11 | Nemotron 3 UltraNVIDIA | USA | 38 | 0.38 | 15 | 143 | ▼855 |
| 12 | MiniMax M3MiniMax | CHN | 45 | 0.14 | 6 | 97 | ▼655 |
| 13 | SFStep 3.7 FlashStepFun | CHN | 31 | 0.09 | 7 | 364 | ▼453 |
| 14 | Nemotron 3 NanoNVIDIA | USA | 15 | 0.02 | 15 | 253 | NEW52 |
| 15 | HNHyperNova 60BMultiverse Computing | ESP | 18 | 0.02 | 7 | 401 | ▲150 |
| 16 | Kimi K2.7 CodeMoonshot AI | CHN | 43 | 0.22 | 5 | 43 | ▼449 |
| 17 | Gemma 4 26B A4BGoogle | USA | 26 | 0.04 | 7 | — | NEW48 |
| 18 | Nemotron 3 SuperNVIDIA | USA | 26 | 0.23 | 15 | 139 | ▼748 |
| 19 | Kimi K2.6Moonshot AI | CHN | 35 | — | 6 | 40 | NEW46 |
| 20 | LCLongCat 2.0LongCat · Meituan | CHN | 34 | 0.12 | 7 | 41 | NEW46 |
Key — what the columns mean
INT
Intelligence Index. Reasoning, knowledge & agentic ability · AA Index v4.1.1.1
COST
Cost per task. USD, includes how many tokens the model uses per job · log scale, cheapest = 100
OPEN
Openness Index. Weights & transparency, scored out of 18
SPD
Output Speed. Provider-median throughput · a proxy for serving weight
PTS
League Points. Blended score · INT 50 · COST 20 · SPD 10 · OPEN 20
Methodology
Every underlying score on this table comes from Artificial Analysis, an independent AI benchmarking organisation. We do not run our own evaluations and we do not adjust theirs.
INT is the Artificial Analysis Intelligence Index (v4.1.1): reasoning, knowledge, coding and agentic ability. COST is their cost per task in USD — price and token appetite combined: what one unit of work actually costs across serving providers. SPD is provider-median output speed in tokens per second — a proxy for how heavy the model is to serve; your own numbers depend on your hardware. OPEN is their Openness Index — weights availability, licence and transparency — scored out of 18.
PTS is the only number we compute. INT and SPD are scored relative to the matchday’s best (best = 100); OPEN is scored against its fixed 18-point maximum; COST is scored on a log scale from the matchday’s cheapest (= 100) to its priciest (= 0), because prices differ by multiples, not margins. A model served free ($0.00) has no market price and takes a dash. The four are blended INT 50% · COST 20% · OPEN 20% · SPD 10% and rounded to whole points — ties are ordered by the unrounded blend. The method is stated here precisely so the table can be checked, argued with, or rebuilt by anyone.
A dash means Artificial Analysis has not published that score yet. A club missing one column is still seated: that column’s weight is redistributed pro-rata across its published scores. A club missing two or more isn’t seated until the data exists.
† marks a provisional seat. A newly released model may be seated before its weights are public when its lab has publicly committed the weights or has an established open-weights record, and Artificial Analysis has published every column except openness. Three matchdays maximum — the weights land or the seat lapses. This rule seated Kimi K3 at launch; its weights shipped on schedule and the mark came off.
Matchday 04 rule change, announced here. Artificial Analysis discontinued its Coding Index this week, so the COD column retires with it — coding remains scored inside INT, where Terminal-Bench v2.1 and SciCode make up 24% of the index. In its place the League now scores COST: Artificial Analysis’ cost per completed task, which prices both the per-token rate and how many tokens a model burns to finish a job. One club per model: a new checkpoint on the same Artificial Analysis model page supersedes the old (as DeepSeek’s 0731 release supersedes the earlier V4 Flash). Point totals are not comparable with earlier matchdays.
Valarian builds no models and holds no stake in any lab on this table. ACRA runs whichever club you pick — so we have no reason to favour one. The table exists because our customers keep asking the same question: which open-weight models are actually good now?
Valarie’s squad
Match-fit today: validated in Valarie, uploaded privately into ACRA, and served from infrastructure you control.
DeepSeek V4 Pro
DeepSeek · CHN
live in ACRADeploy your stack
gpt-oss-120b
OpenAI · USA
live in ACRADeploy your stack
Gemma 4 31B
Google · USA
live in ACRADeploy your stack
Any open-weight club is loadable · Control Plane → Private AI → upload