Introducing the Open-Weight League
Every Friday, Valarian ranks the best open-weight models in the world. Independent data from Artificial Analysis, a published formula, one table. How the League works, in full.
Today we're launching The Open-Weight League — a live league table for open-weight AI models, published at valarian.com/owl and updated every Friday.
Twenty clubs. Four stats — intelligence, cost, openness, speed. One blended score that sets the table, with a frontier at the top and a relegation zone at the bottom. We built it as a form guide rather than a static list, because that's what the open-weight world actually is now: models improve, slip, and get displaced week by week, and a snapshot can't show you form.
We built it because our customers keep asking the same question: which open-weight models are actually good right now? Open weights are the sovereign choice — models you can download and run on infrastructure you control, with keys only you hold. The League is that argument, in table form. Season 2026/27 kicks off on World Cup final weekend, which felt right.
The short version
Every Friday, we rank the best open-weight models in the world. The scores come from Artificial Analysis, an independent AI benchmarking organisation. The ranking formula is ours, and it's published below — in full. Whether a model runs in Valarie has no bearing on its position. That's the whole game.
Who gets in
The League is open weights only. To qualify, a model must have downloadable weights you can run on infrastructure you control — your cluster, your cloud, your air-gapped estate.
Concretely, a club qualifies when:
- Its weights are publicly downloadable (Hugging Face or equivalent),
- Its licence permits self-hosted deployment,
- It is benchmarked by Artificial Analysis.
Closed, API-only models are not eligible, however capable. This isn't a ranking of all AI; it's a ranking of the AI you can actually own. Closed models play a different sport.
Two edge cases, handled in the open. First, the provisional seat (†): a newly released model may be seated before its weights are public when its lab has publicly committed the weights or has an established open-weights record, and Artificial Analysis has published every column except openness. Three matchdays maximum — the weights land or the seat lapses. (Rule formalised 31 Jul 2026; it seated Kimi K3 at launch, whose weights shipped on schedule.) Second, the dash: where Artificial Analysis hasn't published one of a club's four scores, that column's weight is redistributed across the club's published scores, never invented. A model missing two or more columns isn't seated until the data exists.
From all qualifying models, the top twenty by League Points make the table each Matchday. Where a lab has several qualifying models, each model is its own club, scored in its highest published reasoning configuration. Labs can, and do, occupy multiple places in the table.
Where the numbers come from
Every performance score in the League comes from Artificial Analysis, an independent AI benchmarking organisation. They evaluate models continuously across intelligence, price, and speed, and publish their methodology in full.
We chose an external source deliberately. We don't run our own evaluations, we don't adjust their numbers, and we can't nudge a score. You don't have to trust us on the data, only on the formula. And the formula is right here.
Reading the table — the four stats
Every column answers a different question a deployer actually asks: how smart is it (INT), what does it cost to run (COST), how much of it do I truly get (OPEN), and how heavy is it to serve (SPD). All four come directly from Artificial Analysis, which means every number in the table is third-party. The only thing Valarian adds is the blend.
INT — Intelligence Index
Artificial Analysis' composite measure of general capability, the model's overall rating. It answers one question: how smart is this model across the full range of real work?
What goes into it (Intelligence Index v4.1.1 — nine evaluations, weighted by category):
- Agentic tasks (34%) — the largest share: can the model operate in a loop, use tools, and complete long-horizon work? GDPval-AA v2 (real tasks from 44 occupations, where the model must produce actual deliverables: spreadsheets, memos, presentations) and τ³-Banking (multi-step customer-service workflows).
- Coding (24%) — Terminal-Bench v2.1 (completing engineering tasks in a real command-line environment) and SciCode (writing working scientific code).
- Scientific reasoning (24%) — the hardest academic benchmarks: GPQA Diamond (PhD-level science questions), Humanity's Last Exam, CritPt (frontier physics).
- General knowledge & reliability (18%) — including AA-Omniscience, which doesn't just reward correct answers: it penalises hallucinations, so a model that confidently makes things up scores worse than one that knows what it doesn't know. AA-LCR tests long-context reasoning: holding and using large amounts of information at once.
In short: INT rewards models that can reason, act, code, and know their own limits, rather than models tuned to ace one famous benchmark.
COST — Blended Price
Artificial Analysis' cost per completed task for the model, in USD, across the providers serving it. It answers: what does one unit of real work actually cost?
Why market cost, for a league about self-hosting: competitive serving prices track the real compute economics of a model — the thing you inherit when you run it yourself — far better than any spec sheet. And per-task beats per-token: it prices both the rate and how many tokens a model burns to finish a job, so a verbose model can't hide behind a cheap sticker. A model served only for free ($0.00) has no market price and takes a dash.
OPEN — Openness Index
Artificial Analysis' Openness Index, scored out of 18. It measures how much of itself a model actually publishes — not how good the model is, but how yours it can be.
Three dimensions, each worth up to 6 points, summed to a maximum of 18. Within each dimension, two subcomponents are scored 0–3 against defined openness "archetypes" — one for access, one for licence:
- Model availability (0–6) — are the weights downloadable (0–3), and under what licence (0–3)? A permissive licence (use it, modify it, deploy it commercially) scores higher than a restrictive one.
- Data transparency (0–6) — does the lab disclose and license what the model was trained on? Access (0–3) and licence (0–3), scored separately for pre-training and post-training data, then averaged so neither is double-counted.
- Methodology transparency (0–6) — is the build documented and usable? Disclosure (0–3): technical reports, training recipes, tooling. Licence (0–3): whether you may actually use them.
Why this deserves a column: two models can both be "open weights" and sit at opposite ends of this scale. One ships weights, data recipes, and a full technical report under a permissive licence; another drops weights under a restrictive licence and explains nothing. For anyone deploying models on their own infrastructure, that gap is the gap between something you can audit, reproduce, and truly own, and a black box you happen to host. No model evaluated at the Index's launch reached the full 18: full openness is genuinely rare, which is exactly why we award points for it.
SPD — Output Speed
Provider-median output tokens per second, as benchmarked by Artificial Analysis — kept deliberately at the table's smallest weight, because it is a proxy, not a promise: your throughput in your own infrastructure depends on your hardware and serving stack, not on any provider's.
What the proxy is good for: how heavy a model is to serve. When no provider on earth serves a model quickly, that is real information about what your own cluster is in for — and when a model streams fast everywhere, it's light. Read it as serving weight, not as the number you'll get.
How League Points are calculated
PTS is the blended score that sets the table — the one editorial layer we add on top of Artificial Analysis data:
PTS = INT × 50% + COST × 20% + OPEN × 20% + SPD × 10%
The four stats live on different scales, so each is normalised to 0–100 before weighting. Stated exactly, so the table can be rebuilt from this paragraph alone: INT and SPD are scored relative to the matchday's best (leader = 100); OPEN is scored against its fixed 18-point maximum; and COST is scored on a log scale from the matchday's cheapest (= 100) to its priciest (= 0) — logarithmic because prices differ by multiples, not margins. Where a column is unpublished (shown as a dash), its weight is redistributed pro-rata across the club's published scores. Blend, round to whole points; ties are ordered by the unrounded blend.
Why these weights
We could have hidden the weighting. Instead, here's the argument:
- Intelligence, 50% — it's what you're buying. General capability is the product — and coding lives here too: Terminal-Bench v2.1 and SciCode make up 24% of the index.
- Cost, 20% — capability you can't afford to run is capability on paper. Open weights win on economics, so the league scores economics.
- Openness, 20% — this is an open-weight league: a model that publishes more of itself — weights, licence, data, methods — is structurally more sovereign, more auditable, and more yours.
- Speed, 10% — the smallest weight, deliberately: it's a proxy for serving heaviness rather than a promise about your infrastructure, and it's weighted accordingly.
Reasonable people can defend different weights. These are ours, they're published, and they apply identically to every club.
Stability rules
- Any change to the methodology — weights, normalisation, eligibility — is announced before it takes effect, and past tables are archived so they remain comparable.
- When Artificial Analysis revises an index version (e.g. Intelligence Index v4.1 → v5), we adopt the new version at the next Matchday, say so in the changelog, and never mix versions within a single table.
- If a score is corrected upstream by Artificial Analysis, we recompute at the next Matchday. We don't retroactively edit published tables; the archive shows what was known at the time.
Change log — 7 Aug 2026: Artificial Analysis discontinued its standalone Coding Index, so the COD column retired with it (coding remains scored inside INT). The League now scores COST — blended price per 1M tokens — and the blend is INT 50 · COST 20 · OPEN 20 · SPD 10 (COST = AA cost per task). Point totals before and after this date are not comparable.
The zones
FRONTIER (1st). Top of the table. The best open-weight model in the world right now. QUALIFICATION (2nd–5th). The chasing pack: models with a credible claim to the top spot. RELEGATION (18th–20th). On the bubble. A qualifying model outside the table — or a benchmarked model completing its data — can displace them at any Matchday.
Movement arrows (▲▼) show position change versus the previous Matchday. A dash means the club held its position. Matchday 01 shows neither — there is no previous Matchday yet.
Cadence
The table updates every Friday — one Matchday per week, computed from Artificial Analysis' latest published data at the time of ranking. Season 2026/27 runs July 2026 to June 2027.
Independence
Some clubs in the League are validated in Valarie, our sovereign agent, and available to run in ACRA, our platform. This has zero influence on the rankings: positions are computed from Artificial Analysis' published data and the weighting formula above, and nothing else. No pay-to-play, no partner boosts, no exceptions. We built the League because we believe open weights are the sovereign choice. The table is the argument, and it only works if it's honest.
Questions we get
Why isn't [closed model] in the table? Because you can't download it. The League ranks models you can own and run yourself. That includes Meta's Muse Spark line — capable, but API-only, and the first Meta release without open weights. Different sport.
What's a provisional seat (†)? A club seated before its weights are public — allowed when the lab has publicly committed the weights (or has an established open-weights record) and every column except openness is already benchmarked. It lasts at most three matchdays: the weights land or the seat lapses. Kimi K3 is the precedent — seated at launch on Moonshot's public commitment, weights shipped on schedule, mark came off. The † means the League can seat a major release on day one without ever inventing a number.
A club shows a dash in a column — why? Artificial Analysis hasn't published that score yet. Rather than invent a number, the dash's weight is redistributed across the club's published scores, and the cell fills in the day the data exists. Models missing two or more columns aren't seated until then.
A model's benchmark score looks different elsewhere. Different evaluators, different harnesses, different numbers. We use Artificial Analysis exclusively, for consistency; their methodology is linked above.
Can a lab submit a model? No submission needed: if Artificial Analysis benchmarks it and the weights are downloadable, it's automatically eligible.
Who decides all this? The formula and rules on this page decide. We wrote them, we publish them, and we're bound by them like everyone else.
The table is live now at valarian.com/owl. See you Friday.
Source: Artificial Analysis (artificialanalysis.ai). The Open-Weight League is an editorial product of Valarian. Questions or corrections: get in touch.