How We Rank AI Models
Our scores are editorial judgements, but they're built on public evidence, and we show our working.
The short version
Every model in the standings gets a Standings Score from 0 to 100 in each category it competes in. The score blends four inputs:
| Input | Weight | Where it comes from |
|---|---|---|
| Crowd preference | ~35% | Blind head-to-head voting leaderboards such as LMArena |
| Benchmarks | ~30% | Independent indexes like Artificial Analysis, plus the benchmarks labs publish, which we treat with appropriate scepticism |
| Price & access | ~20% | List API prices, free tiers and how easy the model is to actually use |
| Hands-on & momentum | ~15% | Our own testing, reliability reports, and whether a model is improving or stagnating |
Weights shift by category. In Best Value, price counts for much more. In Open-Weight Watch, the license and whether you can realistically run the model matter a lot.
What the scores mean
- 95+ Clear leader. It's the default recommendation for that category.
- 90–94 Elite. Pick based on price, ecosystem or specific strengths.
- 85–89 Very good, and often the smart value choice.
- 80–84 Solid, with real trade-offs.
- Below 80 Niche, a budget pick, or not recommended for most people.
Movement arrows
Arrows (▲▼) compare a model's rank with the previous week's edition. NEW means the model wasn't ranked in that category last week. Our first edition (Week 39, 2026) has no arrows because there was nothing to compare it with yet.
What we don't do
- We don't accept payment for rankings or placement.
- We don't rank models we can't find reliable public information about.
- We don't treat any single benchmark as gospel. Labs choose which results to publish, and leaderboards can be gamed.
Corrections
AI moves fast and we'll sometimes be wrong. If you spot an error, tell us. We correct factual mistakes promptly.
Last updated September 25, 2026