Week 40, 2026Published September 26, 2026
AI Model Power Rankings
Opus 5.5 confirms it: first on the Arena voter boards, too. A quieter week after last Tuesday's pile-up, but the first real votes are in. Claude Opus 5.5 went straight to #1 on LMArena's text and WebDev boards, and GPT-6 Sol proved its value case with voters. Google, meanwhile, says Gemini 4 is in post-training and coming 'as soon as possible'. We'll believe the date when we see it.
Arena WebDev voters put it level with Claude Opus 5 at less than half the price. Its value argument is real, not just OpenAI's word.
Lost the WebDev crown to Opus 5.5 within days, and sits outside the top 25 on the text board. Still elite for automation, but no longer the default.
01Power Rankings
The best general-purpose AI models right now, with quality, price and momentum rolled into one score.
| # | Movement± | Model | Price / 1M tokens | Score |
|---|---|---|---|---|
| 1 | — | Claude Opus 5.5Anthropic Now #1 with the voters as well as the benchmarks, on LMArena's text and WebDev boards. The text lead is provisional on a small sample, but nothing has knocked it off yet. |
$4 in / $20 out | 97 |
| 2 | — | Claude Fable 5.1Anthropic Still in the text top five and #3 on WebDev. Brilliant, but hard to justify at 2.5x Opus 5.5's price. |
$10 in / $50 out | 95 |
| 3 | — | GPT-6 AstraOpenAI Lost WebDev #1 to Opus 5.5 and trails on the text board. It still leads on automation work, which is where it earns its price. |
$10 in / $50 out | 93 |
| 4 | — | GPT-6 SolOpenAI Voters rank it alongside Claude Opus 5 for web coding at a fraction of the cost. The best-value model near the frontier. |
$2 in / $10 out | 92 |
| 5 | — | Gemini 3.8 FlashGoogle Holding a text top-10 spot for 75 cents per million input tokens. The price doubles on January 1. |
$0.75 in / $3.75 out | 90 |
| 6 | — | Muse Spark 1.3Meta The model is fine. Meta's Muse agent being locked out of Amazon is a product problem, not a quality one. |
$1.25 in / $4.25 out | 89 |
| 7 | — | Gemini 3.1 ProGoogle Still in preview and outranked by its Flash sibling. Google says Gemini 4 is in post-training, which is the upgrade this slot needs. |
$2 in / $12 out | 87 |
| 8 | — | Kimi K3Moonshot AI Top 10 on WebDev and level with Gemini 3.1 Pro on the text board. Still the strongest model from outside the US frontier labs. |
Not published | 86 |
| 9 | — | Grok 4.7xAI Its first WebDev showing is respectable, just outside the top 12. Good value, but not in the lead pack. |
$2 in / $6 out | 84 |
| 10 | — | GLM-5.3Z.ai · open Still the best-proven open-weight model, but Xiaomi's MiMo now matches it on the text board. |
Not published | 84 |
| 11 | — | GPT-6 LunaOpenAI Early testers report it keeping pace with Sol on some coding evals at a fraction of the price. We'd like independent numbers before we move it up. |
$0.1 in / $0.5 out | 82 |
| 12 | — | Qwen3.8-FlashAlibaba Cheap and multilingual. Alibaba's voice price cuts this week make its wider stack even better value. |
$0.15 in / $0.47 out | 78 |
02Coding
For developers and 'vibe coders': writing, fixing and shipping real software.
| # | Movement± | Model | Price / 1M tokens | Score |
|---|---|---|---|---|
| 1 | — | Claude Opus 5.5Anthropic #1 on Arena WebDev, well clear of GPT-6 Astra, and top of Terminal-Bench 4.0. The default pick for coding. |
$4 in / $20 out | 98 |
| 2 | — | GPT-6 AstraOpenAI Now #2 on WebDev and still the strongest when the job involves driving a terminal. |
$10 in / $50 out | 94 |
| 3 | — | Claude Fable 5.1Anthropic #3 on WebDev. Excellent, but Opus 5.5 is better and cheaper for most code. |
$10 in / $50 out | 93 |
| 4 | — | GPT-6 SolOpenAI Enters the WebDev top five on real votes, on par with Claude Opus 5. Live in Codex and GitHub Copilot. |
$2 in / $10 out | 91 |
| 5 | — | Kimi K3Moonshot AI Top 10 on WebDev with a big vote count behind it. A popular choice for cheap coding agents. |
Not published | 84 |
| 6 | — | GLM-5.3Z.ai · open The open-weight coding pick, though MiMo-V2.6-Pro scores level with it on WebDev. |
Not published | 83 |
| 7 | — | Gemini 3.8 FlashGoogle Great for quick edits and autocomplete-style work. Wait for Gemini 4 if you need an agent. |
$0.75 in / $3.75 out | 82 |
| 8 | — | Grok 4.7xAI Edges out 4.6 on WebDev, but 38% on Terminal-Bench 4.0 still rules it out for serious agents. |
$2 in / $6 out | 75 |
03Best Value
The most capability per dollar. If you're building an app or watching a budget, start here.
| # | Movement± | Model | Price / 1M tokens | Score |
|---|---|---|---|---|
| 1 | — | Gemini 3.8 FlashGoogle Top-10 text quality at Flash prices, so enjoy it before January. |
$0.75 in / $3.75 out | 95 |
| 2 | ▲1 | GPT-6 SolOpenAI Arena voters now back up the pitch: Opus 5-level web coding for a fifth of Astra's price. |
$2 in / $10 out | 93 |
| 3 | ▼1 | GPT-6 LunaOpenAI $0.10 / $0.50. Early coding results look strong, but they're mostly OpenAI's and testers' numbers so far. |
$0.1 in / $0.5 out | 92 |
| 4 | — | Muse Spark 1.3Meta Strong conversation at $1.25 / $4.25, and the Contributor tier is even cheaper. |
$1.25 in / $4.25 out | 88 |
| 5 | — | Qwen3.8-FlashAlibaba $0.15 / $0.47 with multilingual strength, and Alibaba's audio APIs just got much cheaper too. |
$0.15 in / $0.47 out | 86 |
| 6 | — | Mercury 2.5Inception Very fast diffusion text generation at $0.20 / $0.75. |
$0.2 in / $0.75 out | 80 |
| 7 | — | Grok 4.7xAI Its $6 output tokens are among the cheapest at this quality level. |
$2 in / $6 out | 78 |
04Open-Weight Watch
Models you can download, run privately and customise.
| # | Movement± | Model | Price / 1M tokens | Score |
|---|---|---|---|---|
| 1 | — | GLM-5.3Z.ai · open Still the strongest open-weight model on the most evidence. It has a custom license, so read it. |
Not published | 91 |
| 2 | — | MiMo-V2.6-ProXiaomi · open Its first Arena scores put it level with GLM-5.3 on both text and WebDev. Fewer votes so far, but a serious debut. |
Not published | 88 |
| 3 | — | Ternary Bonsai 2 27BPrismML · open Apache 2.0 and small enough for real-world hardware. |
Not published | 78 |
05Image Generation
Making pictures from text, based on LMArena text-to-image votes plus our hands-on notes.
| # | Movement± | Model | Price / 1M tokens | Score |
|---|---|---|---|---|
| 1 | — | GPT Image 2.5OpenAI Still holds the top three spots on the leaderboard, with the gap unchanged. |
— | 97 |
| 2 | — | MAI-Image 2.6Microsoft AI Microsoft's in-house model remains the best non-OpenAI option. |
— | 88 |
| 3 | — | Reve 2.1Reve Moves up to fifth on the leaderboard. The indie favourite for aesthetics. |
— | 85 |
| 4 | — | Grok Imagine Image 2.0xAI Slipped a place as more votes came in, but still top six on its low setting. |
— | 83 |
| 5 | — | Gemini 3.1 Flash ImageGoogle Ranks lower on raw generation, but it's the best for editing an image through conversation. |
— | 82 |
| 6 | — | Seedream 5.0 ProByteDance Powers the CapCut/Dreamina creator ecosystem, and Alibaba's Qwen-Image 3.0 Pro is now level with it. |
— | 80 |
06Video Generation
Text-to-video, and what to use now that Sora is gone.
| # | Movement± | Model | Price / 1M tokens | Score |
|---|---|---|---|---|
| 1 | — | Gemini Omni 1.1 FlashGoogle Google holds the top two text-to-video spots with Omni Flash and its 1.1 update. |
— | 95 |
| 2 | — | Veo 3.1Google The safest all-rounder, with native audio and 4K. |
— | 92 |
| 3 | — | FLUX 3 VideoBlack Forest Labs Holds #3 on the leaderboard, though on a small vote count. |
— | 90 |
| 4 | — | Kling 3.0Kuaishou The value king at about 11–14 cents per second, plus storyboard mode. |
— | 89 |
| 5 | ▲1 | Grok Imagine Video 1.5xAI Up to #4 in agent mode. Fast and fun. |
— | 87 |
| 6 | ▼1 | Seedance 2.5ByteDance Dipped below its own predecessor as votes piled in. Still the best realistic human motion for creators. |
— | 86 |
| 7 | — | Wan 3.0Alibaba Dropped out of the top five as its vote count doubled. Still the tinkerer's pick. |
— | 83 |
| 8 | — | Runway Gen-4.5Runway Professionals still choose it for camera control. |
— | 82 |
The Wire · what happened this week
Alibaba launches the five-model Qwen-Audio 3.1 stack and cuts voice API prices: roughly 70% off TTS, 85% off Realtime and up to 95% off speech recognition.
Scores are our editorial blend of public leaderboards (LMArena, Artificial Analysis and published benchmarks), price and hands-on use. Read the methodology →