How models answer specifically in English: the score is built from blind comparisons in which people picked the better English reply. This is a collection about the language, not about ability — a model strong in the overall ranking can lose here to one that handles English phrasing better.
| # | Model | Developer | Arena Score, English | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 1,505.8 1,501–1,510 | $5 | $25 | 1,000K |
| 2 | Claude Opus 5 ≈ | Anthropic | 1,500.16 1,482–1,519 | $5 | $25 | 1,000K |
| 3 | Claude Fable 5 ≈ | Anthropic | 1,498.82 1,491–1,507 | $10 | $50 | 1,000K |
| 4 | Claude Opus 4.7 ≈ | Anthropic | 1,487.52 1,482–1,493 | $5 | $25 | 1,000K |
| 5 | Gemini 3.6 Flash ≈ | Google DeepMind | 1,483.77 1,472–1,495 | $1.5 | $7.5 | 1,049K |
| 6 | Muse Spark 1.1 ≈ | Meta AI | 1,480.18 1,471–1,490 | $1.25 | $4.25 | 1,049K |
| 7 | Gemini 3 Pro ≈ | Google DeepMind | 1,479.64 1,474–1,485 | $1.6 | $9.6 | 1,049K |
| 8 | Gemini 3.5 Flash ≈ | Google DeepMind | 1,479.43 1,470–1,488 | $1.5 | $9 | 1,049K |
| 9 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 1,478.03 1,474–1,483 | $2 | $12 | 1,049K |
| 10 | GLM 5.1 ≈ | Zhipu AI / Z.ai | 1,477.83 1,472–1,484 | $0.06 | $0.22 | 205K |
| 11 | MiMo V2.5 Pro ≈ | Xiaomi | 1,477.62 1,472–1,483 | $0.44 | $0.87 | 1,050K |
| 12 | Kimi K3 ≈ | Moonshot AI (Kimi) | 1,477.38 1,461–1,493 | $2 | $8 | 1,049K |
| 13 | ERNIE 5.1 ≈ | Baidu (Ernie) | 1,476.98 1,471–1,483 | $0.75 | $3 | 119K |
| 14 | Claude Sonnet 4.6 ≈ | Anthropic | 1,473.19 1,468–1,478 | $3 | $15 | 1,000K |
| 15 | GLM 5.2 ≈ | Zhipu AI / Z.ai | 1,468.69 1,461–1,476 | $0.42 | $1.32 | 1,049K |
| 16 | Gemini 3 Flash ≈ | Google DeepMind | 1,465.02 1,459–1,471 | $0.4 | $2.4 | 1,049K |
| 17 | GPT-5.6 Sol ≈ | OpenAI | 1,460.56 1,450–1,471 | $5 | $30 | 1,050K |
| 18 | Claude 4.5 Opus ≈ | Anthropic | 1,460.01 1,456–1,464 | $5 | $25 | 200K |
| 19 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 1,458.16 1,452–1,464 | $0.22 | $1.14 | 262K |
| 20 | DeepSeek V4 Pro ≈ | DeepSeek | 1,457.95 1,452–1,463 | $0.44 | $0.87 | 1,049K |
| 21 | Claude Opus 4.8 ≈ | Anthropic | 1,457.59 1,451–1,464 | $5 | $25 | 1,000K |
| 22 | Qwen3.7 Plus ≈ | Alibaba (Qwen) | 1,457.42 1,450–1,464 | $0.5 | $3 | 1,000K |
| 23 | Gemini 2.5 Pro ≈ | Google DeepMind | 1,457.33 1,454–1,461 | $1.25 | $10 | 1,049K |
| 24 | Grok 4.20 (Reasoning) ≈ | xAI | 1,457.27 1,452–1,462 | $2 | $6 | 2,000K |
| 25 | GLM 5 ≈ | Zhipu AI / Z.ai | 1,456.77 1,451–1,463 | $0.48 | $1.9 | 205K |
| 26 | Grok 4.5 ≈ | xAI | 1,455.92 1,446–1,465 | $2 | $6 | 500K |
| 27 | GLM 4.7 ≈ | Zhipu AI / Z.ai | 1,454.08 1,446–1,462 | $0.15 | $0.8 | 205K |
| 28 | MiMo V2 Pro ≈ | Xiaomi | 1,454.04 1,448–1,460 | $0.44 | $0.87 | 1,049K |
| 29 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 1,454.04 1,445–1,464 | $0.7 | $3.1 | 203K |
| 30 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 1,453.91 1,449–1,458 | $0.3 | $1.9 | 256K |
| 31 | Grok 4.20 Multi Agent Beta 0309 ≈ | xAI | 1,452.15 1,447–1,457 | $2 | $6 | 2,000K |
| 32 | Hy3 ≈ | Tencent (Hunyuan) | 1,451.61 1,435–1,468 | $0.13 | $0.53 | 262K |
| 33 | Claude Sonnet 4.5 ≈ | Anthropic | 1,451.19 1,447–1,455 | $3 | $15 | 1,000K |
| 34 | Seed 2.0 Pro ≈ | ByteDance (Doubao) | 1,449.94 1,445–1,455 | $0.5 | $3 | 131K |
| 35 | Qwen3.6 Max Preview ≈ | Alibaba (Qwen) | 1,449.79 1,438–1,462 | $1.3 | $7.8 | 262K |
| 36 | MiniMax M3 ≈ | MiniMax | 1,449.67 1,443–1,456 | $0.3 | $1.2 | 1,049K |
| 37 | GLM 4.6 ≈ | Zhipu AI / Z.ai | 1,448.2 1,443–1,453 | $0.29 | $1.14 | 205K |
| 38 | Claude Sonnet 5 ≈ | Anthropic | 1,448.18 1,440–1,456 | $2 | $10 | 1,000K |
| 39 | Gemma 4 31B ≈ | Google DeepMind | 1,448.17 1,437–1,459 | $0.14 | $0.4 | 262K |
| 40 | MiMo V2.5 ≈ | Xiaomi | 1,447.46 1,442–1,453 | $0.14 | $0.28 | 1,050K |
| 41 | Grok 4.1 ≈ | xAI | 1,446.43 1,442–1,451 | $2 | $10 | 200K |
| 42 | Qwen3.6 Plus ≈ | Alibaba (Qwen) | 1,445.67 1,440–1,451 | $0.5 | $3 | 1,000K |
| 43 | GPT-5.6 Luna ≈ | OpenAI | 1,445.19 1,435–1,455 | $0.2 | $1.2 | 1,050K |
| 44 | GPT-5.2 Chat ≈ | OpenAI | 1,444.41 1,439–1,450 | $1.75 | $14 | 128K |
| 45 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 1,444.26 1,440–1,449 | $0.6 | $3.6 | 262K |
| 46 | GPT-5.6 Terra ≈ | OpenAI | 1,443.56 1,433–1,454 | $2 | $12 | 1,050K |
| 47 | Gemma 4 26B-A4B ≈ | Google DeepMind | 1,442.26 1,431–1,454 | $0.05 | $0.29 | 262K |
| 48 | DeepSeek V4 Flash ≈ | DeepSeek | 1,442.23 1,437–1,448 | $0.14 | $0.28 | 1,049K |
| 49 | Longcat Flash Chat ≈ | Meituan | 1,440.51 1,432–1,449 | — | — | 131K |
| 50 | Qwen3 Max Preview ≈ | Alibaba (Qwen) | 1,440.21 1,434–1,446 | $1.2 | $6 | 256K |
| 51 | Mistral Large 3 ≈ | Mistral AI | 1,439.72 1,435–1,444 | $0.5 | $1.5 | 256K |
| 52 | DeepSeek V3.2 Exp ≈ | DeepSeek | 1,439.48 1,431–1,448 | $0.22 | $0.33 | 164K |
| 53 | Inkling ≈ | Thinking Machines Lab | 1,439.1 1,428–1,450 | $1.87 | $4.68 | 1,049K |
| 54 | MiMo V2 Omni ≈ | Xiaomi | 1,436.91 1,430–1,444 | $0.14 | $0.28 | 262K |
| 55 | Mistral Medium 3.1 ≈ | Mistral AI | 1,436.8 1,433–1,440 | $0.4 | $2 | 262K |
| 56 | DeepSeek V3.2 ≈ | DeepSeek | 1,435.55 1,431–1,440 | $0.28 | $0.4 | 164K |
| 57 | DeepSeek R1 0528 ≈ | DeepSeek | 1,434.45 1,427–1,442 | $0.25 | $0.25 | 164K |
| 58 | Mistral Medium 3.5 ≈ | Mistral AI | 1,433.77 1,425–1,442 | $1.5 | $7.5 | 262K |
| 59 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 1,433.44 1,421–1,446 | $0.3 | $2.5 | 1,049K |
| 60 | GLM 4.5 ≈ | Zhipu AI / Z.ai | 1,433.41 1,427–1,440 | $0.2 | $0.8 | 131K |
| 61 | ChatGPT 4o Latest ≈ | OpenAI | 1,432 1,428–1,436 | $5 | $15 | 128K |
| 62 | MiMo V2 Flash ≈ | Xiaomi | 1,431.13 1,426–1,436 | $0.14 | $0.28 | 262K |
| 63 | DeepSeek V3.1 ≈ | DeepSeek | 1,430.47 1,423–1,438 | $0.2 | $0.7 | 164K |
| 64 | Qwen3 VL 235B A22B Instruct ≈ | Alibaba (Qwen) | 1,430.42 1,422–1,439 | $0.4 | $1.6 | 262K |
| 65 | Claude 4.1 Opus ≈ | Anthropic | 1,428.72 1,425–1,433 | $15 | $75 | 200K |
| 66 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 1,428.63 1,423–1,434 | $0.4 | $3.2 | 262K |
| 67 | Kimi K2 Thinking Turbo ≈ | Moonshot AI (Kimi) | 1,428.12 1,424–1,432 | $1.15 | $8 | 262K |
| 68 | MiniMax M2.7 ≈ | MiniMax | 1,427.74 1,422–1,433 | $0.3 | $1.2 | 205K |
| 69 | Qwen3 235B A22B Thinking-2507 ≈ | Alibaba (Qwen) | 1,427.11 1,418–1,436 | $0.1 | $0.1 | 262K |
| 70 | Qwen3 Next 80B A3B Instruct ≈ | Alibaba (Qwen) | 1,425.02 1,419–1,431 | $0.15 | $1.2 | 262K |
| 71 | Qwen3 235B A22B Instruct 2507 ≈ | Alibaba (Qwen) | 1,424.86 1,421–1,428 | $0.1 | $0.1 | 262K |
| 72 | DeepSeek V3.1 Terminus ≈ | DeepSeek | 1,422.96 1,409–1,437 | $0.21 | $0.79 | 164K |
| 73 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 1,422.64 1,417–1,428 | $0.3 | $2.4 | 262K |
| 74 | Qwen3 Max 2025-09-23 ≈ | Alibaba (Qwen) | 1,420.37 1,412–1,429 | $0.86 | $3.43 | 258K |
| 75 | GPT 4.5 Preview ≈ | OpenAI | 1,419.49 1,412–1,427 | $75 | $150 | 128K |
| 76 | Grok 4 ≈ | xAI | 1,418.4 1,413–1,423 | $3 | $15 | 256K |
| 77 | Step 3.5 Flash ≈ | StepFun | 1,418.24 1,413–1,423 | $0.1 | $0.3 | 262K |
| 78 | GPT-5.5 Instant ≈ | OpenAI | 1,418.15 1,412–1,425 | $5 | $30 | 400K |
| 79 | Gemini 2.5 Flash ≈ | Google DeepMind | 1,418.1 1,415–1,421 | $0.3 | $2.5 | 1,049K |
| 80 | Grok 4.1 Fast Reasoning ≈ | xAI | 1,417.98 1,414–1,422 | $0.2 | $0.5 | 2,000K |
| 81 | Qwen3 VL 235B A22B Thinking ≈ | Alibaba (Qwen) | 1,417.86 1,408–1,427 | $0.4 | $4 | 262K |
| 82 | o3 2025-04-16 ≈ | OpenAI | 1,414.47 1,410–1,419 | $2 | $8 | 200K |
| 83 | Claude Haiku 4.5 ≈ | Anthropic | 1,413.97 1,410–1,418 | $1 | $5 | 200K |
| 84 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 1,412.54 1,408–1,417 | $0.25 | $1.5 | 1,049K |
| 85 | 2025) ≈ | Google DeepMind | 1,411.08 1,406–1,416 | $0.3 | $2.5 | 1,049K |
| 86 | MiMo V2 Flash (Thinking) ≈ | Xiaomi | 1,410.51 1,402–1,419 | $0.1 | $0.31 | 256K |
| 87 | Qwen3.5 35B A3B ≈ | Alibaba (Qwen) | 1,407.94 1,402–1,414 | $0.25 | $2 | 262K |
| 88 | GPT 5 Chat ≈ | OpenAI | 1,407.16 1,402–1,413 | $1.25 | $10 | 128K |
| 89 | Grok 4 Fast Reasoning ≈ | xAI | 1,407.02 1,400–1,414 | $0.2 | $0.5 | 2,000K |
| 90 | Grok 4.3 ≈ | xAI | 1,405.67 1,400–1,411 | $1.25 | $2.5 | 1,000K |
| 91 | Qwen3.5 Flash ≈ | Alibaba (Qwen) | 1,405.65 1,401–1,411 | $0.09 | $0.36 | 1,000K |
| 92 | Nemotron 3 Super ≈ | NVIDIA | 1,400.01 1,390–1,410 | $0.2 | $0.8 | 1,000K |
| 93 | Hunyuan T1 ≈ | Tencent (Hunyuan) | 1,395.04 1,382–1,408 | — | — | 131K |
| 94 | GLM 4.5 Air ≈ | Zhipu AI / Z.ai | 1,394.99 1,390–1,400 | $0.11 | $0.29 | 131K |
| 95 | Qwen3 Next 80B A3B Thinking ≈ | Alibaba (Qwen) | 1,393.03 1,385–1,401 | $0.15 | $1.2 | 262K |
| 96 | GLM 4.6V ≈ | Zhipu AI / Z.ai | 1,392.39 1,376–1,409 | $0.14 | $0.42 | 131K |
| 97 | Qwen3 30B A3B Instruct 2507 ≈ | Alibaba (Qwen) | 1,391.23 1,385–1,397 | $0.05 | $0.19 | 262K |
| 98 | GPT 4.1 2025-04-14 ≈ | OpenAI | 1,390.3 1,386–1,395 | $2 | $8 | 1,048K |
| 99 | GPT-5.3 Chat (latest) ≈ | OpenAI | 1,389.83 1,384–1,396 | $1.75 | $14 | 128K |
| 100 | DeepSeek Chat 0324 ≈ | DeepSeek | 1,386.46 1,382–1,391 | $0.2 | $0.6 | 164K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 99 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 11.6 points — that is about 10 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 13,842: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)