Fiction and free-form writing, rated blindly by people: work of this kind has no correct answer, so preference is the most honest measure available. There is also a second view, from a scored run of set tasks; it measures something else and yields a different order.
One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.
Not every model is measured here. This table is missing 4 of the top ten from the “By solved tasks” tab — GPT 5 2025-08-07, Kimi K2 Instruct, OpenAI o3-pro (2025-06-10), GPT 5 Mini 2025-08-07. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 27 July 2026.
| # | Model | Developer | Arena Score, writing | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 1,510.88 1,481–1,541 | $5 | $25 | 1,000K |
| 2 | Claude Fable 5 ≈ | Anthropic | 1,498.96 1,487–1,511 | $10 | $50 | 1,000K |
| 3 | Gemini 3 Pro ≈ | Google DeepMind | 1,481.7 1,473–1,490 | $1.6 | $9.6 | 1,049K |
| 4 | Claude Opus 4.6 ≈ | Anthropic | 1,481.69 1,475–1,489 | $5 | $25 | 1,000K |
| 5 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 1,481.61 1,475–1,488 | $2 | $12 | 1,049K |
| 6 | Claude Opus 4.7 ≈ | Anthropic | 1,476.58 1,469–1,484 | $5 | $25 | 1,000K |
| 7 | Gemini 3.5 Flash ≈ | Google DeepMind | 1,470.09 1,454–1,486 | $1.5 | $9 | 1,049K |
| 8 | Kimi K3 ≈ | Moonshot AI (Kimi) | 1,464.37 1,440–1,489 | $2 | $8 | 1,049K |
| 9 | Gemini 3.6 Flash ≈ | Google DeepMind | 1,463.55 1,445–1,482 | $1.5 | $7.5 | 1,049K |
| 10 | Gemini 3 Flash ≈ | Google DeepMind | 1,458.54 1,449–1,468 | $0.4 | $2.4 | 1,049K |
| 11 | GLM 5.2 ≈ | Zhipu AI / Z.ai | 1,456.24 1,446–1,467 | $0.42 | $1.32 | 1,049K |
| 12 | Gemini 2.5 Pro ≈ | Google DeepMind | 1,454.49 1,449–1,460 | $1.25 | $10 | 1,049K |
| 13 | GLM 5.1 ≈ | Zhipu AI / Z.ai | 1,454.14 1,445–1,463 | $0.06 | $0.22 | 205K |
| 14 | GPT-5.6 Sol ≈ | OpenAI | 1,453.45 1,437–1,470 | $5 | $30 | 1,050K |
| 15 | DeepSeek V4 Pro ≈ | DeepSeek | 1,443.14 1,435–1,451 | $0.44 | $0.87 | 1,049K |
| 16 | Claude Opus 4.8 ≈ | Anthropic | 1,443.1 1,434–1,452 | $5 | $25 | 1,000K |
| 17 | Claude Sonnet 4.5 ≈ | Anthropic | 1,440.99 1,435–1,447 | $3 | $15 | 1,000K |
| 18 | Claude 4.5 Opus ≈ | Anthropic | 1,440.86 1,434–1,447 | $5 | $25 | 200K |
| 19 | Muse Spark 1.1 ≈ | Meta AI | 1,440.45 1,426–1,455 | $1.25 | $4.25 | 1,049K |
| 20 | Qwen3.7 Plus ≈ | Alibaba (Qwen) | 1,440.34 1,430–1,450 | $0.5 | $3 | 1,000K |
| 21 | ERNIE 5.1 ≈ | Baidu (Ernie) | 1,439.68 1,431–1,448 | $0.75 | $3 | 119K |
| 22 | Grok 4.5 ≈ | xAI | 1,439.21 1,424–1,454 | $2 | $6 | 500K |
| 23 | Claude Sonnet 4.6 ≈ | Anthropic | 1,438.07 1,431–1,445 | $3 | $15 | 1,000K |
| 24 | Grok 4.20 Multi Agent Beta 0309 ≈ | xAI | 1,436.86 1,430–1,444 | $2 | $6 | 2,000K |
| 25 | GLM 5 ≈ | Zhipu AI / Z.ai | 1,436.16 1,427–1,445 | $0.48 | $1.9 | 205K |
| 26 | Qwen3.6 Max Preview ≈ | Alibaba (Qwen) | 1,433.85 1,411–1,457 | $1.3 | $7.8 | 262K |
| 27 | MiMo V2.5 Pro ≈ | Xiaomi | 1,433.77 1,426–1,442 | $0.44 | $0.87 | 1,050K |
| 28 | Grok 4.20 (Reasoning) ≈ | xAI | 1,433.4 1,426–1,440 | $2 | $6 | 2,000K |
| 29 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 1,430.8 1,422–1,439 | $0.22 | $1.14 | 262K |
| 30 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 1,423.84 1,417–1,431 | $0.3 | $1.9 | 256K |
| 31 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 1,420.66 1,403–1,439 | $0.3 | $2.5 | 1,049K |
| 32 | GPT-5.5 Instant ≈ | OpenAI | 1,420.47 1,411–1,430 | $5 | $30 | 400K |
| 33 | Gemma 4 31B ≈ | Google DeepMind | 1,417.17 1,398–1,436 | $0.14 | $0.4 | 262K |
| 34 | MiMo V2 Pro ≈ | Xiaomi | 1,416.27 1,406–1,427 | $0.44 | $0.87 | 1,049K |
| 35 | GLM 4.6 ≈ | Zhipu AI / Z.ai | 1,412.75 1,404–1,421 | $0.29 | $1.14 | 205K |
| 36 | Grok 4.1 ≈ | xAI | 1,412.53 1,406–1,419 | $2 | $10 | 200K |
| 37 | GPT-5.6 Terra ≈ | OpenAI | 1,410.72 1,395–1,426 | $2 | $12 | 1,050K |
| 38 | Hy3 ≈ | Tencent (Hunyuan) | 1,410.4 1,386–1,435 | $0.13 | $0.53 | 262K |
| 39 | Claude 4.1 Opus ≈ | Anthropic | 1,410.22 1,404–1,416 | $15 | $75 | 200K |
| 40 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 1,409.81 1,394–1,425 | $0.7 | $3.1 | 203K |
| 41 | DeepSeek R1 0528 ≈ | DeepSeek | 1,409.54 1,397–1,422 | $0.25 | $0.25 | 164K |
| 42 | ChatGPT 4o Latest ≈ | OpenAI | 1,407.13 1,401–1,413 | $5 | $15 | 128K |
| 43 | Claude Sonnet 5 ≈ | Anthropic | 1,406.98 1,395–1,419 | $2 | $10 | 1,000K |
| 44 | Seed 2.0 Pro ≈ | ByteDance (Doubao) | 1,406.26 1,400–1,413 | $0.5 | $3 | 131K |
| 45 | DeepSeek V3.1 Terminus ≈ | DeepSeek | 1,405.52 1,378–1,433 | $0.21 | $0.79 | 164K |
| 46 | Gemma 4 26B-A4B ≈ | Google DeepMind | 1,404.92 1,386–1,424 | $0.05 | $0.29 | 262K |
| 47 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 1,404.33 1,397–1,411 | $0.6 | $3.6 | 262K |
| 48 | DeepSeek V3.2 Exp ≈ | DeepSeek | 1,403.82 1,389–1,418 | $0.22 | $0.33 | 164K |
| 49 | MiniMax M3 ≈ | MiniMax | 1,403.76 1,394–1,413 | $0.3 | $1.2 | 1,049K |
| 50 | GPT-5.6 Luna ≈ | OpenAI | 1,403.67 1,388–1,419 | $0.2 | $1.2 | 1,050K |
| 51 | Qwen3.6 Plus ≈ | Alibaba (Qwen) | 1,403.57 1,395–1,412 | $0.5 | $3 | 1,000K |
| 52 | Gemini 2.5 Flash ≈ | Google DeepMind | 1,402.43 1,397–1,407 | $0.3 | $2.5 | 1,049K |
| 53 | GLM 4.7 ≈ | Zhipu AI / Z.ai | 1,402.36 1,389–1,416 | $0.15 | $0.8 | 205K |
| 54 | DeepSeek V4 Flash ≈ | DeepSeek | 1,402.07 1,394–1,410 | $0.14 | $0.28 | 1,049K |
| 55 | GPT-5.2 Chat ≈ | OpenAI | 1,401.99 1,393–1,411 | $1.75 | $14 | 128K |
| 56 | Qwen3 Max Preview ≈ | Alibaba (Qwen) | 1,400.79 1,391–1,411 | $1.2 | $6 | 256K |
| 57 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 1,400.66 1,394–1,408 | $0.25 | $1.5 | 1,049K |
| 58 | DeepSeek V3.2 ≈ | DeepSeek | 1,399.13 1,391–1,407 | $0.28 | $0.4 | 164K |
| 59 | Grok 4 ≈ | xAI | 1,398.43 1,390–1,407 | $3 | $15 | 256K |
| 60 | Inkling ≈ | Thinking Machines Lab | 1,395.2 1,379–1,412 | $1.87 | $4.68 | 1,049K |
| 61 | GLM 4.5 ≈ | Zhipu AI / Z.ai | 1,395.04 1,384–1,406 | $0.2 | $0.8 | 131K |
| 62 | GPT 4.5 Preview ≈ | OpenAI | 1,394.29 1,382–1,406 | $75 | $150 | 128K |
| 63 | Grok 4.1 Fast Reasoning ≈ | xAI | 1,393.1 1,386–1,400 | $0.2 | $0.5 | 2,000K |
| 64 | Hunyuan T1 ≈ | Tencent (Hunyuan) | 1,392.73 1,369–1,416 | — | — | 131K |
| 65 | Mistral Medium 3.1 ≈ | Mistral AI | 1,392.2 1,387–1,398 | $0.4 | $2 | 262K |
| 66 | MiMo V2.5 ≈ | Xiaomi | 1,391.79 1,384–1,400 | $0.14 | $0.28 | 1,050K |
| 67 | MiMo V2 Omni ≈ | Xiaomi | 1,391.39 1,380–1,403 | $0.14 | $0.28 | 262K |
| 68 | Mistral Large 3 ≈ | Mistral AI | 1,390.53 1,383–1,398 | $0.5 | $1.5 | 256K |
| 69 | Grok 4.3 ≈ | xAI | 1,388.05 1,380–1,396 | $1.25 | $2.5 | 1,000K |
| 70 | 2025) ≈ | Google DeepMind | 1,387.99 1,379–1,397 | $0.3 | $2.5 | 1,049K |
| 71 | DeepSeek V3.1 ≈ | DeepSeek | 1,387.55 1,374–1,401 | $0.2 | $0.7 | 164K |
| 72 | Qwen3 235B A22B Thinking-2507 ≈ | Alibaba (Qwen) | 1,386.32 1,368–1,404 | $0.1 | $0.1 | 262K |
| 73 | Qwen3 Max 2025-09-23 ≈ | Alibaba (Qwen) | 1,382.56 1,366–1,400 | $0.86 | $3.43 | 258K |
| 74 | MiMo V2 Flash ≈ | Xiaomi | 1,378.57 1,371–1,386 | $0.14 | $0.28 | 262K |
| 75 | Qwen3 235B A22B Instruct 2507 ≈ | Alibaba (Qwen) | 1,375.98 1,370–1,381 | $0.1 | $0.1 | 262K |
| 76 | Kimi K2 Thinking Turbo ≈ | Moonshot AI (Kimi) | 1,374.88 1,368–1,382 | $1.15 | $8 | 262K |
| 77 | Grok 4 Fast Reasoning ≈ | xAI | 1,374.71 1,363–1,386 | $0.2 | $0.5 | 2,000K |
| 78 | Claude 4 Opus ≈ | Anthropic | 1,374.61 1,366–1,383 | $15 | $75 | 200K |
| 79 | Claude Haiku 4.5 ≈ | Anthropic | 1,371.16 1,366–1,377 | $1 | $5 | 200K |
| 80 | Qwen3 VL 235B A22B Instruct ≈ | Alibaba (Qwen) | 1,370.17 1,353–1,387 | $0.4 | $1.6 | 262K |
| 81 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 1,369.14 1,360–1,379 | $0.4 | $3.2 | 262K |
| 82 | Mistral Medium 3.5 ≈ | Mistral AI | 1,368.5 1,354–1,383 | $1.5 | $7.5 | 262K |
| 83 | GPT 5 Chat ≈ | OpenAI | 1,367.89 1,358–1,377 | $1.25 | $10 | 128K |
| 84 | Gemini 2.5 Flash Lite Preview ≈ | Google DeepMind | 1,366.02 1,357–1,375 | $0.1 | $0.4 | 1,049K |
| 85 | DeepSeek Chat 0324 ≈ | DeepSeek | 1,364.71 1,357–1,373 | $0.2 | $0.6 | 164K |
| 86 | GPT 4.1 2025-04-14 ≈ | OpenAI | 1,363.66 1,356–1,371 | $2 | $8 | 1,048K |
| 87 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 1,361.69 1,352–1,371 | $0.3 | $2.4 | 262K |
| 88 | o3 2025-04-16 ≈ | OpenAI | 1,359.4 1,352–1,367 | $2 | $8 | 200K |
| 89 | Hunyuan TurboS ≈ | Tencent (Hunyuan) | 1,358.18 1,343–1,374 | — | — | 131K |
| 90 | Step 3.5 Flash ≈ | StepFun | 1,355.92 1,349–1,363 | $0.1 | $0.3 | 262K |
| 91 | GPT-5.3 Chat (latest) ≈ | OpenAI | 1,355.23 1,346–1,364 | $1.75 | $14 | 128K |
| 92 | DeepSeek Reasoner ≈ | DeepSeek | 1,354.71 1,344–1,365 | $0.55 | $2.19 | 164K |
| 93 | MiniMax M2.7 ≈ | MiniMax | 1,353.29 1,345–1,361 | $0.3 | $1.2 | 205K |
| 94 | Kimi K2 0905 Preview ≈ | Moonshot AI (Kimi) | 1,349.52 1,334–1,365 | $0.6 | $2.5 | 262K |
| 95 | MiMo V2 Flash (Thinking) ≈ | Xiaomi | 1,348.09 1,334–1,362 | $0.1 | $0.31 | 256K |
| 96 | o1 2024-12-17 ≈ | OpenAI | 1,347.89 1,339–1,357 | $15 | $60 | 200K |
| 97 | Qwen3.5 35B A3B ≈ | Alibaba (Qwen) | 1,347.69 1,338–1,357 | $0.25 | $2 | 262K |
| 98 | Longcat Flash Chat ≈ | Meituan | 1,346.08 1,330–1,362 | — | — | 131K |
| 99 | Gemma 3 27B ≈ | Google DeepMind | 1,344.5 1,337–1,352 | $0.03 | $0.11 | 131K |
| 100 | Qwen3 VL 235B A22B Thinking ≈ | Alibaba (Qwen) | 1,344.2 1,326–1,362 | $0.4 | $4 | 262K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 99 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 18.5 points — that is about 11 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 4,577: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)
Not every model is measured here. This table is missing 10 of the top ten from the “By human votes” tab — Claude Opus 5, Claude Fable 5, Gemini 3 Pro, Claude Opus 4.6, Gemini 3.1 Pro Preview and 5 more. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 1 December 2025.
| # | Model | Developer | Score on tasks % | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | GPT 5 2025-08-07 | OpenAI | 86 % | $1.25 | $10 | 272K |
| 2 | Kimi K2 Instruct | Moonshot AI (Kimi) | 85.6 % | $0.1 | $2 | 131K |
| 3 | Claude 4.1 Opus | Anthropic | 84.7 % | $15 | $75 | 200K |
| 4 | OpenAI o3-pro (2025-06-10) | OpenAI | 84.4 % | $20 | $80 | 200K |
| 5 | o3 2025-04-16 | OpenAI | 83.9 % | $2 | $8 | 200K |
| 6 | Gemini 2.5 Pro | Google DeepMind | 83.8 % | $1.25 | $10 | 1,049K |
| 7 | Claude 4 Opus | Anthropic | 83.6 % | $15 | $75 | 200K |
| 8 | GPT 5 Mini 2025-08-07 | OpenAI | 83.1 % | $0.25 | $2 | 272K |
| 9 | DeepSeek Reasoner | DeepSeek | 83 % | $0.55 | $2.19 | 164K |
| 10 | Qwen 3 235b A22B | Alibaba (Qwen) | 83 % | $0.7 | $2.8 | 131K |
| 11 | Qwen3 235B A22B Thinking-2507 | Alibaba (Qwen) | 82.4 % | $0.1 | $0.1 | 262K |
| 12 | DeepSeek R1 0528 | DeepSeek | 81.9 % | $0.25 | $0.25 | 164K |
| 13 | GPT-4o (2024-11-20) | OpenAI | 81.8 % | $2.5 | $10 | 128K |
| 14 | Claude 4 Sonnet | Anthropic | 81.4 % | $3 | $15 | 1,000K |
| 15 | Claude Sonnet 3.7 | Anthropic | 81.1 % | $3 | $15 | 200K |
| 16 | Gemini 2.5 Pro Preview 0506 | Google DeepMind | 80.9 % | $1.25 | $10 | 1,049K |
| 17 | Gemini 2.5 Pro Experimental 0325 | Google DeepMind | 80.5 % | $2.5 | $10 | 1,049K |
| 18 | Claude 3.5 Sonnet 2024-10-22 | Anthropic | 80.3 % | $3 | $15 | 200K |
| 19 | Gemma 3 27B | Google DeepMind | 79.9 % | $0.03 | $0.11 | 131K |
| 20 | Mistral Medium 3 | Mistral AI | 77.3 % | $0.4 | $2 | 131K |
| 21 | GPT OSS 120B | OpenAI | 77.1 % | $0.03 | $0.14 | 131K |
| 22 | DeepSeek Chat 0324 | DeepSeek | 77 % | $0.2 | $0.6 | 164K |
| 23 | Grok 4 | xAI | 76.9 % | $3 | $15 | 256K |
| 24 | Gemini 2.5 Flash Preview | Google DeepMind | 76.5 % | $0.15 | $0.6 | 1,049K |
| 25 | Grok 3 | xAI | 76.4 % | $3 | $15 | 131K |
| 26 | GPT 4.5 Preview | OpenAI | 75.6 % | $75 | $150 | 128K |
| 27 | o4 Mini 2025-04-16 | OpenAI | 75 % | $1.1 | $4.4 | 200K |
| 28 | Gemini 2.0 Flash Thinking 0121 | Google DeepMind | 73.8 % | $0.31 | $1 | 1,000K |
| 29 | Grok 3 Mini | xAI | 73.5 % | $0.3 | $0.5 | 131K |
| 30 | Claude 3.5 Haiku | Anthropic | 73.5 % | $0.25 | $1.25 | 200K |
| 31 | o1 2024-12-17 | OpenAI | 70.2 % | $15 | $60 | 200K |
| 32 | Mistral Large 2407 | Mistral AI | 69 % | $3 | $9 | 131K |
| 33 | OpenAI: GPT-4o-mini (2024-07-18) | OpenAI | 67.2 % | $0.15 | $0.6 | 128K |
| 34 | o1 Mini 2024-09-12 | OpenAI | 64.9 % | $1.1 | $4.4 | 128K |
| 35 | Grok 2 1212 | xAI | 63.6 % | $2 | $10 | 131K |
| 36 | Phi 4 | Microsoft | 62.6 % | $0.06 | $0.14 | 128K |
| 37 | Llama 4 Maverick 17b 128e Instruct | Meta AI | 62 % | $0.05 | $0.1 | 1,049K |
| 38 | o3 Mini 2025-01-31 | OpenAI | 61.7 % | $1.1 | $4.4 | 200K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
Epoch AI
— license CC-BY · primary source
По данным Epoch AI