Prompts about layout and web applications, rated blindly by people. It differs from the programming collection in what is measured: there, solved tasks are counted; here, whose result people preferred. For the web that is closer to the point, because how the result looks is part of what gets judged.
| # | Model | Developer | Arena Score, web | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 1,704.67 1,689–1,720 | $5 | $25 | 1,000K |
| 2 | Kimi K3 | Moonshot AI (Kimi) | 1,675.56 1,664–1,687 | $2 | $8 | 1,049K |
| 3 | Claude Fable 5 | Anthropic | 1,630.02 1,621–1,639 | $10 | $50 | 1,000K |
| 4 | GPT-5.6 Sol ≈ | OpenAI | 1,620.27 1,611–1,630 | $5 | $30 | 1,050K |
| 5 | GLM 5.2 | Zhipu AI / Z.ai | 1,586.29 1,577–1,595 | $0.42 | $1.32 | 1,049K |
| 6 | DeepSeek V4 Flash ≈ | DeepSeek | 1,576.54 1,559–1,594 | $0.14 | $0.28 | 1,049K |
| 7 | Claude Opus 4.7 ≈ | Anthropic | 1,560.63 1,554–1,567 | $5 | $25 | 1,000K |
| 8 | Grok 4.5 ≈ | xAI | 1,549.24 1,539–1,560 | $2 | $6 | 500K |
| 9 | Claude Sonnet 5 ≈ | Anthropic | 1,541.89 1,532–1,552 | $2 | $10 | 1,000K |
| 10 | Claude Opus 4.8 ≈ | Anthropic | 1,538.84 1,531–1,547 | $5 | $25 | 1,000K |
| 11 | Claude Opus 4.6 ≈ | Anthropic | 1,537.93 1,532–1,544 | $5 | $25 | 1,000K |
| 12 | Muse Spark 1.1 ≈ | Meta AI | 1,535.95 1,525–1,546 | $1.25 | $4.25 | 1,049K |
| 13 | Gemini 3.6 Flash ≈ | Google DeepMind | 1,532.54 1,521–1,544 | $1.5 | $7.5 | 1,049K |
| 14 | GPT-5.6 Luna ≈ | OpenAI | 1,522.94 1,512–1,534 | $0.2 | $1.2 | 1,050K |
| 15 | Claude Sonnet 4.6 ≈ | Anthropic | 1,522.64 1,517–1,528 | $3 | $15 | 1,000K |
| 16 | GPT-5.6 Terra ≈ | OpenAI | 1,521.61 1,510–1,533 | $2 | $12 | 1,050K |
| 17 | Qwen3.7 Max ≈ | Alibaba (Qwen) | 1,516.97 1,509–1,525 | $2.5 | $7.5 | 1,000K |
| 18 | Hy3 ≈ | Tencent (Hunyuan) | 1,516.53 1,500–1,533 | $0.13 | $0.53 | 262K |
| 19 | GLM 5.1 ≈ | Zhipu AI / Z.ai | 1,516.16 1,508–1,524 | $0.06 | $0.22 | 205K |
| 20 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 1,509.51 1,502–1,517 | $0.22 | $1.14 | 262K |
| 21 | MiniMax M3 | MiniMax | 1,490.74 1,483–1,499 | $0.3 | $1.2 | 1,049K |
| 22 | Gemini 3.5 Flash ≈ | Google DeepMind | 1,486.18 1,478–1,494 | $1.5 | $9 | 1,049K |
| 23 | Qwen3.6 Max Preview ≈ | Alibaba (Qwen) | 1,478.18 1,465–1,491 | $1.3 | $7.8 | 262K |
| 24 | MiMo V2.5 Pro ≈ | Xiaomi | 1,474.02 1,467–1,481 | $0.44 | $0.87 | 1,050K |
| 25 | Kimi K2.7 Code ≈ | Moonshot AI (Kimi) | 1,472.57 1,463–1,482 | $0.28 | $1.1 | 262K |
| 26 | Claude 4.5 Opus ≈ | Anthropic | 1,466.99 1,460–1,474 | $5 | $25 | 200K |
| 27 | Qwen3.6 Plus ≈ | Alibaba (Qwen) | 1,457.86 1,452–1,464 | $0.5 | $3 | 1,000K |
| 28 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 1,451.88 1,408–1,496 | $0.3 | $2.5 | 1,049K |
| 29 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 1,446.83 1,441–1,452 | $2 | $12 | 1,049K |
| 30 | DeepSeek V4 Pro ≈ | DeepSeek | 1,446 1,439–1,453 | $0.44 | $0.87 | 1,049K |
| 31 | Gemini 3 Pro ≈ | Google DeepMind | 1,438.15 1,430–1,447 | $1.6 | $9.6 | 1,049K |
| 32 | Gemini 3 Flash ≈ | Google DeepMind | 1,438.01 1,429–1,447 | $0.4 | $2.4 | 1,049K |
| 33 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 1,435.77 1,430–1,441 | $0.3 | $1.9 | 256K |
| 34 | MiMo V2.5 ≈ | Xiaomi | 1,435.5 1,428–1,443 | $0.14 | $0.28 | 1,050K |
| 35 | GLM 5 ≈ | Zhipu AI / Z.ai | 1,434.74 1,426–1,443 | $0.48 | $1.9 | 205K |
| 36 | GLM 4.7 ≈ | Zhipu AI / Z.ai | 1,433.46 1,421–1,446 | $0.15 | $0.8 | 205K |
| 37 | MiMo V2 Pro ≈ | Xiaomi | 1,433 1,425–1,441 | $0.44 | $0.87 | 1,049K |
| 38 | Inkling | Thinking Machines Lab | 1,411.46 1,402–1,421 | $1.87 | $4.68 | 1,049K |
| 39 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 1,399.59 1,394–1,405 | $0.6 | $3.6 | 262K |
| 40 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 1,399.35 1,385–1,414 | $0.7 | $3.1 | 203K |
| 41 | MiniMax M2.7 ≈ | MiniMax | 1,398.49 1,392–1,405 | $0.3 | $1.2 | 205K |
| 42 | Claude 4.1 Opus ≈ | Anthropic | 1,388.74 1,378–1,400 | $15 | $75 | 200K |
| 43 | MiniMax M2.5 ≈ | MiniMax | 1,386.27 1,378–1,395 | $0.3 | $1.2 | 1,000K |
| 44 | Claude Sonnet 4.5 ≈ | Anthropic | 1,385.13 1,378–1,392 | $3 | $15 | 1,000K |
| 45 | Grok 4.20 (Reasoning) ≈ | xAI | 1,373.68 1,367–1,380 | $2 | $6 | 2,000K |
| 46 | Gemma 4 26B-A4B ≈ | Google DeepMind | 1,366.59 1,349–1,384 | $0.05 | $0.29 | 262K |
| 47 | Gemma 4 31B ≈ | Google DeepMind | 1,363.39 1,355–1,371 | $0.14 | $0.4 | 262K |
| 48 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 1,359.5 1,352–1,367 | $0.4 | $3.2 | 262K |
| 49 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 1,357.62 1,349–1,366 | $0.3 | $2.4 | 262K |
| 50 | Grok 4.3 ≈ | xAI | 1,356.93 1,350–1,364 | $1.25 | $2.5 | 1,000K |
| 51 | Laguna M.1 ≈ | Poolside | 1,348.81 1,339–1,359 | $0.2 | $0.4 | 262K |
| 52 | GLM 4.6 ≈ | Zhipu AI / Z.ai | 1,339.03 1,328–1,350 | $0.29 | $1.14 | 205K |
| 53 | MiMo V2 Flash ≈ | Xiaomi | 1,330.69 1,321–1,341 | $0.14 | $0.28 | 262K |
| 54 | Claude Haiku 4.5 ≈ | Anthropic | 1,325.04 1,320–1,330 | $1 | $5 | 200K |
| 55 | DeepSeek V3.2 ≈ | DeepSeek | 1,323.38 1,315–1,331 | $0.28 | $0.4 | 164K |
| 56 | Kimi K2 Thinking Turbo ≈ | Moonshot AI (Kimi) | 1,322 1,315–1,329 | $1.15 | $8 | 262K |
| 57 | Laguna XS.2 ≈ | Poolside | 1,303.56 1,292–1,315 | $0.2 | $0.4 | 262K |
| 58 | MiniMax M2 ≈ | MiniMax | 1,296.69 1,286–1,308 | $0.3 | $1.2 | 205K |
| 59 | MiMo V2 Flash (Thinking) ≈ | Xiaomi | 1,291.23 1,275–1,308 | $0.1 | $0.31 | 256K |
| 60 | Qwen3 Coder 480B A35B ≈ | Alibaba (Qwen) | 1,272.22 1,264–1,280 | $1.5 | $7.5 | 262K |
| 61 | DeepSeek V3.2 Exp ≈ | DeepSeek | 1,271.78 1,258–1,285 | $0.22 | $0.33 | 164K |
| 62 | Mistral Medium 3.5 ≈ | Mistral AI | 1,267.07 1,251–1,283 | $1.5 | $7.5 | 262K |
| 63 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 1,256.23 1,249–1,263 | $0.25 | $1.5 | 1,049K |
| 64 | KAT-Coder-Pro V1 ≈ | Kuaishou (Kling, KAT) | 1,255.01 1,235–1,275 | $0.3 | $1.2 | 256K |
| 65 | Qwen3.5 35B A3B ≈ | Alibaba (Qwen) | 1,250.46 1,233–1,268 | $0.25 | $2 | 262K |
| 66 | Grok 4.1 Fast Reasoning ≈ | xAI | 1,240.06 1,229–1,251 | $0.2 | $0.5 | 2,000K |
| 67 | Trinity Large Thinking ≈ | Arcee AI | 1,238.93 1,218–1,260 | $0.22 | $0.85 | 262K |
| 68 | Qwen3.5 Flash ≈ | Alibaba (Qwen) | 1,237.96 1,218–1,258 | $0.09 | $0.36 | 1,000K |
| 69 | Mistral Large 3 ≈ | Mistral AI | 1,230.02 1,204–1,256 | $0.5 | $1.5 | 256K |
| 70 | Gemini 2.5 Pro ≈ | Google DeepMind | 1,224.23 1,208–1,240 | $1.25 | $10 | 1,049K |
| 71 | Grok 4.1 ≈ | xAI | 1,210.33 1,185–1,235 | $2 | $10 | 200K |
| 72 | Granite 4.1 8B ≈ | IBM | 1,193.98 1,175–1,213 | $0.05 | $0.1 | 131K |
| 73 | Devstral 2 ≈ | Mistral AI | 1,193.89 1,173–1,215 | $0.4 | $2 | 256K |
| 74 | Mercury 2 ≈ | Inception Labs | 1,166.34 1,141–1,192 | $0.25 | $0.75 | 128K |
| 75 | Grok Code Fast 1 ≈ | xAI | 1,163.41 1,135–1,192 | $0.2 | $1.5 | 256K |
| 76 | Grok 4 Fast Reasoning ≈ | xAI | 1,160.41 1,133–1,188 | $0.2 | $0.5 | 2,000K |
| 77 | Devstral Medium | Mistral AI | 1,079.26 1,048–1,111 | $0.4 | $2 | 128K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 70 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 19.2 points — that is about 2 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 6,310: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)