AI ranking

Best AI models for web development

Prompts about layout and web applications, rated blindly by people. It differs from the programming collection in what is measured: there, solved tasks are counted; here, whose result people preferred. For the web that is closer to the point, because how the result looks is part of what gets judged.

77 models in the table
77 considered in the calculation
1 August 2026 measurement date at the source
27 July 2026 newest model released here
7 August 2026 last recalculated
# Model Developer Arena Score, web Sign in$ / 1M Output$ / 1M Contexttokens
1 Claude Opus 5 Anthropic 1,704.67 1,689–1,720 $5 $25 1,000K
2 Kimi K3 Moonshot AI (Kimi) 1,675.56 1,664–1,687 $2 $8 1,049K
3 Claude Fable 5 Anthropic 1,630.02 1,621–1,639 $10 $50 1,000K
4 GPT-5.6 Sol OpenAI 1,620.27 1,611–1,630 $5 $30 1,050K
5 GLM 5.2 Zhipu AI / Z.ai 1,586.29 1,577–1,595 $0.42 $1.32 1,049K
6 DeepSeek V4 Flash DeepSeek 1,576.54 1,559–1,594 $0.14 $0.28 1,049K
7 Claude Opus 4.7 Anthropic 1,560.63 1,554–1,567 $5 $25 1,000K
8 Grok 4.5 xAI 1,549.24 1,539–1,560 $2 $6 500K
9 Claude Sonnet 5 Anthropic 1,541.89 1,532–1,552 $2 $10 1,000K
10 Claude Opus 4.8 Anthropic 1,538.84 1,531–1,547 $5 $25 1,000K
11 Claude Opus 4.6 Anthropic 1,537.93 1,532–1,544 $5 $25 1,000K
12 Muse Spark 1.1 Meta AI 1,535.95 1,525–1,546 $1.25 $4.25 1,049K
13 Gemini 3.6 Flash Google DeepMind 1,532.54 1,521–1,544 $1.5 $7.5 1,049K
14 GPT-5.6 Luna OpenAI 1,522.94 1,512–1,534 $0.2 $1.2 1,050K
15 Claude Sonnet 4.6 Anthropic 1,522.64 1,517–1,528 $3 $15 1,000K
16 GPT-5.6 Terra OpenAI 1,521.61 1,510–1,533 $2 $12 1,050K
17 Qwen3.7 Max Alibaba (Qwen) 1,516.97 1,509–1,525 $2.5 $7.5 1,000K
18 Hy3 Tencent (Hunyuan) 1,516.53 1,500–1,533 $0.13 $0.53 262K
19 GLM 5.1 Zhipu AI / Z.ai 1,516.16 1,508–1,524 $0.06 $0.22 205K
20 Kimi K2.6 Moonshot AI (Kimi) 1,509.51 1,502–1,517 $0.22 $1.14 262K
21 MiniMax M3 MiniMax 1,490.74 1,483–1,499 $0.3 $1.2 1,049K
22 Gemini 3.5 Flash Google DeepMind 1,486.18 1,478–1,494 $1.5 $9 1,049K
23 Qwen3.6 Max Preview Alibaba (Qwen) 1,478.18 1,465–1,491 $1.3 $7.8 262K
24 MiMo V2.5 Pro Xiaomi 1,474.02 1,467–1,481 $0.44 $0.87 1,050K
25 Kimi K2.7 Code Moonshot AI (Kimi) 1,472.57 1,463–1,482 $0.28 $1.1 262K
26 Claude 4.5 Opus Anthropic 1,466.99 1,460–1,474 $5 $25 200K
27 Qwen3.6 Plus Alibaba (Qwen) 1,457.86 1,452–1,464 $0.5 $3 1,000K
28 Gemini 3.5 Flash Lite Google DeepMind 1,451.88 1,408–1,496 $0.3 $2.5 1,049K
29 Gemini 3.1 Pro Preview Google DeepMind 1,446.83 1,441–1,452 $2 $12 1,049K
30 DeepSeek V4 Pro DeepSeek 1,446 1,439–1,453 $0.44 $0.87 1,049K
31 Gemini 3 Pro Google DeepMind 1,438.15 1,430–1,447 $1.6 $9.6 1,049K
32 Gemini 3 Flash Google DeepMind 1,438.01 1,429–1,447 $0.4 $2.4 1,049K
33 Kimi K2.5 Thinking TEE Moonshot AI (Kimi) 1,435.77 1,430–1,441 $0.3 $1.9 256K
34 MiMo V2.5 Xiaomi 1,435.5 1,428–1,443 $0.14 $0.28 1,050K
35 GLM 5 Zhipu AI / Z.ai 1,434.74 1,426–1,443 $0.48 $1.9 205K
36 GLM 4.7 Zhipu AI / Z.ai 1,433.46 1,421–1,446 $0.15 $0.8 205K
37 MiMo V2 Pro Xiaomi 1,433 1,425–1,441 $0.44 $0.87 1,049K
38 Inkling Thinking Machines Lab 1,411.46 1,402–1,421 $1.87 $4.68 1,049K
39 Qwen3.5 397B-A17B Alibaba (Qwen) 1,399.59 1,394–1,405 $0.6 $3.6 262K
40 GLM 5V Turbo Zhipu AI / Z.ai 1,399.35 1,385–1,414 $0.7 $3.1 203K
41 MiniMax M2.7 MiniMax 1,398.49 1,392–1,405 $0.3 $1.2 205K
42 Claude 4.1 Opus Anthropic 1,388.74 1,378–1,400 $15 $75 200K
43 MiniMax M2.5 MiniMax 1,386.27 1,378–1,395 $0.3 $1.2 1,000K
44 Claude Sonnet 4.5 Anthropic 1,385.13 1,378–1,392 $3 $15 1,000K
45 Grok 4.20 (Reasoning) xAI 1,373.68 1,367–1,380 $2 $6 2,000K
46 Gemma 4 26B-A4B Google DeepMind 1,366.59 1,349–1,384 $0.05 $0.29 262K
47 Gemma 4 31B Google DeepMind 1,363.39 1,355–1,371 $0.14 $0.4 262K
48 Qwen3.5 122B-A10B Alibaba (Qwen) 1,359.5 1,352–1,367 $0.4 $3.2 262K
49 Qwen3.5 27B Alibaba (Qwen) 1,357.62 1,349–1,366 $0.3 $2.4 262K
50 Grok 4.3 xAI 1,356.93 1,350–1,364 $1.25 $2.5 1,000K
51 Laguna M.1 Poolside 1,348.81 1,339–1,359 $0.2 $0.4 262K
52 GLM 4.6 Zhipu AI / Z.ai 1,339.03 1,328–1,350 $0.29 $1.14 205K
53 MiMo V2 Flash Xiaomi 1,330.69 1,321–1,341 $0.14 $0.28 262K
54 Claude Haiku 4.5 Anthropic 1,325.04 1,320–1,330 $1 $5 200K
55 DeepSeek V3.2 DeepSeek 1,323.38 1,315–1,331 $0.28 $0.4 164K
56 Kimi K2 Thinking Turbo Moonshot AI (Kimi) 1,322 1,315–1,329 $1.15 $8 262K
57 Laguna XS.2 Poolside 1,303.56 1,292–1,315 $0.2 $0.4 262K
58 MiniMax M2 MiniMax 1,296.69 1,286–1,308 $0.3 $1.2 205K
59 MiMo V2 Flash (Thinking) Xiaomi 1,291.23 1,275–1,308 $0.1 $0.31 256K
60 Qwen3 Coder 480B A35B Alibaba (Qwen) 1,272.22 1,264–1,280 $1.5 $7.5 262K
61 DeepSeek V3.2 Exp DeepSeek 1,271.78 1,258–1,285 $0.22 $0.33 164K
62 Mistral Medium 3.5 Mistral AI 1,267.07 1,251–1,283 $1.5 $7.5 262K
63 Gemini 3.1 Flash Lite Preview Google DeepMind 1,256.23 1,249–1,263 $0.25 $1.5 1,049K
64 KAT-Coder-Pro V1 Kuaishou (Kling, KAT) 1,255.01 1,235–1,275 $0.3 $1.2 256K
65 Qwen3.5 35B A3B Alibaba (Qwen) 1,250.46 1,233–1,268 $0.25 $2 262K
66 Grok 4.1 Fast Reasoning xAI 1,240.06 1,229–1,251 $0.2 $0.5 2,000K
67 Trinity Large Thinking Arcee AI 1,238.93 1,218–1,260 $0.22 $0.85 262K
68 Qwen3.5 Flash Alibaba (Qwen) 1,237.96 1,218–1,258 $0.09 $0.36 1,000K
69 Mistral Large 3 Mistral AI 1,230.02 1,204–1,256 $0.5 $1.5 256K
70 Gemini 2.5 Pro Google DeepMind 1,224.23 1,208–1,240 $1.25 $10 1,049K
71 Grok 4.1 xAI 1,210.33 1,185–1,235 $2 $10 200K
72 Granite 4.1 8B IBM 1,193.98 1,175–1,213 $0.05 $0.1 131K
73 Devstral 2 Mistral AI 1,193.89 1,173–1,215 $0.4 $2 256K
74 Mercury 2 Inception Labs 1,166.34 1,141–1,192 $0.25 $0.75 128K
75 Grok Code Fast 1 xAI 1,163.41 1,135–1,192 $0.2 $1.5 256K
76 Grok 4 Fast Reasoning xAI 1,160.41 1,133–1,188 $0.2 $0.5 2,000K
77 Devstral Medium Mistral AI 1,079.26 1,048–1,111 $0.4 $2 128K

The table scrolls sideways: not all columns fit.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 70 models — their order among themselves is not determined by the available data.

The confidence interval for half the models is wider than 19.2 points — that is about 2 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 6,310: the fewer there are, the wider the interval.

Arena (LMArena) — license CC-BY-4.0 · primary source
По данным Arena (LMArena)