How many programming tasks the model actually solved on fixed benchmark sets — a verifiable result, not an opinion. The collection has a second view, based on blind human ratings: the order there is different, and the gap between the two lists says more than either list on its own.
One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.
Not every model is measured here. This table is missing 3 of the top ten from the “By human votes” tab — MiMo V2.5 Pro, Gemini 3.6 Flash, Muse Spark 1.1. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 24 July 2026.
| # | Model | Developer | Task score 0–100, adjusted for coverage | Sign in$ / 1M | Output$ / 1M | Contexttokens | Aider Polyglot | Terminal-Bench | SWE-bench Verified | GSO-Bench | FrontierCode | Task setsmeasured |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.7 | Anthropic | 59.79 64.22 | $5 | $25 | 1,000K | — | 90.2 | 83.5 | 44.1 | 38.5 | 5 |
| 2 | Claude Opus 4.6 | Anthropic | 59.43 66.57 | $5 | $25 | 1,000K | — | 79.8 | 78.7 | 41.2 | — | 3 |
| 3 | GPT 5.4 2026-03-05 | OpenAI | 57.49 63.34 | $2.5 | $15 | 1,050K | — | 81.8 | 76.9 | 31.4 | — | 3 |
| 4 | Gemini 3.5 Flash | Google DeepMind | 56.64 64.57 | $1.5 | $9 | 1,049K | — | — | 79.3 | — | — | 2 |
| 5 | Claude Fable 5 | Anthropic | 55.96 63.2 | $10 | $50 | 1,000K | — | — | — | — | 53.5 | 2 |
| 6 | GLM 5 | Zhipu AI / Z.ai | 55.48 62.24 | $0.48 | $1.9 | 205K | — | 52.4 | 72.1 | — | — | 2 |
| 7 | Kimi K2.6 | Moonshot AI (Kimi) | 55.42 62.13 | $0.22 | $1.14 | 262K | — | — | 76.7 | — | — | 2 |
| 8 | Claude Opus 5 | Anthropic | 55.21 61.7 | $5 | $25 | 1,000K | — | — | — | — | 53.4 | 2 |
| 9 | Claude Sonnet 4.6 | Anthropic | 55.01 59.2 | $3 | $15 | 1,000K | — | 53.4 | 75.2 | — | — | 3 |
| 10 | o3 2025-04-16 | OpenAI | 53.98 56.61 | $2 | $8 | 200K | 81.3 | — | 62.3 | 8.8 | — | 4 |
| 11 | o1 2024-12-17 | OpenAI | 53.78 58.85 | $15 | $60 | 200K | 61.7 | — | — | — | — | 2 |
| 12 | GPT-5.6 Sol | OpenAI | 53.03 57.35 | $5 | $30 | 1,050K | — | — | — | — | 47.5 | 2 |
| 13 | Claude 4.5 Opus | Anthropic | 52.74 55.42 | $5 | $25 | 200K | — | 63.1 | 76.7 | 26.5 | — | 3 |
| 14 | GPT 5 2025-08-07 | OpenAI | 52.58 54.51 | $1.25 | $10 | 272K | 88 | 49.6 | 73.6 | 6.9 | — | 4 |
| 15 | Claude 4.1 Opus | Anthropic | 52.2 55.68 | $15 | $75 | 200K | — | 38 | 73.4 | — | — | 2 |
| 16 | Gemini 3 Pro Preview | Google DeepMind | 51.67 53.64 | $2 | $12 | 1,049K | — | 69.4 | 72.9 | 18.6 | — | 3 |
| 17 | Grok 4 | xAI | 51.06 53.4 | $3 | $15 | 256K | 79.6 | 27.2 | — | — | — | 2 |
| 18 | GLM 5.2 | Zhipu AI / Z.ai | 51.05 52.6 | $0.42 | $1.32 | 1,049K | — | — | 78.7 | — | 24.5 | 3 |
| 19 | Claude Opus 4.8 | Anthropic | 50.96 52.45 | $5 | $25 | 1,000K | — | — | — | 47.1 | 46.5 | 3 |
| 20 | Gemini 3.1 Pro Preview | Google DeepMind | 50.05 51.38 | $2 | $12 | 1,049K | — | 80.2 | — | 22.6 | — | 2 |
| 21 | Claude 4 Opus | Anthropic | 49.4 49.85 | $15 | $75 | 200K | 72 | — | 70.7 | 6.9 | — | 3 |
| 22 | Gemini 3 Flash Preview | Google DeepMind | 49.39 49.84 | $0.5 | $3 | 1,049K | — | 64.3 | 75.4 | 9.8 | — | 3 |
| 23 | Kimi K2.5 | Moonshot AI (Kimi) | 49.26 49.62 | $0.35 | $1.7 | 262K | — | 43.2 | 73.8 | — | — | 3 |
| 24 | GPT 5 Mini 2025-08-07 | OpenAI | 49.23 49.74 | $0.25 | $2 | 272K | — | 34.8 | 64.7 | — | — | 2 |
| 25 | DeepSeek V4 Pro | DeepSeek | 48.17 47.62 | $0.44 | $0.87 | 1,049K | — | — | 77.6 | — | 17.6 | 2 |
| 26 | GPT 4.1 2025-04-14 | OpenAI | 48.07 47.65 | $2 | $8 | 1,048K | 52.4 | — | 48.5 | — | — | 3 |
| 27 | Claude Sonnet 5 | Anthropic | 47.72 47.05 | $2 | $10 | 1,000K | — | — | — | 37.3 | 42.7 | 3 |
| 28 | o4 Mini 2025-04-16 | OpenAI | 47.01 45.87 | $1.1 | $4.4 | 200K | 72 | — | — | 3.6 | — | 3 |
| 29 | Gemini 2.5 Pro | Google DeepMind | 46.9 45.08 | $1.25 | $10 | 1,049K | — | 32.6 | 57.6 | — | — | 2 |
| 30 | Claude Sonnet 3.7 | Anthropic | 46.85 45.91 | $3 | $15 | 200K | 64.9 | — | 61 | 3.8 | — | 4 |
| 31 | Gemini 2.5 Pro Preview 0605 | Google DeepMind | 46.11 43.51 | $1.13 | $9 | 1,049K | 83.1 | — | — | 3.9 | — | 2 |
| 32 | Claude Sonnet 4.5 | Anthropic | 45.98 44.16 | $3 | $15 | 1,000K | — | 46.5 | 71.3 | 14.7 | — | 3 |
| 33 | o3 Mini 2025-01-31 | OpenAI | 42.63 38.57 | $1.1 | $4.4 | 200K | 60.4 | — | — | 1.3 | — | 3 |
| 34 | Claude 4 Sonnet | Anthropic | 40.91 33.1 | $3 | $15 | 1,000K | 61.3 | — | — | 4.9 | — | 2 |
| 35 | Claude 3.5 Sonnet 2024-10-22 | Anthropic | 40.33 34.73 | $3 | $15 | 200K | 51.6 | — | — | 4.6 | — | 3 |
| 36 | GPT OSS 120B | OpenAI | 39.48 30.25 | $0.03 | $0.14 | 131K | 41.8 | 18.7 | — | — | — | 2 |
| 37 | Claude 3.5 Haiku | Anthropic | 39.36 30 | $0.25 | $1.25 | 200K | 28 | — | — | — | — | 2 |
| 38 | Kimi K2 Instruct | Moonshot AI (Kimi) | 37.85 30.6 | $0.1 | $2 | 131K | 59.1 | 27.8 | — | 4.9 | — | 3 |
| 39 | GPT-4o (2024-08-06) | OpenAI | 36.63 24.55 | $2.5 | $10 | 128K | 23.1 | — | — | — | — | 2 |
| 40 | OpenAI GPT-4.1 Mini | OpenAI | 36.46 24.2 | $0.4 | $1.6 | 1,048K | 32.4 | — | — | — | — | 2 |
| 41 | GPT-4o (2024-11-20) | OpenAI | 29.32 16.4 | $2.5 | $10 | 128K | 18.2 | — | 31 | 0 | — | 3 |
The table scrolls sideways: not all columns fit.
The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero. Showing the five most complete task sets out of 7; the rest are on the model page.
The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
Epoch AI
— license CC-BY · primary source
По данным Epoch AI
Not every model is measured here. This table is missing 1 of the top ten from the “By solved tasks” tab — GPT 5.4 2026-03-05. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 27 July 2026.
| # | Model | Developer | Arena Score, coding | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 1,534.96 1,529–1,541 | $5 | $25 | 1,000K |
| 2 | Claude Opus 5 ≈ | Anthropic | 1,534.15 1,510–1,559 | $5 | $25 | 1,000K |
| 3 | Claude Fable 5 ≈ | Anthropic | 1,518.95 1,509–1,529 | $10 | $50 | 1,000K |
| 4 | Claude Opus 4.7 ≈ | Anthropic | 1,516.56 1,510–1,523 | $5 | $25 | 1,000K |
| 5 | Claude Sonnet 4.6 | Anthropic | 1,502.48 1,497–1,508 | $3 | $15 | 1,000K |
| 6 | Kimi K3 ≈ | Moonshot AI (Kimi) | 1,499.63 1,480–1,519 | $2 | $8 | 1,049K |
| 7 | MiMo V2.5 Pro ≈ | Xiaomi | 1,498.44 1,492–1,505 | $0.44 | $0.87 | 1,050K |
| 8 | Claude 4.5 Opus ≈ | Anthropic | 1,498.02 1,493–1,503 | $5 | $25 | 200K |
| 9 | Gemini 3.6 Flash ≈ | Google DeepMind | 1,495.38 1,481–1,510 | $1.5 | $7.5 | 1,049K |
| 10 | Muse Spark 1.1 ≈ | Meta AI | 1,495.35 1,484–1,507 | $1.25 | $4.25 | 1,049K |
| 11 | GLM 5.1 ≈ | Zhipu AI / Z.ai | 1,493.04 1,486–1,500 | $0.06 | $0.22 | 205K |
| 12 | Gemini 3.5 Flash ≈ | Google DeepMind | 1,492.01 1,480–1,504 | $1.5 | $9 | 1,049K |
| 13 | ERNIE 5.1 ≈ | Baidu (Ernie) | 1,488.69 1,482–1,496 | $0.75 | $3 | 119K |
| 14 | GPT-5.6 Terra ≈ | OpenAI | 1,488.59 1,476–1,501 | $2 | $12 | 1,050K |
| 15 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 1,486.74 1,480–1,494 | $0.22 | $1.14 | 262K |
| 16 | GPT-5.6 Sol ≈ | OpenAI | 1,486.07 1,473–1,499 | $5 | $30 | 1,050K |
| 17 | Claude Opus 4.8 ≈ | Anthropic | 1,485.6 1,478–1,493 | $5 | $25 | 1,000K |
| 18 | Claude Sonnet 5 ≈ | Anthropic | 1,485.59 1,476–1,495 | $2 | $10 | 1,000K |
| 19 | Claude Sonnet 4.5 ≈ | Anthropic | 1,485.37 1,480–1,490 | $3 | $15 | 1,000K |
| 20 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 1,483.89 1,479–1,489 | $2 | $12 | 1,049K |
| 21 | Gemini 3 Pro ≈ | Google DeepMind | 1,482.68 1,476–1,490 | $1.6 | $9.6 | 1,049K |
| 22 | Qwen3.7 Plus ≈ | Alibaba (Qwen) | 1,481.93 1,474–1,490 | $0.5 | $3 | 1,000K |
| 23 | Grok 4.5 ≈ | xAI | 1,479.12 1,468–1,491 | $2 | $6 | 500K |
| 24 | GLM 5.2 ≈ | Zhipu AI / Z.ai | 1,477.24 1,469–1,486 | $0.42 | $1.32 | 1,049K |
| 25 | MiMo V2 Pro ≈ | Xiaomi | 1,476.32 1,469–1,484 | $0.44 | $0.87 | 1,049K |
| 26 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 1,473.6 1,468–1,479 | $0.3 | $1.9 | 256K |
| 27 | Claude 4.1 Opus ≈ | Anthropic | 1,473.09 1,468–1,478 | $15 | $75 | 200K |
| 28 | MiniMax M3 ≈ | MiniMax | 1,472.57 1,465–1,480 | $0.3 | $1.2 | 1,049K |
| 29 | GPT-5.6 Luna ≈ | OpenAI | 1,471.79 1,460–1,484 | $0.2 | $1.2 | 1,050K |
| 30 | Seed 2.0 Pro ≈ | ByteDance (Doubao) | 1,471.1 1,465–1,477 | $0.5 | $3 | 131K |
| 31 | DeepSeek V4 Pro ≈ | DeepSeek | 1,470.43 1,464–1,477 | $0.44 | $0.87 | 1,049K |
| 32 | MiMo V2.5 ≈ | Xiaomi | 1,469.48 1,463–1,476 | $0.14 | $0.28 | 1,050K |
| 33 | Longcat Flash Chat ≈ | Meituan | 1,467.98 1,455–1,481 | — | — | 131K |
| 34 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 1,467.51 1,456–1,479 | $0.7 | $3.1 | 203K |
| 35 | Hy3 ≈ | Tencent (Hunyuan) | 1,467.42 1,448–1,487 | $0.13 | $0.53 | 262K |
| 36 | Qwen3.6 Max Preview ≈ | Alibaba (Qwen) | 1,467.25 1,452–1,483 | $1.3 | $7.8 | 262K |
| 37 | Inkling ≈ | Thinking Machines Lab | 1,467.21 1,454–1,480 | $1.87 | $4.68 | 1,049K |
| 38 | Qwen3.6 Plus ≈ | Alibaba (Qwen) | 1,465.74 1,459–1,472 | $0.5 | $3 | 1,000K |
| 39 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 1,465.71 1,460–1,471 | $0.6 | $3.6 | 262K |
| 40 | MiMo V2 Omni ≈ | Xiaomi | 1,465.04 1,456–1,474 | $0.14 | $0.28 | 262K |
| 41 | Gemini 3 Flash ≈ | Google DeepMind | 1,462.61 1,455–1,470 | $0.4 | $2.4 | 1,049K |
| 42 | GLM 5 ≈ | Zhipu AI / Z.ai | 1,460.36 1,453–1,468 | $0.48 | $1.9 | 205K |
| 43 | Mistral Medium 3.5 ≈ | Mistral AI | 1,460.3 1,449–1,471 | $1.5 | $7.5 | 262K |
| 44 | Grok 4.20 (Reasoning) ≈ | xAI | 1,458.81 1,453–1,465 | $2 | $6 | 2,000K |
| 45 | Grok 4.20 Multi Agent Beta 0309 ≈ | xAI | 1,457.78 1,452–1,464 | $2 | $6 | 2,000K |
| 46 | Qwen3 Max Preview ≈ | Alibaba (Qwen) | 1,456.92 1,449–1,465 | $1.2 | $6 | 256K |
| 47 | DeepSeek V4 Flash ≈ | DeepSeek | 1,456.12 1,450–1,463 | $0.14 | $0.28 | 1,049K |
| 48 | GLM 4.7 ≈ | Zhipu AI / Z.ai | 1,455.74 1,444–1,468 | $0.15 | $0.8 | 205K |
| 49 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 1,455.6 1,441–1,470 | $0.3 | $2.5 | 1,049K |
| 50 | Gemma 4 31B ≈ | Google DeepMind | 1,454.93 1,440–1,470 | $0.14 | $0.4 | 262K |
| 51 | Kimi K2 Thinking Turbo ≈ | Moonshot AI (Kimi) | 1,453.31 1,448–1,459 | $1.15 | $8 | 262K |
| 52 | Gemini 2.5 Pro ≈ | Google DeepMind | 1,452.45 1,448–1,457 | $1.25 | $10 | 1,049K |
| 53 | MiniMax M2.7 ≈ | MiniMax | 1,451.3 1,445–1,457 | $0.3 | $1.2 | 205K |
| 54 | Claude Haiku 4.5 ≈ | Anthropic | 1,450.74 1,446–1,455 | $1 | $5 | 200K |
| 55 | GLM 4.6 ≈ | Zhipu AI / Z.ai | 1,449.4 1,442–1,457 | $0.29 | $1.14 | 205K |
| 56 | DeepSeek V3.2 ≈ | DeepSeek | 1,448.64 1,442–1,455 | $0.28 | $0.4 | 164K |
| 57 | GPT-5.2 Chat ≈ | OpenAI | 1,446.66 1,440–1,454 | $1.75 | $14 | 128K |
| 58 | Qwen3 235B A22B Instruct 2507 ≈ | Alibaba (Qwen) | 1,444.78 1,440–1,449 | $0.1 | $0.1 | 262K |
| 59 | Grok 4.1 ≈ | xAI | 1,444.4 1,439–1,450 | $2 | $10 | 200K |
| 60 | Gemma 4 26B-A4B ≈ | Google DeepMind | 1,444.01 1,429–1,459 | $0.05 | $0.29 | 262K |
| 61 | Mistral Large 3 ≈ | Mistral AI | 1,443.94 1,438–1,450 | $0.5 | $1.5 | 256K |
| 62 | Qwen3 Next 80B A3B Instruct ≈ | Alibaba (Qwen) | 1,441.15 1,433–1,450 | $0.15 | $1.2 | 262K |
| 63 | MiMo V2 Flash ≈ | Xiaomi | 1,441.11 1,435–1,447 | $0.14 | $0.28 | 262K |
| 64 | Qwen3 VL 235B A22B Instruct ≈ | Alibaba (Qwen) | 1,440.09 1,427–1,453 | $0.4 | $1.6 | 262K |
| 65 | Qwen3 Max 2025-09-23 ≈ | Alibaba (Qwen) | 1,439.19 1,426–1,452 | $0.86 | $3.43 | 258K |
| 66 | Step 3.5 Flash ≈ | StepFun | 1,436.75 1,431–1,443 | $0.1 | $0.3 | 262K |
| 67 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 1,436.07 1,429–1,443 | $0.4 | $3.2 | 262K |
| 68 | Mistral Medium 3.1 ≈ | Mistral AI | 1,434.76 1,430–1,439 | $0.4 | $2 | 262K |
| 69 | GPT-5.5 Instant ≈ | OpenAI | 1,434.45 1,427–1,442 | $5 | $30 | 400K |
| 70 | GLM 4.5 ≈ | Zhipu AI / Z.ai | 1,433.35 1,425–1,442 | $0.2 | $0.8 | 131K |
| 71 | DeepSeek V3.2 Exp ≈ | DeepSeek | 1,433.29 1,422–1,445 | $0.22 | $0.33 | 164K |
| 72 | Qwen3 VL 235B A22B Thinking ≈ | Alibaba (Qwen) | 1,427.7 1,413–1,442 | $0.4 | $4 | 262K |
| 73 | DeepSeek R1 0528 ≈ | DeepSeek | 1,426.99 1,416–1,438 | $0.25 | $0.25 | 164K |
| 74 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 1,424.82 1,417–1,432 | $0.3 | $2.4 | 262K |
| 75 | Qwen3 235B A22B Thinking-2507 ≈ | Alibaba (Qwen) | 1,423.79 1,409–1,438 | $0.1 | $0.1 | 262K |
| 76 | Gemini 2.5 Flash ≈ | Google DeepMind | 1,423.41 1,419–1,428 | $0.3 | $2.5 | 1,049K |
| 77 | Grok 4 Fast Reasoning ≈ | xAI | 1,418.56 1,409–1,428 | $0.2 | $0.5 | 2,000K |
| 78 | Qwen3 30B A3B Instruct 2507 ≈ | Alibaba (Qwen) | 1,417.48 1,409–1,426 | $0.05 | $0.19 | 262K |
| 79 | MiMo V2 Flash (Thinking) ≈ | Xiaomi | 1,417.25 1,405–1,429 | $0.1 | $0.31 | 256K |
| 80 | Grok 4.3 ≈ | xAI | 1,415.7 1,409–1,422 | $1.25 | $2.5 | 1,000K |
| 81 | DeepSeek V3.1 ≈ | DeepSeek | 1,414.76 1,403–1,426 | $0.2 | $0.7 | 164K |
| 82 | ChatGPT 4o Latest ≈ | OpenAI | 1,414.41 1,409–1,419 | $5 | $15 | 128K |
| 83 | Qwen3.5 Flash ≈ | Alibaba (Qwen) | 1,413.61 1,408–1,419 | $0.09 | $0.36 | 1,000K |
| 84 | Qwen3 Coder 480B A35B ≈ | Alibaba (Qwen) | 1,412.31 1,403–1,421 | $1.5 | $7.5 | 262K |
| 85 | Grok 4.1 Fast Reasoning ≈ | xAI | 1,411.95 1,406–1,418 | $0.2 | $0.5 | 2,000K |
| 86 | Qwen3.5 35B A3B ≈ | Alibaba (Qwen) | 1,409.26 1,402–1,417 | $0.25 | $2 | 262K |
| 87 | Grok 4 ≈ | xAI | 1,408.98 1,402–1,416 | $3 | $15 | 256K |
| 88 | o3 2025-04-16 ≈ | OpenAI | 1,408.19 1,402–1,414 | $2 | $8 | 200K |
| 89 | DeepSeek V3.1 Terminus ≈ | DeepSeek | 1,406.59 1,386–1,427 | $0.21 | $0.79 | 164K |
| 90 | GPT-5.3 Chat (latest) ≈ | OpenAI | 1,406.23 1,399–1,413 | $1.75 | $14 | 128K |
| 91 | Nemotron 3 Super ≈ | NVIDIA | 1,405.97 1,392–1,420 | $0.2 | $0.8 | 1,000K |
| 92 | NVIDIA Nemotron 3 Super 120B A12B ≈ | NVIDIA | 1,405.97 1,392–1,420 | $0.15 | $0.65 | 262K |
| 93 | Claude 4 Opus ≈ | Anthropic | 1,401.8 1,395–1,409 | $15 | $75 | 200K |
| 94 | 2025) ≈ | Google DeepMind | 1,401.55 1,394–1,409 | $0.3 | $2.5 | 1,049K |
| 95 | GPT 5 Chat ≈ | OpenAI | 1,400.27 1,392–1,408 | $1.25 | $10 | 128K |
| 96 | Kimi K2 0905 Preview ≈ | Moonshot AI (Kimi) | 1,399.71 1,387–1,412 | $0.6 | $2.5 | 262K |
| 97 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 1,399.63 1,394–1,406 | $0.25 | $1.5 | 1,049K |
| 98 | GPT 4.5 Preview ≈ | OpenAI | 1,397.05 1,384–1,410 | $75 | $150 | 128K |
| 99 | GLM 4.5 Air ≈ | Zhipu AI / Z.ai | 1,396.36 1,389–1,404 | $0.11 | $0.29 | 131K |
| 100 | Mercury 2 ≈ | Inception Labs | 1,392.87 1,372–1,413 | $0.25 | $0.75 | 128K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 98 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 15.2 points — that is about 11 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 7,317: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)