The share of solved mathematical problems, from olympiad to research level. The percentages are lower than in other collections, and that is a property of the problems, not of the models: comparing them with the numbers from the knowledge collection is meaningless. The second view comes from blind human ratings.
One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.
Not every model is measured here. This table is missing 2 of the top ten from the “By human votes” tab — Gemini 3.6 Flash, ERNIE 5.1. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 27 July 2026.
| # | Model | Developer | Task score 0–100, adjusted for coverage | Sign in$ / 1M | Output$ / 1M | Contexttokens | Mock AIME 2024–2025 | FrontierMath, level 4 | FrontierMath, levels 1–3 | MATH, difficulty level five | Task setsmeasured |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 74.77 91.52 | $10 | $50 | 1,000K | 99.7 | 87.8 | 87 | — | 3 |
| 2 | GPT-5.6 Sol | OpenAI | 72.1 94.56 | $5 | $30 | 1,050K | 100 | — | 89.1 | — | 2 |
| 3 | Claude Opus 5 | Anthropic | 71.39 85.89 | $5 | $25 | 1,000K | 98.9 | 73.2 | 85.6 | — | 3 |
| 4 | GPT-5.6 Terra | OpenAI | 71.14 85.47 | $2 | $12 | 1,050K | 99.7 | 70.7 | 86 | — | 3 |
| 5 | o3 2025-04-16 | OpenAI | 70.23 90.82 | $2 | $8 | 200K | 83.9 | — | — | 97.8 | 2 |
| 6 | GPT-5.6 Luna | OpenAI | 68.14 80.47 | $0.2 | $1.2 | 1,050K | 98.3 | 61 | 82.1 | — | 3 |
| 7 | Qwen3 Max 2025-09-23 | Alibaba (Qwen) | 67.43 85.22 | $0.86 | $3.43 | 258K | 73.3 | — | — | 97.1 | 2 |
| 8 | Grok 3 Mini | xAI | 67 84.35 | $0.3 | $0.5 | 131K | 77.8 | — | — | 90.9 | 2 |
| 9 | o1 2024-12-17 | OpenAI | 66.83 84.01 | $15 | $60 | 200K | 73.3 | — | — | 94.7 | 2 |
| 10 | Claude Opus 4.8 | Anthropic | 66.74 78.14 | $5 | $25 | 1,000K | 98.3 | 56.1 | 80 | — | 3 |
| 11 | Claude Haiku 4.5 | Anthropic | 65.57 81.5 | $1 | $5 | 200K | 66.6 | — | — | 96.4 | 2 |
| 12 | DeepSeek R1 0528 | DeepSeek | 65.57 81.5 | $0.25 | $0.25 | 164K | 66.4 | — | — | 96.6 | 2 |
| 13 | GPT 5.4 2026-03-05 | OpenAI | 64.93 75.13 | $2.5 | $15 | 1,050K | 97.8 | 49 | 78.6 | — | 3 |
| 14 | Claude 4 Sonnet | Anthropic | 63.69 77.73 | $3 | $15 | 1,000K | 71.1 | — | — | 84.4 | 2 |
| 15 | Claude 4 Opus | Anthropic | 62.19 74.73 | $15 | $75 | 200K | 64.4 | — | — | 85.1 | 2 |
| 16 | Claude Sonnet 3.7 | Anthropic | 62.05 74.45 | $3 | $15 | 200K | 57.7 | — | — | 91.2 | 2 |
| 17 | Kimi K3 | Moonshot AI (Kimi) | 61.54 69.47 | $2 | $8 | 1,049K | 97.2 | 39 | 72.2 | — | 3 |
| 18 | DeepSeek Reasoner | DeepSeek | 61.41 73.17 | $0.55 | $2.19 | 164K | 53.3 | — | — | 93.1 | 2 |
| 19 | GPT 5 2025-08-07 | OpenAI | 61.03 66.73 | $1.25 | $10 | 272K | 91.4 | 22 | 55.4 | 98.1 | 4 |
| 20 | Grok 3 | xAI | 60.89 72.13 | $3 | $15 | 131K | 55.5 | — | — | 88.8 | 2 |
| 21 | GPT 5.4 Pro 2026-03-05 | OpenAI | 60.07 70.5 | $30 | $180 | 1,050K | — | 58.5 | 82.5 | — | 2 |
| 22 | Claude Opus 4.7 | Anthropic | 59.8 66.56 | $5 | $25 | 1,000K | 97.8 | 31.7 | 70.2 | — | 3 |
| 23 | o1 Mini 2024-09-12 | OpenAI | 58.84 68.04 | $1.1 | $4.4 | 128K | 46.9 | — | — | 89.2 | 2 |
| 24 | Qwen3.7 Max | Alibaba (Qwen) | 58.6 64.57 | $2.5 | $7.5 | 1,000K | 95 | 34.2 | 64.6 | — | 3 |
| 25 | OpenAI GPT-4.1 Mini | OpenAI | 57.81 65.98 | $0.4 | $1.6 | 1,048K | 44.7 | — | — | 87.3 | 2 |
| 26 | Claude Sonnet 5 | Anthropic | 57.78 63.2 | $2 | $10 | 1,000K | 94.7 | 29.3 | 65.6 | — | 3 |
| 27 | Claude Opus 4.6 | Anthropic | 57.31 62.41 | $5 | $25 | 1,000K | 94.4 | 26.8 | 66 | — | 3 |
| 28 | GPT 5 Mini 2025-08-07 | OpenAI | 57.11 60.84 | $0.25 | $2 | 272K | 86.7 | 12.2 | 46.7 | 97.9 | 4 |
| 29 | Gemini 3.5 Flash | Google DeepMind | 56.9 61.73 | $1.5 | $9 | 1,049K | 95.6 | 26.8 | 62.8 | — | 3 |
| 30 | Gemini 3.1 Pro Preview | Google DeepMind | 56.27 60.69 | $2 | $12 | 1,049K | 95.6 | 26.8 | 59.7 | — | 3 |
| 31 | Grok 4.5 | xAI | 55.73 59.79 | $2 | $6 | 500K | 97.8 | 24.4 | 57.2 | — | 3 |
| 32 | Kimi K2.6 | Moonshot AI (Kimi) | 55.65 59.65 | $0.22 | $1.14 | 262K | 96.1 | 25.6 | 57.2 | — | 3 |
| 33 | GPT 4.1 2025-04-14 | OpenAI | 55.14 60.64 | $2 | $8 | 1,048K | 38.3 | — | — | 83 | 2 |
| 34 | GLM 5.2 | Zhipu AI / Z.ai | 54.83 58.29 | $0.42 | $1.32 | 1,049K | 86.4 | 29.3 | 59.2 | — | 3 |
| 35 | GPT 5.2 Pro 2025-12-11 | OpenAI | 54.82 60 | $21 | $168 | 272K | — | 46 | 74 | — | 2 |
| 36 | GPT 4.5 Preview | OpenAI | 53.91 58.18 | $75 | $150 | 128K | 37.7 | — | — | 78.6 | 2 |
| 37 | o4 Mini 2025-04-16 | OpenAI | 53.3 55.13 | $1.1 | $4.4 | 200K | 81.7 | 4.9 | 36.1 | 97.8 | 4 |
| 38 | Mistral Medium 3 | Mistral AI | 53.27 56.89 | $0.4 | $2 | 131K | 32.2 | — | — | 81.6 | 2 |
| 39 | DeepSeek Chat 0324 | DeepSeek | 53.14 56.64 | $0.2 | $0.6 | 164K | 37.7 | — | — | 75.6 | 2 |
| 40 | o1 Preview 2024-09-12 | OpenAI | 53 56.35 | $15 | $60 | 128K | 31 | — | — | 81.7 | 2 |
| 41 | Kimi K2.7 Code | Moonshot AI (Kimi) | 52.38 54.21 | $0.28 | $1.1 | 262K | 96.4 | 12.2 | 54 | — | 3 |
| 42 | Gemini 3 Flash Preview | Google DeepMind | 52.07 53.69 | $0.5 | $3 | 1,049K | 92.8 | 17.1 | 51.2 | — | 3 |
| 43 | Grok 4.20 (Reasoning) | xAI | 50.7 51.4 | $1.25 | $2.5 | 1,000K | 92.2 | 17.1 | 44.9 | — | 3 |
| 44 | Claude Sonnet 4.5 | Anthropic | 50.18 50.45 | $3 | $15 | 1,000K | 77.8 | 2.4 | 23.9 | 97.7 | 4 |
| 45 | Grok 4.3 | xAI | 50.01 50.26 | $1.25 | $2.5 | 1,000K | 93.3 | 14.6 | 42.8 | — | 3 |
| 46 | GPT 5 Nano 2025-08-07 | OpenAI | 49.68 49.69 | $0.05 | $0.4 | 272K | 81.1 | 2.4 | 20 | 95.2 | 4 |
| 47 | GPT 4.1 Nano 2025-04-14 | OpenAI | 49.53 49.41 | $0.1 | $0.4 | 1,048K | 28.8 | — | — | 70 | 2 |
| 48 | GPT 5.4 Mini 2026-03-17 | OpenAI | 49.5 49.4 | $0.75 | $4.5 | 272K | 87.2 | 9.8 | 51.2 | — | 3 |
| 49 | GPT 5.4 Nano 2026-03-17 | OpenAI | 48.83 48.29 | $0.2 | $1.25 | 272K | 87.8 | 12.2 | 44.9 | — | 3 |
| 50 | DeepSeek V4 Pro | DeepSeek | 48.73 48.12 | $0.44 | $0.87 | 1,049K | 96.7 | 2.4 | 45.3 | — | 3 |
| 51 | o3 Mini 2025-01-31 | OpenAI | 48.55 48 | $1.1 | $4.4 | 200K | 76.9 | 0 | 18.6 | 96.5 | 4 |
| 52 | Gemma 3 27B | Google DeepMind | 48.24 46.84 | $0.03 | $0.11 | 131K | 19.6 | — | — | 74 | 2 |
| 53 | Llama 4 Maverick 17b 128e Instruct | Meta AI | 48.2 46.75 | $0.05 | $0.1 | 1,049K | 20.5 | — | — | 73 | 2 |
| 54 | Qwen2.5 Max 2025-01-25 | Alibaba (Qwen) | 45.63 41.61 | — | — | 128K | 16 | — | — | 67.2 | 2 |
| 55 | DeepSeek V3 | DeepSeek | 44.97 40.3 | $0.27 | $1.1 | 164K | 15.8 | — | — | 64.9 | 2 |
| 56 | Claude 4.5 Opus | Anthropic | 44.93 41.79 | $5 | $25 | 200K | 86.1 | 4.9 | 34.4 | — | 3 |
| 57 | Phi 4 | Microsoft | 44.47 39.3 | $0.06 | $0.14 | 128K | 13.7 | — | — | 64.9 | 2 |
| 58 | GPT 5 Pro 2025-10-06 | OpenAI | 43.65 37.65 | $15 | $120 | 400K | — | 19.5 | 55.8 | — | 2 |
| 59 | Grok 2 1212 | xAI | 43.56 37.48 | $2 | $10 | 131K | 11.4 | — | — | 63.5 | 2 |
| 60 | Qwen2.5 72B Instruct | Alibaba (Qwen) | 42.61 35.57 | $1.4 | $5.6 | 131K | 8 | — | — | 63.2 | 2 |
| 61 | Llama 4 Scout 17B 16E Instruct | Meta AI | 42.31 34.98 | $0.05 | $0.1 | 10,000K | 7.7 | — | — | 62.3 | 2 |
| 62 | Gemini 2.5 Pro | Google DeepMind | 41.71 36.42 | $1.25 | $10 | 1,049K | 84.7 | 0 | 24.6 | — | 3 |
| 63 | Claude 3.5 Sonnet 2024-10-22 | Anthropic | 41.16 32.67 | $3 | $15 | 200K | 8.4 | — | — | 57 | 2 |
| 64 | GPT-4o (2024-08-06) | OpenAI | 39.72 29.79 | $2.5 | $10 | 128K | 6.3 | — | — | 53.3 | 2 |
| 65 | OpenAI: GPT-4o-mini (2024-07-18) | OpenAI | 39.69 29.74 | $0.15 | $0.6 | 128K | 6.9 | — | — | 52.6 | 2 |
| 66 | Llama 3.1 405B Instruct | Meta AI | 39.67 29.7 | $0.12 | $0.3 | 128K | 9.6 | — | — | 49.8 | 2 |
| 67 | Claude Sonnet 3.5 | Anthropic | 39.35 29.06 | $2.6 | $13 | 200K | 6.4 | — | — | 51.7 | 2 |
| 68 | Mistral Large | Mistral AI | 39.32 28.99 | $2 | $6 | 131K | 7.7 | — | — | 50.3 | 2 |
| 69 | GPT-4o (2024-05-13) | OpenAI | 39.13 28.61 | $5 | $15 | 128K | 6.2 | — | — | 51.1 | 2 |
| 70 | GPT-4o (2024-11-20) | OpenAI | 38.81 27.97 | $2.5 | $10 | 128K | 6.2 | — | — | 49.8 | 2 |
| 71 | GPT 4 Turbo 2024-04-09 | OpenAI | 38.15 26.65 | $10 | $30 | 128K | 6.6 | — | — | 46.7 | 2 |
| 72 | Mistral Large 2407 | Mistral AI | 38.12 26.6 | $3 | $9 | 131K | 8.4 | — | — | 44.8 | 2 |
| 73 | Mistral Small 3.1 | Mistral AI | 37.95 26.26 | $0.1 | $0.3 | 128K | 5.7 | — | — | 46.8 | 2 |
| 74 | Claude 3.5 Haiku | Anthropic | 37.47 25.29 | $0.25 | $1.25 | 200K | 4.2 | — | — | 46.4 | 2 |
| 75 | Claude 4.1 Opus | Anthropic | 36.64 27.98 | $15 | $75 | 200K | 68.9 | 2.4 | 12.6 | — | 3 |
| 76 | Llama 3.3 70B Instruct | Meta AI | 36.48 23.32 | $0.05 | $0.23 | 131K | 5 | — | — | 41.6 | 2 |
| 77 | Claude 3 Opus 2024-02-29 | Anthropic | 35.35 21.06 | $15 | $75 | 200K | 4.6 | — | — | 37.5 | 2 |
| 78 | Llama 3.2 90B Vision Instruct | Meta AI | 35.32 20.99 | $0.35 | $0.4 | 128K | 2.5 | — | — | 39.4 | 2 |
| 79 | Llama 3.1 70B | Meta AI | 34.87 20.1 | $0.12 | $0.3 | 131K | 3.5 | — | — | 36.7 | 2 |
| 80 | Google: Gemma 2 27B | Google DeepMind | 32.12 14.59 | $0.65 | $0.65 | 8K | 1.3 | — | — | 27.9 | 2 |
| 81 | Llama 3 70B Instruct | Meta AI | 31.52 13.39 | $0.12 | $0.3 | 8K | 4.2 | — | — | 22.6 | 2 |
| 82 | Mistral Large 2402 | Mistral AI | 31.4 13.16 | $4 | $12 | 32K | 1.9 | — | — | 24.5 | 2 |
| 83 | Llama 3.1 8B | Meta AI | 31.14 12.64 | $0.02 | $0.03 | 131K | 2.4 | — | — | 22.9 | 2 |
| 84 | GPT 4 0613 | OpenAI | 30.82 11.99 | $30 | $60 | 8K | 1 | — | — | 23 | 2 |
| 85 | Claude 3 Sonnet | Anthropic | 29.97 10.29 | $3 | $15 | 200K | 2.4 | — | — | 18.2 | 2 |
| 86 | Anthropic: Claude 3 Haiku | Anthropic | 28.97 8.3 | $0.25 | $1.25 | 200K | 1.7 | — | — | 14.9 | 2 |
| 87 | Meta: Llama 3 8B Instruct | Meta AI | 26.54 3.43 | $0.03 | $0.04 | 8K | 0.7 | — | — | 6.1 | 2 |
| 88 | Llama 2 70B Chat HF | Meta AI | 25.65 1.65 | $1 | $1 | 4K | 0 | — | — | 3.3 | 2 |
The table scrolls sideways: not all columns fit.
The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero.
The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
Epoch AI
— license CC-BY · primary source
По данным Epoch AI
| # | Model | Developer | Arena Score, math | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 1,544.36 1,507–1,582 | $5 | $25 | 1,000K |
| 2 | Claude Fable 5 ≈ | Anthropic | 1,534.02 1,513–1,555 | $10 | $50 | 1,000K |
| 3 | Gemini 3.5 Flash ≈ | Google DeepMind | 1,523.8 1,498–1,549 | $1.5 | $9 | 1,049K |
| 4 | Gemini 3.6 Flash ≈ | Google DeepMind | 1,508.21 1,473–1,543 | $1.5 | $7.5 | 1,049K |
| 5 | Claude Opus 4.6 ≈ | Anthropic | 1,507.89 1,498–1,518 | $5 | $25 | 1,000K |
| 6 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 1,489.07 1,480–1,498 | $2 | $12 | 1,049K |
| 7 | Claude Opus 4.7 ≈ | Anthropic | 1,488.38 1,476–1,500 | $5 | $25 | 1,000K |
| 8 | Grok 4.5 ≈ | xAI | 1,484.15 1,456–1,513 | $2 | $6 | 500K |
| 9 | ERNIE 5.1 ≈ | Baidu (Ernie) | 1,479.13 1,465–1,493 | $0.75 | $3 | 119K |
| 10 | GLM 5.1 ≈ | Zhipu AI / Z.ai | 1,478.76 1,464–1,493 | $0.06 | $0.22 | 205K |
| 11 | Inkling ≈ | Thinking Machines Lab | 1,478.45 1,446–1,511 | $1.87 | $4.68 | 1,049K |
| 12 | MiMo V2.5 Pro ≈ | Xiaomi | 1,477.73 1,465–1,491 | $0.44 | $0.87 | 1,050K |
| 13 | Gemini 3 Pro ≈ | Google DeepMind | 1,476.06 1,465–1,487 | $1.6 | $9.6 | 1,049K |
| 14 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 1,474.56 1,461–1,488 | $0.22 | $1.14 | 262K |
| 15 | Gemini 3 Flash ≈ | Google DeepMind | 1,473.99 1,461–1,487 | $0.4 | $2.4 | 1,049K |
| 16 | GPT-5.6 Terra ≈ | OpenAI | 1,472.82 1,444–1,502 | $2 | $12 | 1,050K |
| 17 | GLM 5.2 ≈ | Zhipu AI / Z.ai | 1,471.38 1,453–1,490 | $0.42 | $1.32 | 1,049K |
| 18 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 1,470.79 1,461–1,481 | $0.3 | $1.9 | 256K |
| 19 | Muse Spark 1.1 ≈ | Meta AI | 1,470.72 1,443–1,499 | $1.25 | $4.25 | 1,049K |
| 20 | Qwen3.7 Plus ≈ | Alibaba (Qwen) | 1,469.29 1,451–1,487 | $0.5 | $3 | 1,000K |
| 21 | Claude Opus 4.8 ≈ | Anthropic | 1,466.9 1,452–1,482 | $5 | $25 | 1,000K |
| 22 | Gemma 4 26B-A4B ≈ | Google DeepMind | 1,466.5 1,438–1,495 | $0.05 | $0.29 | 262K |
| 23 | Gemma 4 31B ≈ | Google DeepMind | 1,466.07 1,438–1,494 | $0.14 | $0.4 | 262K |
| 24 | Claude Sonnet 5 ≈ | Anthropic | 1,462.99 1,441–1,485 | $2 | $10 | 1,000K |
| 25 | Qwen3.6 Max Preview ≈ | Alibaba (Qwen) | 1,462.55 1,433–1,492 | $1.3 | $7.8 | 262K |
| 26 | Claude Sonnet 4.6 ≈ | Anthropic | 1,461.43 1,450–1,472 | $3 | $15 | 1,000K |
| 27 | Claude 4.5 Opus ≈ | Anthropic | 1,459.45 1,450–1,469 | $5 | $25 | 200K |
| 28 | Grok 4.20 (Reasoning) ≈ | xAI | 1,457.07 1,446–1,468 | $2 | $6 | 2,000K |
| 29 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 1,454.59 1,428–1,481 | $0.7 | $3.1 | 203K |
| 30 | GPT-5.6 Sol ≈ | OpenAI | 1,453.21 1,424–1,482 | $5 | $30 | 1,050K |
| 31 | Qwen3.6 Plus ≈ | Alibaba (Qwen) | 1,450.87 1,439–1,463 | $0.5 | $3 | 1,000K |
| 32 | Gemini 2.5 Pro ≈ | Google DeepMind | 1,450.58 1,444–1,458 | $1.25 | $10 | 1,049K |
| 33 | GPT-5.6 Luna ≈ | OpenAI | 1,450.15 1,421–1,479 | $0.2 | $1.2 | 1,050K |
| 34 | Qwen3 Max Preview ≈ | Alibaba (Qwen) | 1,449.83 1,435–1,465 | $1.2 | $6 | 256K |
| 35 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 1,448.89 1,438–1,459 | $0.6 | $3.6 | 262K |
| 36 | MiMo V2 Pro ≈ | Xiaomi | 1,447.85 1,433–1,463 | $0.44 | $0.87 | 1,049K |
| 37 | DeepSeek V4 Pro ≈ | DeepSeek | 1,442.56 1,430–1,455 | $0.44 | $0.87 | 1,049K |
| 38 | MiniMax M3 ≈ | MiniMax | 1,442.28 1,426–1,458 | $0.3 | $1.2 | 1,049K |
| 39 | Qwen3 Next 80B A3B Instruct ≈ | Alibaba (Qwen) | 1,442.26 1,426–1,459 | $0.15 | $1.2 | 262K |
| 40 | Longcat Flash Chat ≈ | Meituan | 1,440.87 1,419–1,463 | — | — | 131K |
| 41 | MiMo V2.5 ≈ | Xiaomi | 1,439.64 1,427–1,452 | $0.14 | $0.28 | 1,050K |
| 42 | Grok 4.20 Multi Agent Beta 0309 ≈ | xAI | 1,439.32 1,428–1,450 | $2 | $6 | 2,000K |
| 43 | Seed 2.0 Pro ≈ | ByteDance (Doubao) | 1,438.75 1,429–1,449 | $0.5 | $3 | 131K |
| 44 | Qwen3 Max 2025-09-23 ≈ | Alibaba (Qwen) | 1,438.73 1,415–1,462 | $0.86 | $3.43 | 258K |
| 45 | GLM 5 ≈ | Zhipu AI / Z.ai | 1,438.25 1,424–1,453 | $0.48 | $1.9 | 205K |
| 46 | DeepSeek V3.2 ≈ | DeepSeek | 1,436.18 1,425–1,447 | $0.28 | $0.4 | 164K |
| 47 | GPT-5.2 Chat ≈ | OpenAI | 1,434.04 1,421–1,447 | $1.75 | $14 | 128K |
| 48 | Mistral Medium 3.5 ≈ | Mistral AI | 1,433.51 1,408–1,459 | $1.5 | $7.5 | 262K |
| 49 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 1,433.27 1,419–1,448 | $0.3 | $2.4 | 262K |
| 50 | GLM 4.6 ≈ | Zhipu AI / Z.ai | 1,432.57 1,420–1,445 | $0.29 | $1.14 | 205K |
| 51 | MiMo V2 Omni ≈ | Xiaomi | 1,432.03 1,412–1,452 | $0.14 | $0.28 | 262K |
| 52 | Qwen3 235B A22B Instruct 2507 ≈ | Alibaba (Qwen) | 1,431.34 1,424–1,439 | $0.1 | $0.1 | 262K |
| 53 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 1,431.12 1,421–1,442 | $0.25 | $1.5 | 1,049K |
| 54 | Kimi K2 Thinking Turbo ≈ | Moonshot AI (Kimi) | 1,429.15 1,419–1,439 | $1.15 | $8 | 262K |
| 55 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 1,428.55 1,415–1,443 | $0.4 | $3.2 | 262K |
| 56 | DeepSeek V4 Flash ≈ | DeepSeek | 1,427.56 1,415–1,440 | $0.14 | $0.28 | 1,049K |
| 57 | GLM 4.5 ≈ | Zhipu AI / Z.ai | 1,427.09 1,412–1,443 | $0.2 | $0.8 | 131K |
| 58 | Qwen3 VL 235B A22B Instruct ≈ | Alibaba (Qwen) | 1,425.84 1,403–1,448 | $0.4 | $1.6 | 262K |
| 59 | o3 2025-04-16 ≈ | OpenAI | 1,424.48 1,415–1,434 | $2 | $8 | 200K |
| 60 | Grok 4 ≈ | xAI | 1,424.24 1,412–1,437 | $3 | $15 | 256K |
| 61 | DeepSeek V3.2 Exp ≈ | DeepSeek | 1,424.1 1,404–1,445 | $0.22 | $0.33 | 164K |
| 62 | GLM 4.7 ≈ | Zhipu AI / Z.ai | 1,423.31 1,402–1,444 | $0.15 | $0.8 | 205K |
| 63 | GPT-5.5 Instant ≈ | OpenAI | 1,423.07 1,407–1,439 | $5 | $30 | 400K |
| 64 | Claude 4.1 Opus ≈ | Anthropic | 1,421.84 1,413–1,430 | $15 | $75 | 200K |
| 65 | Claude Sonnet 4.5 ≈ | Anthropic | 1,421.66 1,413–1,430 | $3 | $15 | 1,000K |
| 66 | DeepSeek V3.1 ≈ | DeepSeek | 1,421.33 1,403–1,440 | $0.2 | $0.7 | 164K |
| 67 | MiniMax M2.7 ≈ | MiniMax | 1,420.2 1,408–1,432 | $0.3 | $1.2 | 205K |
| 68 | Grok 4.1 ≈ | xAI | 1,417.5 1,408–1,427 | $2 | $10 | 200K |
| 69 | 2025) ≈ | Google DeepMind | 1,416.44 1,403–1,430 | $0.3 | $2.5 | 1,049K |
| 70 | Mistral Large 3 ≈ | Mistral AI | 1,415.23 1,405–1,426 | $0.5 | $1.5 | 256K |
| 71 | Qwen3 VL 235B A22B Thinking ≈ | Alibaba (Qwen) | 1,414.82 1,387–1,443 | $0.4 | $4 | 262K |
| 72 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 1,414.68 1,378–1,451 | $0.3 | $2.5 | 1,049K |
| 73 | Qwen3 235B A22B Thinking-2507 ≈ | Alibaba (Qwen) | 1,412.99 1,389–1,437 | $0.1 | $0.1 | 262K |
| 74 | GPT 4.5 Preview ≈ | OpenAI | 1,411.89 1,397–1,427 | $75 | $150 | 128K |
| 75 | Mistral Medium 3.1 ≈ | Mistral AI | 1,410.6 1,403–1,418 | $0.4 | $2 | 262K |
| 76 | Gemini 2.5 Flash ≈ | Google DeepMind | 1,410.32 1,404–1,417 | $0.3 | $2.5 | 1,049K |
| 77 | Hunyuan T1 ≈ | Tencent (Hunyuan) | 1,409.68 1,372–1,448 | — | — | 131K |
| 78 | Qwen3.5 Flash ≈ | Alibaba (Qwen) | 1,409.11 1,398–1,420 | $0.09 | $0.36 | 1,000K |
| 79 | GPT 5 Chat ≈ | OpenAI | 1,408.99 1,395–1,423 | $1.25 | $10 | 128K |
| 80 | Step 3.5 Flash ≈ | StepFun | 1,406.89 1,396–1,418 | $0.1 | $0.3 | 262K |
| 81 | ChatGPT 4o Latest ≈ | OpenAI | 1,405.81 1,398–1,414 | $5 | $15 | 128K |
| 82 | Grok 4.1 Fast Reasoning ≈ | xAI | 1,405.28 1,395–1,416 | $0.2 | $0.5 | 2,000K |
| 83 | Qwen3.5 35B A3B ≈ | Alibaba (Qwen) | 1,405.11 1,391–1,419 | $0.25 | $2 | 262K |
| 84 | DeepSeek R1 0528 ≈ | DeepSeek | 1,403.81 1,384–1,424 | $0.25 | $0.25 | 164K |
| 85 | Grok 4 Fast Reasoning ≈ | xAI | 1,403.28 1,386–1,421 | $0.2 | $0.5 | 2,000K |
| 86 | DeepSeek V3.1 Terminus ≈ | DeepSeek | 1,398.82 1,361–1,437 | $0.21 | $0.79 | 164K |
| 87 | Qwen3 32B ≈ | Alibaba (Qwen) | 1,398.27 1,368–1,428 | $0.7 | $2.8 | 131K |
| 88 | GLM 4.5 Air ≈ | Zhipu AI / Z.ai | 1,397.64 1,383–1,412 | $0.11 | $0.29 | 131K |
| 89 | Kimi K2 0905 Preview ≈ | Moonshot AI (Kimi) | 1,397.5 1,377–1,418 | $0.6 | $2.5 | 262K |
| 90 | MiMo V2 Flash ≈ | Xiaomi | 1,395.77 1,384–1,407 | $0.14 | $0.28 | 262K |
| 91 | Qwen3 Next 80B A3B Thinking ≈ | Alibaba (Qwen) | 1,395.51 1,376–1,415 | $0.15 | $1.2 | 262K |
| 92 | Qwen 3 235b A22B ≈ | Alibaba (Qwen) | 1,395.48 1,382–1,409 | $0.7 | $2.8 | 131K |
| 93 | Qwen3 30B A3B Instruct 2507 ≈ | Alibaba (Qwen) | 1,394.86 1,380–1,410 | $0.05 | $0.19 | 262K |
| 94 | Llama 3.3 Nemotron Super 49B v1.5 ≈ | Meta AI | 1,394.18 1,355–1,433 | $0.05 | $0.25 | 131K |
| 95 | Claude Haiku 4.5 ≈ | Anthropic | 1,392.77 1,385–1,401 | $1 | $5 | 200K |
| 96 | DeepSeek Reasoner ≈ | DeepSeek | 1,392.4 1,378–1,406 | $0.55 | $2.19 | 164K |
| 97 | Grok 4.3 ≈ | xAI | 1,390.9 1,378–1,404 | $1.25 | $2.5 | 1,000K |
| 98 | o1 2024-12-17 ≈ | OpenAI | 1,388.33 1,378–1,399 | $15 | $60 | 200K |
| 99 | GPT OSS 120B ≈ | OpenAI | 1,388.15 1,374–1,402 | $0.03 | $0.14 | 131K |
| 100 | o4 Mini 2025-04-16 ≈ | OpenAI | 1,386.94 1,376–1,398 | $1.1 | $4.4 | 200K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 99 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 28.2 points — that is about 18 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 1,661: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)