AI ranking

Best AI models for mathematics

The share of solved mathematical problems, from olympiad to research level. The percentages are lower than in other collections, and that is a property of the problems, not of the models: comparing them with the numbers from the knowledge collection is meaningless. The second view comes from blind human ratings.

One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.

88 models in the table
88 considered in the calculation
27 July 2026 newest model released here
7 August 2026 last recalculated

Not every model is measured here. This table is missing 2 of the top ten from the “By human votes” tab — Gemini 3.6 Flash, ERNIE 5.1. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 27 July 2026.

# Model Developer Task score 0–100, adjusted for coverage Sign in$ / 1M Output$ / 1M Contexttokens Mock AIME 2024–2025 FrontierMath, level 4 FrontierMath, levels 1–3 MATH, difficulty level five Task setsmeasured
1 Claude Fable 5 Anthropic 74.77 91.52 $10 $50 1,000K 99.7 87.8 87 3
2 GPT-5.6 Sol OpenAI 72.1 94.56 $5 $30 1,050K 100 89.1 2
3 Claude Opus 5 Anthropic 71.39 85.89 $5 $25 1,000K 98.9 73.2 85.6 3
4 GPT-5.6 Terra OpenAI 71.14 85.47 $2 $12 1,050K 99.7 70.7 86 3
5 o3 2025-04-16 OpenAI 70.23 90.82 $2 $8 200K 83.9 97.8 2
6 GPT-5.6 Luna OpenAI 68.14 80.47 $0.2 $1.2 1,050K 98.3 61 82.1 3
7 Qwen3 Max 2025-09-23 Alibaba (Qwen) 67.43 85.22 $0.86 $3.43 258K 73.3 97.1 2
8 Grok 3 Mini xAI 67 84.35 $0.3 $0.5 131K 77.8 90.9 2
9 o1 2024-12-17 OpenAI 66.83 84.01 $15 $60 200K 73.3 94.7 2
10 Claude Opus 4.8 Anthropic 66.74 78.14 $5 $25 1,000K 98.3 56.1 80 3
11 Claude Haiku 4.5 Anthropic 65.57 81.5 $1 $5 200K 66.6 96.4 2
12 DeepSeek R1 0528 DeepSeek 65.57 81.5 $0.25 $0.25 164K 66.4 96.6 2
13 GPT 5.4 2026-03-05 OpenAI 64.93 75.13 $2.5 $15 1,050K 97.8 49 78.6 3
14 Claude 4 Sonnet Anthropic 63.69 77.73 $3 $15 1,000K 71.1 84.4 2
15 Claude 4 Opus Anthropic 62.19 74.73 $15 $75 200K 64.4 85.1 2
16 Claude Sonnet 3.7 Anthropic 62.05 74.45 $3 $15 200K 57.7 91.2 2
17 Kimi K3 Moonshot AI (Kimi) 61.54 69.47 $2 $8 1,049K 97.2 39 72.2 3
18 DeepSeek Reasoner DeepSeek 61.41 73.17 $0.55 $2.19 164K 53.3 93.1 2
19 GPT 5 2025-08-07 OpenAI 61.03 66.73 $1.25 $10 272K 91.4 22 55.4 98.1 4
20 Grok 3 xAI 60.89 72.13 $3 $15 131K 55.5 88.8 2
21 GPT 5.4 Pro 2026-03-05 OpenAI 60.07 70.5 $30 $180 1,050K 58.5 82.5 2
22 Claude Opus 4.7 Anthropic 59.8 66.56 $5 $25 1,000K 97.8 31.7 70.2 3
23 o1 Mini 2024-09-12 OpenAI 58.84 68.04 $1.1 $4.4 128K 46.9 89.2 2
24 Qwen3.7 Max Alibaba (Qwen) 58.6 64.57 $2.5 $7.5 1,000K 95 34.2 64.6 3
25 OpenAI GPT-4.1 Mini OpenAI 57.81 65.98 $0.4 $1.6 1,048K 44.7 87.3 2
26 Claude Sonnet 5 Anthropic 57.78 63.2 $2 $10 1,000K 94.7 29.3 65.6 3
27 Claude Opus 4.6 Anthropic 57.31 62.41 $5 $25 1,000K 94.4 26.8 66 3
28 GPT 5 Mini 2025-08-07 OpenAI 57.11 60.84 $0.25 $2 272K 86.7 12.2 46.7 97.9 4
29 Gemini 3.5 Flash Google DeepMind 56.9 61.73 $1.5 $9 1,049K 95.6 26.8 62.8 3
30 Gemini 3.1 Pro Preview Google DeepMind 56.27 60.69 $2 $12 1,049K 95.6 26.8 59.7 3
31 Grok 4.5 xAI 55.73 59.79 $2 $6 500K 97.8 24.4 57.2 3
32 Kimi K2.6 Moonshot AI (Kimi) 55.65 59.65 $0.22 $1.14 262K 96.1 25.6 57.2 3
33 GPT 4.1 2025-04-14 OpenAI 55.14 60.64 $2 $8 1,048K 38.3 83 2
34 GLM 5.2 Zhipu AI / Z.ai 54.83 58.29 $0.42 $1.32 1,049K 86.4 29.3 59.2 3
35 GPT 5.2 Pro 2025-12-11 OpenAI 54.82 60 $21 $168 272K 46 74 2
36 GPT 4.5 Preview OpenAI 53.91 58.18 $75 $150 128K 37.7 78.6 2
37 o4 Mini 2025-04-16 OpenAI 53.3 55.13 $1.1 $4.4 200K 81.7 4.9 36.1 97.8 4
38 Mistral Medium 3 Mistral AI 53.27 56.89 $0.4 $2 131K 32.2 81.6 2
39 DeepSeek Chat 0324 DeepSeek 53.14 56.64 $0.2 $0.6 164K 37.7 75.6 2
40 o1 Preview 2024-09-12 OpenAI 53 56.35 $15 $60 128K 31 81.7 2
41 Kimi K2.7 Code Moonshot AI (Kimi) 52.38 54.21 $0.28 $1.1 262K 96.4 12.2 54 3
42 Gemini 3 Flash Preview Google DeepMind 52.07 53.69 $0.5 $3 1,049K 92.8 17.1 51.2 3
43 Grok 4.20 (Reasoning) xAI 50.7 51.4 $1.25 $2.5 1,000K 92.2 17.1 44.9 3
44 Claude Sonnet 4.5 Anthropic 50.18 50.45 $3 $15 1,000K 77.8 2.4 23.9 97.7 4
45 Grok 4.3 xAI 50.01 50.26 $1.25 $2.5 1,000K 93.3 14.6 42.8 3
46 GPT 5 Nano 2025-08-07 OpenAI 49.68 49.69 $0.05 $0.4 272K 81.1 2.4 20 95.2 4
47 GPT 4.1 Nano 2025-04-14 OpenAI 49.53 49.41 $0.1 $0.4 1,048K 28.8 70 2
48 GPT 5.4 Mini 2026-03-17 OpenAI 49.5 49.4 $0.75 $4.5 272K 87.2 9.8 51.2 3
49 GPT 5.4 Nano 2026-03-17 OpenAI 48.83 48.29 $0.2 $1.25 272K 87.8 12.2 44.9 3
50 DeepSeek V4 Pro DeepSeek 48.73 48.12 $0.44 $0.87 1,049K 96.7 2.4 45.3 3
51 o3 Mini 2025-01-31 OpenAI 48.55 48 $1.1 $4.4 200K 76.9 0 18.6 96.5 4
52 Gemma 3 27B Google DeepMind 48.24 46.84 $0.03 $0.11 131K 19.6 74 2
53 Llama 4 Maverick 17b 128e Instruct Meta AI 48.2 46.75 $0.05 $0.1 1,049K 20.5 73 2
54 Qwen2.5 Max 2025-01-25 Alibaba (Qwen) 45.63 41.61 128K 16 67.2 2
55 DeepSeek V3 DeepSeek 44.97 40.3 $0.27 $1.1 164K 15.8 64.9 2
56 Claude 4.5 Opus Anthropic 44.93 41.79 $5 $25 200K 86.1 4.9 34.4 3
57 Phi 4 Microsoft 44.47 39.3 $0.06 $0.14 128K 13.7 64.9 2
58 GPT 5 Pro 2025-10-06 OpenAI 43.65 37.65 $15 $120 400K 19.5 55.8 2
59 Grok 2 1212 xAI 43.56 37.48 $2 $10 131K 11.4 63.5 2
60 Qwen2.5 72B Instruct Alibaba (Qwen) 42.61 35.57 $1.4 $5.6 131K 8 63.2 2
61 Llama 4 Scout 17B 16E Instruct Meta AI 42.31 34.98 $0.05 $0.1 10,000K 7.7 62.3 2
62 Gemini 2.5 Pro Google DeepMind 41.71 36.42 $1.25 $10 1,049K 84.7 0 24.6 3
63 Claude 3.5 Sonnet 2024-10-22 Anthropic 41.16 32.67 $3 $15 200K 8.4 57 2
64 GPT-4o (2024-08-06) OpenAI 39.72 29.79 $2.5 $10 128K 6.3 53.3 2
65 OpenAI: GPT-4o-mini (2024-07-18) OpenAI 39.69 29.74 $0.15 $0.6 128K 6.9 52.6 2
66 Llama 3.1 405B Instruct Meta AI 39.67 29.7 $0.12 $0.3 128K 9.6 49.8 2
67 Claude Sonnet 3.5 Anthropic 39.35 29.06 $2.6 $13 200K 6.4 51.7 2
68 Mistral Large Mistral AI 39.32 28.99 $2 $6 131K 7.7 50.3 2
69 GPT-4o (2024-05-13) OpenAI 39.13 28.61 $5 $15 128K 6.2 51.1 2
70 GPT-4o (2024-11-20) OpenAI 38.81 27.97 $2.5 $10 128K 6.2 49.8 2
71 GPT 4 Turbo 2024-04-09 OpenAI 38.15 26.65 $10 $30 128K 6.6 46.7 2
72 Mistral Large 2407 Mistral AI 38.12 26.6 $3 $9 131K 8.4 44.8 2
73 Mistral Small 3.1 Mistral AI 37.95 26.26 $0.1 $0.3 128K 5.7 46.8 2
74 Claude 3.5 Haiku Anthropic 37.47 25.29 $0.25 $1.25 200K 4.2 46.4 2
75 Claude 4.1 Opus Anthropic 36.64 27.98 $15 $75 200K 68.9 2.4 12.6 3
76 Llama 3.3 70B Instruct Meta AI 36.48 23.32 $0.05 $0.23 131K 5 41.6 2
77 Claude 3 Opus 2024-02-29 Anthropic 35.35 21.06 $15 $75 200K 4.6 37.5 2
78 Llama 3.2 90B Vision Instruct Meta AI 35.32 20.99 $0.35 $0.4 128K 2.5 39.4 2
79 Llama 3.1 70B Meta AI 34.87 20.1 $0.12 $0.3 131K 3.5 36.7 2
80 Google: Gemma 2 27B Google DeepMind 32.12 14.59 $0.65 $0.65 8K 1.3 27.9 2
81 Llama 3 70B Instruct Meta AI 31.52 13.39 $0.12 $0.3 8K 4.2 22.6 2
82 Mistral Large 2402 Mistral AI 31.4 13.16 $4 $12 32K 1.9 24.5 2
83 Llama 3.1 8B Meta AI 31.14 12.64 $0.02 $0.03 131K 2.4 22.9 2
84 GPT 4 0613 OpenAI 30.82 11.99 $30 $60 8K 1 23 2
85 Claude 3 Sonnet Anthropic 29.97 10.29 $3 $15 200K 2.4 18.2 2
86 Anthropic: Claude 3 Haiku Anthropic 28.97 8.3 $0.25 $1.25 200K 1.7 14.9 2
87 Meta: Llama 3 8B Instruct Meta AI 26.54 3.43 $0.03 $0.04 8K 0.7 6.1 2
88 Llama 2 70B Chat HF Meta AI 25.65 1.65 $1 $1 4K 0 3.3 2

The table scrolls sideways: not all columns fit.

The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero.

What each task set checks
FrontierMath, level 4
research-grade problems models measured: 39
FrontierMath, levels 1–3
models measured: 39
MATH, difficulty level five
models measured: 72
Mock AIME 2024–2025
olympiad problems models measured: 110

The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

Epoch AI — license CC-BY · primary source
По данным Epoch AI