AI ranking

Best AI models for programming

How many programming tasks the model actually solved on fixed benchmark sets — a verifiable result, not an opinion. The collection has a second view, based on blind human ratings: the order there is different, and the gap between the two lists says more than either list on its own.

One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.

41 models in the table
41 considered in the calculation
24 July 2026 newest model released here
7 August 2026 last recalculated

Not every model is measured here. This table is missing 3 of the top ten from the “By human votes” tab — MiMo V2.5 Pro, Gemini 3.6 Flash, Muse Spark 1.1. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 24 July 2026.

# Model Developer Task score 0–100, adjusted for coverage Sign in$ / 1M Output$ / 1M Contexttokens Aider Polyglot Terminal-Bench SWE-bench Verified GSO-Bench FrontierCode Task setsmeasured
1 Claude Opus 4.7 Anthropic 59.79 64.22 $5 $25 1,000K 90.2 83.5 44.1 38.5 5
2 Claude Opus 4.6 Anthropic 59.43 66.57 $5 $25 1,000K 79.8 78.7 41.2 3
3 GPT 5.4 2026-03-05 OpenAI 57.49 63.34 $2.5 $15 1,050K 81.8 76.9 31.4 3
4 Gemini 3.5 Flash Google DeepMind 56.64 64.57 $1.5 $9 1,049K 79.3 2
5 Claude Fable 5 Anthropic 55.96 63.2 $10 $50 1,000K 53.5 2
6 GLM 5 Zhipu AI / Z.ai 55.48 62.24 $0.48 $1.9 205K 52.4 72.1 2
7 Kimi K2.6 Moonshot AI (Kimi) 55.42 62.13 $0.22 $1.14 262K 76.7 2
8 Claude Opus 5 Anthropic 55.21 61.7 $5 $25 1,000K 53.4 2
9 Claude Sonnet 4.6 Anthropic 55.01 59.2 $3 $15 1,000K 53.4 75.2 3
10 o3 2025-04-16 OpenAI 53.98 56.61 $2 $8 200K 81.3 62.3 8.8 4
11 o1 2024-12-17 OpenAI 53.78 58.85 $15 $60 200K 61.7 2
12 GPT-5.6 Sol OpenAI 53.03 57.35 $5 $30 1,050K 47.5 2
13 Claude 4.5 Opus Anthropic 52.74 55.42 $5 $25 200K 63.1 76.7 26.5 3
14 GPT 5 2025-08-07 OpenAI 52.58 54.51 $1.25 $10 272K 88 49.6 73.6 6.9 4
15 Claude 4.1 Opus Anthropic 52.2 55.68 $15 $75 200K 38 73.4 2
16 Gemini 3 Pro Preview Google DeepMind 51.67 53.64 $2 $12 1,049K 69.4 72.9 18.6 3
17 Grok 4 xAI 51.06 53.4 $3 $15 256K 79.6 27.2 2
18 GLM 5.2 Zhipu AI / Z.ai 51.05 52.6 $0.42 $1.32 1,049K 78.7 24.5 3
19 Claude Opus 4.8 Anthropic 50.96 52.45 $5 $25 1,000K 47.1 46.5 3
20 Gemini 3.1 Pro Preview Google DeepMind 50.05 51.38 $2 $12 1,049K 80.2 22.6 2
21 Claude 4 Opus Anthropic 49.4 49.85 $15 $75 200K 72 70.7 6.9 3
22 Gemini 3 Flash Preview Google DeepMind 49.39 49.84 $0.5 $3 1,049K 64.3 75.4 9.8 3
23 Kimi K2.5 Moonshot AI (Kimi) 49.26 49.62 $0.35 $1.7 262K 43.2 73.8 3
24 GPT 5 Mini 2025-08-07 OpenAI 49.23 49.74 $0.25 $2 272K 34.8 64.7 2
25 DeepSeek V4 Pro DeepSeek 48.17 47.62 $0.44 $0.87 1,049K 77.6 17.6 2
26 GPT 4.1 2025-04-14 OpenAI 48.07 47.65 $2 $8 1,048K 52.4 48.5 3
27 Claude Sonnet 5 Anthropic 47.72 47.05 $2 $10 1,000K 37.3 42.7 3
28 o4 Mini 2025-04-16 OpenAI 47.01 45.87 $1.1 $4.4 200K 72 3.6 3
29 Gemini 2.5 Pro Google DeepMind 46.9 45.08 $1.25 $10 1,049K 32.6 57.6 2
30 Claude Sonnet 3.7 Anthropic 46.85 45.91 $3 $15 200K 64.9 61 3.8 4
31 Gemini 2.5 Pro Preview 0605 Google DeepMind 46.11 43.51 $1.13 $9 1,049K 83.1 3.9 2
32 Claude Sonnet 4.5 Anthropic 45.98 44.16 $3 $15 1,000K 46.5 71.3 14.7 3
33 o3 Mini 2025-01-31 OpenAI 42.63 38.57 $1.1 $4.4 200K 60.4 1.3 3
34 Claude 4 Sonnet Anthropic 40.91 33.1 $3 $15 1,000K 61.3 4.9 2
35 Claude 3.5 Sonnet 2024-10-22 Anthropic 40.33 34.73 $3 $15 200K 51.6 4.6 3
36 GPT OSS 120B OpenAI 39.48 30.25 $0.03 $0.14 131K 41.8 18.7 2
37 Claude 3.5 Haiku Anthropic 39.36 30 $0.25 $1.25 200K 28 2
38 Kimi K2 Instruct Moonshot AI (Kimi) 37.85 30.6 $0.1 $2 131K 59.1 27.8 4.9 3
39 GPT-4o (2024-08-06) OpenAI 36.63 24.55 $2.5 $10 128K 23.1 2
40 OpenAI GPT-4.1 Mini OpenAI 36.46 24.2 $0.4 $1.6 1,048K 32.4 2
41 GPT-4o (2024-11-20) OpenAI 29.32 16.4 $2.5 $10 128K 18.2 31 0 3

The table scrolls sideways: not all columns fit.

The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero. Showing the five most complete task sets out of 7; the rest are on the model page.

What each task set checks
Aider Polyglot
code edits in six languages models measured: 43
CadEval
building CAD models with code models measured: 13
CursorBench
edits in the editor models measured: 12
FrontierCode
patches fit to be merged into a project models measured: 15
GSO-Bench
code optimisation models measured: 24
SWE-bench Verified
fixing bugs in repositories models measured: 31
SWE-bench Verified, console only
Terminal-Bench
working in the command line models measured: 35

The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

Epoch AI — license CC-BY · primary source
По данным Epoch AI