AI ranking

Best AI models for agent tasks

Here the model does not answer a question but does the work: edits files in an operating system, runs searches, completes tasks in a working environment. That is why the percentages are low even for strong models, and they cannot be compared with scores from the knowledge collection — it is different work, not different strength.

32 models in the table
32 considered in the calculation
9 June 2026 newest model released here
7 August 2026 last recalculated
# Model Developer Task score 0–100, adjusted for coverage Sign in$ / 1M Output$ / 1M Contexttokens APEX-Agents DeepResearch Bench BALROG Cybench The Agent Company Task setsmeasured
1 Claude Opus 4.6 Anthropic 40.67 46.22 $5 $25 1,000K 32.4 55.3 93 4
2 Claude 4.5 Opus Anthropic 40.11 43.63 $5 $25 200K 20.7 43.5 82 6
3 Claude Sonnet 4.5 Anthropic 39.89 44.02 $3 $15 1,000K 52.6 60 5
4 Claude 4.1 Opus Anthropic 38.89 45.1 $15 $75 200K 49.7 42 3
5 Gemini 3.1 Pro Preview Google DeepMind 37.42 45.25 $2 $12 1,049K 33.5 57 2
6 Claude 4 Opus Anthropic 36.54 43.5 $15 $75 200K 49 38 2
7 Claude Sonnet 4.6 Anthropic 36.52 39.99 $3 $15 1,000K 23.7 54.9 4
8 Kimi K2.5 Moonshot AI (Kimi) 34.22 38.85 $0.35 $1.7 262K 14.4 2
9 Grok 4 xAI 32.85 34.16 $3 $15 256K 15.2 47.9 43.6 43 5
10 Gemini 3 Flash Preview Google DeepMind 32.82 36.05 $0.5 $3 1,049K 24 48.1 2
11 Claude 4 Sonnet Anthropic 32.61 33.82 $3 $15 1,000K 9.3 47.8 35 33.1 5
12 GPT 5.4 2026-03-05 OpenAI 32.56 35.55 $2.5 $15 1,050K 36 35.1 2
13 Gemini 3 Pro Preview Google DeepMind 31.72 32.79 $2 $12 1,049K 31.5 58.1 4
14 Claude Sonnet 3.7 Anthropic 31.58 32.58 $3 $15 200K 43.6 20 30.9 4
15 Gemini 2.5 Pro Preview 0506 Google DeepMind 30.34 31.1 $1.25 $10 1,049K 31.9 30.3 2
16 Claude Opus 4.8 Anthropic 30.14 30.42 $5 $25 1,000K 42.5 50.2 4
17 Claude Fable 5 Anthropic 30.07 30.55 $10 $50 1,000K 45 2
18 GPT 5.4 Mini 2026-03-17 OpenAI 30.01 30.43 $0.75 $4.5 272K 24.6 36.3 2
19 o3 2025-04-16 OpenAI 29.46 29.4 $2 $8 200K 17.2 46.6 4
20 Claude 3.5 Sonnet 2024-10-22 Anthropic 28.94 28.3 $3 $15 200K 32.6 24 2
21 GPT 5 2025-08-07 OpenAI 28.84 28.54 $1.25 $10 272K 18.3 55.1 32.8 5
22 Claude Opus 4.7 Anthropic 27.82 26.05 $5 $25 1,000K 33.9 2
23 Gemini 2.5 Pro Google DeepMind 27.75 26.54 $1.25 $10 1,049K 6.6 49.7 3
24 Gemini 3.1 Flash Lite Google DeepMind 27.14 24.7 $0.25 $1.5 1,049K 13 36.4 2
25 Claude Haiku 4.5 Anthropic 24.82 20.05 $1 $5 200K 8.9 31.2 2
26 Gemini 2.5 Flash Google DeepMind 23.62 17.65 $0.3 $2.5 1,049K 1.8 33.5 2
27 Llama 3.1 70B Meta AI 23.49 17.4 $0.12 $0.3 131K 27.9 6.9 2
28 Grok 3 xAI 22.69 15.8 $3 $15 131K 2.1 29.5 2
29 Kimi K2.6 Moonshot AI (Kimi) 20.67 11.75 $0.22 $1.14 262K 18.9 2
30 Qwen2.5 72B Instruct Alibaba (Qwen) 20.27 10.95 $1.4 $5.6 131K 16.2 5.7 2
31 Llama 3.1 405B Instruct Meta AI 18.52 7.45 $0.12 $0.3 128K 7.5 7.4 2
32 GPT-4o (2024-11-20) OpenAI 15.21 8.03 $2.5 $10 128K 1.1 12.5 8.6 4

The table scrolls sideways: not all columns fit.

The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero. Showing the five most complete task sets out of 9; the rest are on the model page.

What each task set checks
APEX-Agents
models measured: 47
BALROG
game environments models measured: 22
Cybench
cybersecurity tasks models measured: 19
DeepResearch Bench
deep research models measured: 22
GDPval
tasks from real occupations models measured: 11
OSWorld
working inside an operating system, first version models measured: 8
OSWorld 2.0
working inside an operating system models measured: 6
Remote Labor Index
jobs from a freelance marketplace models measured: 10
The Agent Company
work tasks in an office environment models measured: 13

The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

Epoch AI — license CC-BY · primary source
По данным Epoch AI