Here the model does not answer a question but does the work: edits files in an operating system, runs searches, completes tasks in a working environment. That is why the percentages are low even for strong models, and they cannot be compared with scores from the knowledge collection — it is different work, not different strength.
| # | Model | Developer | Task score 0–100, adjusted for coverage | Sign in$ / 1M | Output$ / 1M | Contexttokens | APEX-Agents | DeepResearch Bench | BALROG | Cybench | The Agent Company | Task setsmeasured |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 40.67 46.22 | $5 | $25 | 1,000K | 32.4 | 55.3 | — | 93 | — | 4 |
| 2 | Claude 4.5 Opus | Anthropic | 40.11 43.63 | $5 | $25 | 200K | 20.7 | — | 43.5 | 82 | — | 6 |
| 3 | Claude Sonnet 4.5 | Anthropic | 39.89 44.02 | $3 | $15 | 1,000K | — | 52.6 | — | 60 | — | 5 |
| 4 | Claude 4.1 Opus | Anthropic | 38.89 45.1 | $15 | $75 | 200K | — | 49.7 | — | 42 | — | 3 |
| 5 | Gemini 3.1 Pro Preview | Google DeepMind | 37.42 45.25 | $2 | $12 | 1,049K | 33.5 | — | 57 | — | — | 2 |
| 6 | Claude 4 Opus | Anthropic | 36.54 43.5 | $15 | $75 | 200K | — | 49 | — | 38 | — | 2 |
| 7 | Claude Sonnet 4.6 | Anthropic | 36.52 39.99 | $3 | $15 | 1,000K | 23.7 | 54.9 | — | — | — | 4 |
| 8 | Kimi K2.5 | Moonshot AI (Kimi) | 34.22 38.85 | $0.35 | $1.7 | 262K | 14.4 | — | — | — | — | 2 |
| 9 | Grok 4 | xAI | 32.85 34.16 | $3 | $15 | 256K | 15.2 | 47.9 | 43.6 | 43 | — | 5 |
| 10 | Gemini 3 Flash Preview | Google DeepMind | 32.82 36.05 | $0.5 | $3 | 1,049K | 24 | — | 48.1 | — | — | 2 |
| 11 | Claude 4 Sonnet | Anthropic | 32.61 33.82 | $3 | $15 | 1,000K | 9.3 | 47.8 | — | 35 | 33.1 | 5 |
| 12 | GPT 5.4 2026-03-05 | OpenAI | 32.56 35.55 | $2.5 | $15 | 1,050K | 36 | 35.1 | — | — | — | 2 |
| 13 | Gemini 3 Pro Preview | Google DeepMind | 31.72 32.79 | $2 | $12 | 1,049K | 31.5 | — | 58.1 | — | — | 4 |
| 14 | Claude Sonnet 3.7 | Anthropic | 31.58 32.58 | $3 | $15 | 200K | — | 43.6 | — | 20 | 30.9 | 4 |
| 15 | Gemini 2.5 Pro Preview 0506 | Google DeepMind | 30.34 31.1 | $1.25 | $10 | 1,049K | — | 31.9 | — | — | 30.3 | 2 |
| 16 | Claude Opus 4.8 | Anthropic | 30.14 30.42 | $5 | $25 | 1,000K | 42.5 | 50.2 | — | — | — | 4 |
| 17 | Claude Fable 5 | Anthropic | 30.07 30.55 | $10 | $50 | 1,000K | 45 | — | — | — | — | 2 |
| 18 | GPT 5.4 Mini 2026-03-17 | OpenAI | 30.01 30.43 | $0.75 | $4.5 | 272K | 24.6 | 36.3 | — | — | — | 2 |
| 19 | o3 2025-04-16 | OpenAI | 29.46 29.4 | $2 | $8 | 200K | 17.2 | 46.6 | — | — | — | 4 |
| 20 | Claude 3.5 Sonnet 2024-10-22 | Anthropic | 28.94 28.3 | $3 | $15 | 200K | — | — | 32.6 | — | 24 | 2 |
| 21 | GPT 5 2025-08-07 | OpenAI | 28.84 28.54 | $1.25 | $10 | 272K | 18.3 | 55.1 | 32.8 | — | — | 5 |
| 22 | Claude Opus 4.7 | Anthropic | 27.82 26.05 | $5 | $25 | 1,000K | 33.9 | — | — | — | — | 2 |
| 23 | Gemini 2.5 Pro | Google DeepMind | 27.75 26.54 | $1.25 | $10 | 1,049K | 6.6 | 49.7 | — | — | — | 3 |
| 24 | Gemini 3.1 Flash Lite | Google DeepMind | 27.14 24.7 | $0.25 | $1.5 | 1,049K | 13 | 36.4 | — | — | — | 2 |
| 25 | Claude Haiku 4.5 | Anthropic | 24.82 20.05 | $1 | $5 | 200K | 8.9 | — | 31.2 | — | — | 2 |
| 26 | Gemini 2.5 Flash | Google DeepMind | 23.62 17.65 | $0.3 | $2.5 | 1,049K | 1.8 | — | 33.5 | — | — | 2 |
| 27 | Llama 3.1 70B | Meta AI | 23.49 17.4 | $0.12 | $0.3 | 131K | — | — | 27.9 | — | 6.9 | 2 |
| 28 | Grok 3 | xAI | 22.69 15.8 | $3 | $15 | 131K | 2.1 | — | 29.5 | — | — | 2 |
| 29 | Kimi K2.6 | Moonshot AI (Kimi) | 20.67 11.75 | $0.22 | $1.14 | 262K | 18.9 | — | — | — | — | 2 |
| 30 | Qwen2.5 72B Instruct | Alibaba (Qwen) | 20.27 10.95 | $1.4 | $5.6 | 131K | — | — | 16.2 | — | 5.7 | 2 |
| 31 | Llama 3.1 405B Instruct | Meta AI | 18.52 7.45 | $0.12 | $0.3 | 128K | — | — | — | 7.5 | 7.4 | 2 |
| 32 | GPT-4o (2024-11-20) | OpenAI | 15.21 8.03 | $2.5 | $10 | 128K | 1.1 | — | — | 12.5 | 8.6 | 4 |
The table scrolls sideways: not all columns fit.
The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero. Showing the five most complete task sets out of 9; the rest are on the model page.
The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
Epoch AI
— license CC-BY · primary source
По данным Epoch AI