AI ranking

Best AI models by breadth of knowledge

A test of factual knowledge on questions that have a single correct answer. It differs from the neighbouring collections in that the model does not reason or write code here — it recalls. The second view on the same topic comes from blind human ratings.

One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.

79 models in the table
79 considered in the calculation
27 July 2026 newest model released here
7 August 2026 last recalculated

Not every model is measured here. This table is missing 1 of the top ten from the “By human votes” tab — MiMo V2.5 Pro. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 27 July 2026.

# Model Developer Task score 0–100, adjusted for coverage Sign in$ / 1M Output$ / 1M Contexttokens GPQA Diamond SimpleQA Verified CritPt Humanity’s Last Exam Task setsmeasured
1 Gemini 3 Flash Preview Google DeepMind 55.84 72.51 $0.5 $3 1,049K 77.6 67.4 2
2 Qwen3.6 Max Preview Alibaba (Qwen) 55.18 71.19 $1.3 $7.8 262K 85.4 56.9 2
3 GPT-5.6 Sol OpenAI 54.72 65.08 $5 $30 1,050K 91.3 71.6 32.3 3
4 Qwen3 Max 2025-09-23 Alibaba (Qwen) 52.32 65.47 $0.86 $3.43 258K 63.5 67.5 2
5 Grok 4 xAI 52.23 65.29 $3 $15 256K 82.7 47.9 2
6 Gemini 3.1 Pro Preview Google DeepMind 51.54 57.72 $2 $12 1,049K 92.1 77.3 17.7 43.7 4
7 Claude Opus 5 Anthropic 51.21 59.23 $5 $25 1,000K 91.8 56.7 29.1 3
8 Gemini 3.5 Flash Google DeepMind 50.06 57.31 $1.5 $9 1,049K 90.4 68.4 13.1 3
9 Grok 4.20 (Reasoning) xAI 49.81 60.44 $1.25 $2.5 1,000K 85.8 35.1 2
10 Claude Opus 4.6 Anthropic 48.67 55 $5 $25 1,000K 87.4 46.5 31.1 3
11 GPT-5.6 Terra OpenAI 48.51 54.73 $2 $12 1,050K 91.1 43.1 30 3
12 GPT 5.4 Pro 2026-03-05 OpenAI 48.41 53.03 $30 $180 1,050K 92.8 47.8 30 41.5 4
13 Qwen3.7 Max Alibaba (Qwen) 47.82 53.59 $2.5 $7.5 1,000K 88.8 58.5 13.4 3
14 Grok 4.5 xAI 47.71 53.39 $2 $6 500K 91.3 53.5 15.4 3
15 Kimi K2 Thinking Turbo Moonshot AI (Kimi) 47.23 55.28 $1.15 $8 262K 79 31.6 2
16 Gemini 3 Pro Preview Google DeepMind 47.11 51.08 $2 $12 1,049K 90.2 72.9 6.9 34.4 4
17 Kimi K3 Moonshot AI (Kimi) 47.06 52.31 $2 $8 1,049K 90.8 42.7 23.4 3
18 GLM 4.7 Zhipu AI / Z.ai 46.91 54.64 $0.15 $0.8 205K 77.8 31.5 2
19 DeepSeek V4 Pro DeepSeek 46.88 52.02 $0.44 $0.87 1,049K 86.2 57 12.9 3
20 DeepSeek Reasoner DeepSeek 45.94 52.7 $0.14 $0.28 1,000K 77.9 27.5 2
21 GPT-5.6 Luna OpenAI 45.89 50.37 $0.2 $1.2 1,050K 88.8 41.7 20.6 3
22 Qwen3.5 Plus Alibaba (Qwen) 45.81 52.44 $0.4 $2.4 1,000K 78.9 26 2
23 Claude Opus 4.8 Anthropic 45.35 49.47 $5 $25 1,000K 88.1 39.5 20.9 3
24 GLM 5.2 Zhipu AI / Z.ai 45.29 49.37 $0.42 $1.32 1,049K 89.1 38.1 20.9 3
25 GPT 5.4 2026-03-05 OpenAI 45.12 48.09 $2.5 $15 1,050K 91.1 44.8 23.4 33 4
26 Qwen3.6 Flash Alibaba (Qwen) 44.69 50.2 $0.19 $1.13 1,000K 79.2 21.2 2
27 Claude 4.5 Opus Anthropic 44.6 48.21 $5 $25 200K 81.4 41.8 21.4 3
28 Qwen3.5 Flash Alibaba (Qwen) 44.14 49.11 $0.09 $0.36 1,000K 78.5 19.8 2
29 Claude Fable 5 Anthropic 43.81 48.44 $10 $50 1,000K 68.3 28.6 2
30 DeepSeek R1 0528 DeepSeek 43.55 47.92 $0.25 $0.25 164K 68.4 27.4 2
31 Claude Opus 4.7 Anthropic 43.47 45.61 $5 $25 1,000K 86.9 50.6 12 33 4
32 Gemini 2.5 Pro Google DeepMind 43.35 46.13 $1.25 $10 1,049K 80.4 56 2 3
33 Kimi K2.5 Moonshot AI (Kimi) 43.26 45.98 $0.35 $1.7 262K 83.5 33.9 20.6 3
34 Kimi K2.7 Code Moonshot AI (Kimi) 42.72 45.08 $0.28 $1.1 262K 86 39.2 10 3
35 Gemini 2.5 Pro Experimental 0325 Google DeepMind 42.71 46.24 $2.5 $10 1,049K 78.5 14 2
36 Qwen3.6 Plus Alibaba (Qwen) 42.7 45.04 $0.5 $3 1,000K 83.2 49.1 2.9 3
37 Kimi K2.6 Moonshot AI (Kimi) 42.55 44.8 $0.22 $1.14 262K 87.7 38.7 8 3
38 Grok 3 Mini xAI 41.95 44.73 $0.3 $0.5 131K 68.4 21.1 2
39 Grok 4.3 xAI 41.89 43.7 $1.25 $2.5 1,000K 85.1 38 8 3
40 Claude Sonnet 5 Anthropic 41.52 43.08 $2 $10 1,000K 87.4 25 16.9 3
41 GPT 5 2025-08-07 OpenAI 40.78 41.58 $1.25 $10 272K 81.6 50.6 12.6 21.6 4
42 Qwen3 235B A22B Thinking-2507 Alibaba (Qwen) 40.37 41.17 $0.1 $0.1 262K 73.4 50.1 0 3
43 GLM 5.1 Zhipu AI / Z.ai 40.17 40.83 $0.06 $0.22 205K 80.6 37.3 4.6 3
44 GPT 5.4 Mini 2026-03-17 OpenAI 39.01 38.9 $0.75 $4.5 272K 78.1 28.6 10 3
45 Claude Sonnet 4.6 Anthropic 38.73 38.44 $3 $15 1,000K 83.2 29 3.1 3
46 Claude Sonnet 3.7 Anthropic 38.68 38.19 $3 $15 200K 73 3.4 2
47 Claude 4.1 Opus Anthropic 37.98 37.19 $15 $75 200K 69.7 34.8 7.1 3
48 o1 2024-12-17 OpenAI 37.67 36.17 $15 $60 200K 69 3.3 2
49 o3 2025-04-16 OpenAI 37.47 36.62 $2 $8 200K 75.8 53 1.4 16.3 4
50 GPT 5 Nano 2025-08-07 OpenAI 37.45 35.73 $0.05 $0.4 272K 59.3 12.2 2
51 Gemini 2.5 Pro Preview 0506 Google DeepMind 36.89 34.61 $1.25 $10 1,049K 55.6 13.7 2
52 DeepSeek Reasoner DeepSeek 35.44 31.7 $0.55 $2.19 164K 62.3 1.1 2
53 GPT 4.5 Preview OpenAI 34.32 29.46 $75 $150 128K 58.3 0.7 2
54 GPT 5.4 Nano 2026-03-17 OpenAI 34.18 30.85 $0.2 $1.25 272K 71.3 12 9.3 3
55 DeepSeek Chat 0324 DeepSeek 33.79 28.41 $0.2 $0.6 164K 56.8 0 2
56 GPT 4.1 2025-04-14 OpenAI 33.72 28.26 $2 $8 1,048K 55.9 0.6 2
57 OpenAI GPT-4.1 Mini OpenAI 33.2 27.23 $0.4 $1.6 1,048K 54.5 0 2
58 GPT OSS 120B OpenAI 32.21 27.57 $0.03 $0.14 131K 67.7 13.9 1.1 3
59 o4 Mini 2025-04-16 OpenAI 31.6 27.82 $1.1 $4.4 200K 72.8 23.9 0.6 14 4
60 Claude Sonnet 4.5 Anthropic 31.48 27.64 $3 $15 1,000K 76.4 23.6 1.1 9.4 4
61 Mistral Medium 3 Mistral AI 31.1 23.02 $0.4 $2 131K 46 0 2
62 Claude 4 Sonnet Anthropic 30.8 25.22 $3 $15 1,000K 72.3 0.3 3.1 3
63 Gemini 2.0 Flash Thinking 0121 Google DeepMind 30.74 22.31 $0.31 $1 1,000K 42.8 1.9 2
64 Claude 4 Opus Anthropic 30.64 24.96 $15 $75 200K 68.4 0.3 6.2 3
65 GPT 5 Mini 2025-08-07 OpenAI 30.23 25.76 $0.25 $2 272K 66.7 21 0 15.4 4
66 DeepSeek V3 DeepSeek 30.1 21.03 $0.27 $1.1 164K 42.1 0 2
67 Claude 3.5 Sonnet 2024-10-22 Anthropic 29.69 20.2 $3 $15 200K 40.4 0 2
68 Claude Haiku 4.5 Anthropic 29.17 22.51 $1 $5 200K 61.6 5.9 0 3
69 Llama 4 Scout 17B 16E Instruct Meta AI 28.53 17.89 $0.05 $0.1 10,000K 35.8 0 2
70 GPT 4.1 Nano 2025-04-14 OpenAI 27.56 15.95 $0.1 $0.4 1,048K 31.9 0 2
71 Gemma 3 27B Google DeepMind 27.54 15.91 $0.03 $0.11 131K 31.8 0 2
72 Mistral Small 3.1 Mistral AI 27.08 14.99 $0.1 $0.3 128K 30 0 2
73 Llama 3.3 70B Instruct Meta AI 27.07 14.96 $0.05 $0.23 131K 29.9 0 2
74 Llama 4 Maverick 17b 128e Instruct Meta AI 27.05 18.97 $0.05 $0.1 1,049K 56 0 0.9 3
75 Gemma 4 31B Google DeepMind 22.34 5.5 $0.1 $0.3 262K 9.6 1.4 2
76 GPT-4o (2024-11-20) OpenAI 21.77 10.17 $2.5 $10 128K 30.5 0 0 3
77 Gemini 3.1 Flash Lite Google DeepMind 20.88 2.59 $0.25 $1.5 1,049K 1.1 4 2
78 Claude 3.5 Haiku Anthropic 20.51 8.07 $0.25 $1.25 200K 17.5 6.7 0 3
79 Llama 3.1 8B Meta AI 19.9 0.63 $0.02 $0.03 131K 1.3 0 2

The table scrolls sideways: not all columns fit.

The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero.

What each task set checks
CritPt
physics problems models measured: 63
GPQA Diamond
graduate-level questions models measured: 122
Humanity’s Last Exam
expert-level questions models measured: 35
SimpleQA Verified
factual accuracy models measured: 57

The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

Epoch AI — license CC-BY · primary source
По данным Epoch AI