AI ranking

Best AI models for logic and reasoning

Reasoning tasks where the answer is either right or wrong: what counts is the share of solved problems, not of pleasing answers. Next to it there is a second view, from blind human ratings. Logic is where preference and a solved task part ways the most: a well-worded chain of reasoning pleases people even when its answer is wrong.

One question, two ways to answer it. The two metrics measure different things and do not add up into a single score.

91 models in the table
91 considered in the calculation
27 July 2026 newest model released here
7 August 2026 last recalculated

Not every model is measured here. This table is missing 2 of the top ten from the “By human votes” tab — Muse Spark 1.1, MiMo V2.5 Pro. They have no score at all on this tab's task sets: an empty place means "not measured", not "performed badly". The newest model this task set has reached was released on 27 July 2026.

# Model Developer Task score 0–100, adjusted for coverage Sign in$ / 1M Output$ / 1M Contexttokens WeirdML SimpleBench ARC-AGI ARC-AGI-2 Chess puzzles Task setsmeasured
1 GPT-5.6 Sol OpenAI 71.53 92.92 $5 $30 1,050K 88.8 97.5 92.5 3
2 Claude Opus 5 Anthropic 65.91 79.15 $5 $25 1,000K 91.8 97.5 88.3 39 4
3 Gemini 3.1 Pro Preview Google DeepMind 64.89 75.07 $2 $12 1,049K 72.1 75.5 98 77.1 52.7 5
4 GPT 5.4 Pro 2026-03-05 OpenAI 62.79 72.13 $30 $180 1,050K 57.4 68.9 94.5 83.3 56.4 5
5 GPT-5.6 Terra OpenAI 61.12 69.79 $2 $12 1,050K 78.3 38.7 96.5 83.9 51.6 5
6 GPT 5.4 2026-03-05 OpenAI 60.88 71.6 $2.5 $15 1,050K 77.7 93.7 74 41.1 4
7 Gemini 3.5 Flash Google DeepMind 60.79 69.33 $1.5 $9 1,049K 62.6 72 92.5 72.1 47.4 5
8 GPT-5.6 Sol Pro OpenAI 59.3 72.53 $5 $30 1,050K 89.4 66 62.1 3
9 Claude Opus 4.8 Anthropic 59.24 67.16 $5 $25 1,000K 82.9 57.8 92.5 72.1 30.6 5
10 Claude Opus 4.7 Anthropic 58.06 65.51 $5 $25 1,000K 76.4 55.5 93.5 75.8 26.4 5
11 Claude Fable 5 Anthropic 57.41 69.38 $10 $50 1,000K 91.9 78.3 37.9 3
12 Grok 4.20 xAI 57.16 68.97 $1.25 $2.5 2,000K 52.3 89.5 65.1 3
13 Gemini 2.5 Pro Preview 0605 Google DeepMind 56.37 73.29 $1.13 $9 1,049K 54.9 2
14 Claude Opus 4.6 Anthropic 56.26 62.98 $5 $25 1,000K 78 61.1 94 69.2 12.7 5
15 GPT 5.2 Pro 2025-12-11 OpenAI 54.49 64.51 $21 $168 272K 48.9 90.5 54.2 3
16 Grok 4.5 xAI 51.68 56.58 $2 $6 500K 46.4 64 87.2 52.6 32.7 5
17 GPT-5.6 Luna OpenAI 51.47 56.29 $0.2 $1.2 1,050K 60.9 36.2 88 59.5 36.9 5
18 Gemini 3 Pro Preview Google DeepMind 50.57 55.02 $2 $12 1,049K 69.9 71.7 75 31.1 27.4 5
19 Claude Sonnet 4.6 Anthropic 50.06 55.36 $3 $15 1,000K 66.1 86.5 60.4 8.5 4
20 OpenAI o3-pro (2025-06-10) OpenAI 49.74 54.89 $20 $80 200K 58.2 59.3 4.9 4
21 Kimi K3 Moonshot AI (Kimi) 49.32 59.2 $2 $8 1,049K 82.6 35.8 2
22 GPT 5 2025-08-07 OpenAI 49.26 52.54 $1.25 $10 272K 60.7 48 65.7 9.9 33.7 6
23 Grok 4 xAI 47.31 49.94 $3 $15 256K 45.7 52.6 66.7 16 24.2 6
24 GPT 5 Pro 2025-10-06 OpenAI 46.96 50.71 $15 $120 400K 60.4 53.9 70.2 18.3 4
25 Gemini 2.5 Pro Experimental 0325 Google DeepMind 46.88 54.31 $2.5 $10 1,049K 41.9 2
26 Claude Sonnet 5 Anthropic 46.4 51.04 $2 $10 1,000K 68.8 52.7 31.6 3
27 Claude 4.5 Opus Anthropic 46.01 48.63 $5 $25 200K 63.7 54.4 80 37.6 7.4 5
28 Grok 4 Fast xAI 44.99 47.76 $0.2 $0.5 2,000K 42.9 48.5 5.3 4
29 o3 2025-04-16 OpenAI 44.31 45.93 $2 $8 200K 52.4 43.7 60.8 6.5 23.2 6
30 GLM 5.2 Zhipu AI / Z.ai 44.28 46.7 $0.42 $1.32 1,049K 70.1 77 22.8 16.9 4
31 GPT 5 Mini 2025-08-07 OpenAI 43.29 45.21 $0.25 $2 272K 52.7 54.3 4.4 4
32 Kimi K2.5 Moonshot AI (Kimi) 41.41 42.07 $0.35 $1.7 262K 45.6 36.2 65.3 11.8 7.4 6
33 Gemini 3 Flash Preview Google DeepMind 40.53 40.96 $0.5 $3 1,049K 61.6 53.3 21.5 33.6 34.8 5
34 Qwen3 235B A22B Thinking-2507 Alibaba (Qwen) 40.47 41.15 $0.1 $0.1 262K 41 7.4 3
35 o4 Mini 2025-04-16 OpenAI 40.33 40.63 $1.1 $4.4 200K 52.6 26.4 58.7 6.1 22.1 6
36 Qwen3.7 Max Alibaba (Qwen) 40.33 41.21 $2.5 $7.5 1,000K 64.5 17.9 2
37 Qwen 3 235b A22B Alibaba (Qwen) 40.21 40.73 $0.7 $2.8 131K 37.3 17.2 3
38 Gemini 2.5 Flash Preview Google DeepMind 39.87 40.15 $0.15 $0.6 1,049K 41 32.3 3
39 Kimi K2.7 Code Moonshot AI (Kimi) 39.87 40.16 $0.28 $1.1 262K 54.1 49.5 16.9 3
40 Claude 4 Opus Anthropic 39.75 39.87 $15 $75 200K 43.4 50.6 35.7 8.6 5
41 DeepSeek V4 Pro DeepSeek 39.41 39.39 $0.44 $0.87 1,049K 48.9 53.4 15.8 3
42 GPT 5.4 Mini 2026-03-17 OpenAI 39.25 39.15 $0.75 $4.5 272K 60.3 63.7 18.9 13.7 4
43 Kimi K2.6 Moonshot AI (Kimi) 39.22 39 $0.22 $1.14 262K 55.9 22.1 2
44 GLM 5.1 Zhipu AI / Z.ai 39.17 38.98 $0.06 $0.22 205K 57.1 46.1 13.7 3
45 Gemini 2.5 Flash 0520 Google DeepMind 38.91 38.65 $0.15 $0.6 1,049K 41 33.3 2.5 4
46 Kimi K2 Instruct Moonshot AI (Kimi) 38.18 37.34 $0.1 $2 131K 39.4 11.6 3
47 o1 2024-12-17 OpenAI 38.14 37.62 $15 $60 200K 43.8 28.1 30.7 2.2 5
48 Claude Sonnet 3.7 Anthropic 37.9 37.12 $3 $15 200K 35.7 28.6 0.9 4
49 Grok 4.3 xAI 37.47 35.49 $1.25 $2.5 1,000K 49.9 21.1 2
50 Claude 3.5 Sonnet 2024-10-22 Anthropic 37.14 34.83 $3 $15 200K 40 29.7 2
51 Gemini 2.0 Flash Thinking 0121 Google DeepMind 37.13 34.82 $0.31 $1 1,000K 16.8 2
52 MiniMax M2.5 MiniMax 36.86 34.27 $0.3 $1.2 1,000K 63.7 4.9 2
53 Qwen3.6 Max Preview Alibaba (Qwen) 36.79 34.14 $1.3 $7.8 262K 55.6 12.7 2
54 Claude Sonnet 4.5 Anthropic 36.64 35.51 $3 $15 1,000K 47.7 45.2 63.7 13.6 7.4 5
55 Qwen3 Max 2025-09-23 Alibaba (Qwen) 36.4 33.35 $0.86 $3.43 258K 0 2
56 Claude 4 Sonnet Anthropic 36.06 34.71 $3 $15 1,000K 46.1 34.6 40 5.9 5
57 DeepSeek Chat 0324 DeepSeek 35.52 32.91 $0.2 $0.6 164K 36.1 12.6 3
58 GPT 5.4 Nano 2026-03-17 OpenAI 35.28 33.19 $0.2 $1.25 272K 49.2 51.5 5.7 26.4 4
59 DeepSeek R1 0528 DeepSeek 35.26 33.58 $0.25 $0.25 164K 41.6 29 21.2 1.1 5
60 Claude 4.1 Opus Anthropic 35.16 32.3 $15 $75 200K 42.8 52 2.2 3
61 DeepSeek V3.2 DeepSeek 34.98 30.52 $0.28 $0.4 164K 57 4 2
62 o1 Preview 2024-09-12 OpenAI 34.9 31.87 $15 $60 128K 47.6 30 18 3
63 Grok 3 Mini xAI 34.18 31.55 $0.3 $0.5 131K 42.6 16.5 0.4 4
64 Gemini 2.5 Pro Google DeepMind 32.44 28.93 $1.25 $10 1,049K 54 41 4.9 15.8 4
65 GPT OSS 120B OpenAI 32.3 28.73 $0.03 $0.14 131K 48.2 6.5 15.8 4
66 GLM 5 Zhipu AI / Z.ai 32.25 29.37 $0.48 $1.9 205K 48.2 43.8 44.7 4.9 5.3 5
67 DeepSeek Reasoner DeepSeek 31.28 28.01 $0.55 $2.19 164K 36.5 17.1 15.8 1.3 5
68 Claude Sonnet 3.5 Anthropic 30.72 21.99 $2.6 $13 200K 31 13 2
69 GPT 4.5 Preview OpenAI 30.67 27.15 $75 $150 128K 39.4 21.4 10.3 0.8 5
70 Qwen3 235B A22B Instruct 2507 Alibaba (Qwen) 30.46 25.96 $0.1 $0.1 262K 38.7 11 1.3 4
71 Claude Haiku 4.5 Anthropic 29.87 25.08 $1 $5 200K 45.4 47.7 4 3.2 4
72 GLM 4.7 Zhipu AI / Z.ai 29.31 19.17 $0.15 $0.8 205K 37.2 1.1 2
73 Grok 3 xAI 29.04 24.87 $3 $15 131K 37.2 23.3 5.5 0 5
74 o3 Mini 2025-01-31 OpenAI 28.76 25.2 $1.1 $4.4 200K 43.7 7.4 34.5 3 12.7 6
75 GPT 4.1 2025-04-14 OpenAI 28.59 24.25 $2 $8 1,048K 39 12.4 5.5 0.4 5
76 Qwen3.6 Flash Alibaba (Qwen) 28.5 17.55 $0.19 $1.13 1,000K 22.2 12.9 2
77 GPT 5 Nano 2025-08-07 OpenAI 27.89 23.27 $0.05 $0.4 272K 38.1 20.7 2.6 10.6 5
78 Grok 2 1212 xAI 27.09 14.74 $2 $10 131K 22.2 7.2 2
79 Llama 3.1 405B Instruct Meta AI 26.97 14.49 $0.12 $0.3 128K 21.4 7.6 2
80 Llama 3.3 70B Instruct Meta AI 26.1 17.21 $0.05 $0.23 131K 14.4 3.9 3
81 Llama 4 Maverick 17b 128e Instruct Meta AI 23.89 17.66 $0.05 $0.1 1,049K 24.5 13.2 4.4 0 5
82 OpenAI GPT-4.1 Mini OpenAI 23.79 17.53 $0.4 $1.6 1,048K 37.6 3.5 0 2.2 5
83 Llama 4 Scout 17B 16E Instruct Meta AI 23.08 12.17 $0.05 $0.1 10,000K 0.5 0 3
84 GPT-4o (2024-08-06) OpenAI 22.18 4.91 $2.5 $10 128K 1.4 8.5 2
85 Claude 3 Opus 2024-02-29 Anthropic 22.06 10.47 $15 $75 200K 23.2 8.2 0 3
86 o1 Mini 2024-09-12 OpenAI 21.96 13.22 $1.1 $4.4 128K 36.3 1.7 14 0.8 4
87 GPT-4o (2024-11-20) OpenAI 21.7 9.87 $2.5 $10 128K 25.1 4.5 0 3
88 GPT 4 Turbo 2024-04-09 OpenAI 21.62 9.74 $10 $30 128K 18 10.1 1.1 3
89 Magistral Small 2506 Mistral AI 20.97 2.5 $0.5 $1.5 33K 5 0 2
90 GPT 4.1 Nano 2025-04-14 OpenAI 20.48 11 $0.1 $0.4 1,048K 19 0 0 4
91 OpenAI: GPT-4o-mini (2024-07-18) OpenAI 15.11 2.94 $0.15 $0.6 128K 11.8 0 0 0 4

The table scrolls sideways: not all columns fit.

The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero. Showing the five most complete task sets out of 6; the rest are on the model page.

What each task set checks
ARC-AGI
generalising to unseen patterns models measured: 64
ARC-AGI-2
models measured: 62
Chess puzzles
models measured: 61
Fiction.LiveBench
holding a long context models measured: 43
SimpleBench
trick questions models measured: 70
WeirdML
unusual machine learning tasks models measured: 101

The "task sets" column shows how many task sets this model's score is based on. The more there are, the more reliable the figure: a score from two sets is more scattered than one from six, and the coverage adjustment accounts for exactly that.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

Epoch AI — license CC-BY · primary source
По данным Epoch AI