AI ranking

Best models with open weights

The same overall score as in the main ranking, but only among models with open weights — the ones you can download and run on your own hardware. The collection shows how far open models trail the closed ones, and in which tasks they do not trail at all.

100 models in the table
107 considered in the calculation
27 July 2026 newest model released here
7 August 2026 last recalculated
# Model Developer Overall score Sign in$ / 1M Output$ / 1M Contexttokens code25 % language20 % math15 % knowledge20 % reasoning20 %
1 MiMo V2.5 Pro Xiaomi 99.89 98–100 $0.44 $0.87 1,050K 99.7 100 99.8 100 100
2 Kimi K3 Moonshot AI (Kimi) 99.59 95–100 $2 $8 1,049K 100 99.9 98.5 99.9
3 GLM 5.1 Zhipu AI / Z.ai 98.44 97–99 $0.06 $0.22 205K 98.6 100 100 96.2 97.9
4 Kimi K2.6 Moonshot AI (Kimi) 96.94 95–98 $0.22 $1.14 262K 97.2 95 99 97.4 96.6
5 GLM 5.2 Zhipu AI / Z.ai 96.2 94–98 $0.42 $1.32 1,049K 95.1 97.7 98.2 93.5 97.2
6 MiMo V2 Pro Xiaomi 94.27 92–96 $0.44 $0.87 1,049K 94.9 94 92.6 95.2 94.1
7 DeepSeek V4 Pro DeepSeek 93.55 92–95 $0.44 $0.87 1,049K 93.7 95 91.3 92.6 94.7
8 Inkling Thinking Machines Lab 93.46 90–96 $1.87 $4.68 1,049K 93 90.2 99.9 93.3 92.7
9 MiniMax M3 MiniMax 92.88 91–95 $0.3 $1.2 1,049K 94.1 92.9 91.2 93.4 92.1
10 Hy3 Tencent (Hunyuan) 92.62 88–97 $0.13 $0.53 262K 93 93.4 91.9 92.1
11 Qwen3.5 397B-A17B Alibaba (Qwen) 92.41 91–94 $0.6 $3.6 262K 92.6 91.5 92.8 93 92.2
12 GLM 5V Turbo Zhipu AI / Z.ai 92.31 89–95 $0.7 $3.1 203K 93 94 94.2 89.7 91
13 MiMo V2.5 Xiaomi 92.28 91–94 $0.14 $0.28 1,050K 93.4 92.3 90.6 92 92.3
14 Qwen3.6 Plus Alibaba (Qwen) 92.15 91–94 $0.5 $3 1,000K 92.6 91.9 93.3 91.2 92
15 GLM 5 Zhipu AI / Z.ai 92.08 90–94 $0.48 $1.9 205K 91.5 94.7 90.2 91.1 92.7
16 Gemma 4 31B Google DeepMind 91.9 88–95 $0.14 $0.4 262K 90.3 92.5 96.9 90.1 91.4
17 Gemma 4 26B-A4B Google DeepMind 90.64 87–94 $0.05 $0.29 262K 87.9 91 97 89.4 90.1
18 MiMo V2 Omni Xiaomi 90.3 88–93 $0.14 $0.28 262K 92.5 89.6 88.7 88.8 90.9
19 DeepSeek V4 Flash DeepSeek 89.9 88–92 $0.14 $0.28 1,049K 90.5 91 87.7 88.6 91
20 GLM 4.6 Zhipu AI / Z.ai 89.49 88–91 $0.29 $1.14 205K 89.1 92.5 88.9 86.7 90.2
21 GLM 4.7 Zhipu AI / Z.ai 89.4 86–92 $0.15 $0.8 205K 90.5 94 86.6 84 91
22 Mistral Medium 3.5 Mistral AI 89.24 86–92 $1.5 $7.5 262K 91.4 88.8 89.1 86.8 89.4
23 DeepSeek V3.2 DeepSeek 88.68 87–90 $0.28 $0.4 164K 88.9 89.3 89.7 86.9 88.7
24 Kimi K2 Thinking Turbo Moonshot AI (Kimi) 88.14 87–90 $1.15 $8 262K 89.9 87.4 88 87.5 87.4
25 Qwen3 235B A22B Instruct 2507 Alibaba (Qwen) 88.11 87–89 $0.1 $0.1 262K 88.1 86.6 88.6 88.7 88.8
26 Qwen3 VL 235B A22B Instruct Alibaba (Qwen) 87.87 85–91 $0.4 $1.6 262K 87.1 88 87.2 89.1 88
27 MiniMax M2.7 MiniMax 87.45 86–89 $0.3 $1.2 205K 89.5 87.3 85.9 87.4 86.3
28 GLM 4.5 Zhipu AI / Z.ai 86.92 85–89 $0.2 $0.8 131K 85.6 88.8 87.5 85.5 87.7
29 Qwen3.5 122B-A10B Alibaba (Qwen) 86.81 85–89 $0.4 $3.2 262K 86.2 87.5 87.9 86.7 86.1
30 Qwen3 Next 80B A3B Instruct Alibaba (Qwen) 86.77 84–89 $0.15 $1.2 262K 87.3 86.6 91.2 82.1 87.6
31 DeepSeek V3.2 Exp DeepSeek 86.12 83–89 $0.22 $0.33 164K 85.6 90.3 86.8 80.3 88
32 Qwen3 235B A22B Thinking-2507 Alibaba (Qwen) 86.08 82–90 $0.1 $0.1 262K 83.5 87.2 84.2 91 84.8
33 MiMo V2 Flash Xiaomi 85.66 84–87 $0.14 $0.28 262K 87.3 88.2 80 85.3 85.8
34 Qwen3.5 27B Alibaba (Qwen) 85.37 83–87 $0.3 $2.4 262K 83.7 86 89 85 84.4
35 Step 3.5 Flash StepFun 84.67 83–86 $0.1 $0.3 262K 86.3 84.9 82.7 84.6 83.9
36 Qwen3 VL 235B A22B Thinking Alibaba (Qwen) 84.17 80–88 $0.4 $4 262K 84.4 84.8 84.6 84 83.1
37 DeepSeek R1 0528 DeepSeek 84.04 81–87 $0.25 $0.25 164K 84.2 89 81.9 79.6 84.9
38 DeepSeek V3.1 DeepSeek 83.99 81–87 $0.2 $0.7 164K 81.5 88 86.2 80.7 84.7
39 DeepSeek V3.1 Terminus DeepSeek 82.26 78–87 $0.21 $0.79 164K 79.8 86.1 80.7 82.7
40 Qwen3.5 35B A3B Alibaba (Qwen) 81.48 80–83 $0.25 $2 262K 80.3 82.3 82.3 81.5 81.5
41 Qwen3 30B A3B Instruct 2507 Alibaba (Qwen) 79.97 78–82 $0.05 $0.19 262K 82.1 78.1 79.8 78.4 80.9
42 Nemotron 3 Super NVIDIA 79.17 76–83 $0.2 $0.8 1,000K 79.6 80.3 76.9 79.6 78.8
43 NVIDIA Nemotron 3 Super 120B A12B NVIDIA 78.9 75–83 $0.15 $0.65 262K 79.6 76.9 79.6 78.8
44 GLM 4.5 Air Zhipu AI / Z.ai 77.42 75–80 $0.11 $0.29 131K 77.5 79 80.5 74.2 76.6
45 Kimi K2 0905 Preview Moonshot AI (Kimi) 77 74–80 $0.6 $2.5 262K 78.3 75.6 80.4 73.7 77.6
46 Qwen3 Next 80B A3B Thinking Alibaba (Qwen) 76.62 74–80 $0.15 $1.2 262K 76.5 78.5 79.9 74.1 74.9
47 GLM 4.6V Zhipu AI / Z.ai 76.38 72–81 $0.14 $0.42 131K 76.6 78.4 74.1
48 MiniMax M2.5 MiniMax 75.03 73–77 $0.3 $1.2 1,000K 73.6 76 76.5 75.1 74.7
49 Qwen 3 235b A22B Alibaba (Qwen) 74.32 72–76 $0.7 $2.8 131K 75.2 74.4 79.9 70.1 73.1
50 GPT OSS 120B OpenAI 74.08 72–76 $0.03 $0.14 131K 74.1 73.7 78.2 72 73.5
51 Qwen3 Coder 480B A35B Alibaba (Qwen) 74.07 72–76 $1.5 $7.5 262K 81 71.1 72.8 68 75.3
52 Nemotron 3 Nano 30B A3B NVIDIA 73.51 71–76 $0.05 $0.2 262K 74 74.8 74.5 73.9 70.5
53 DeepSeek Reasoner DeepSeek 73.41 71–76 $0.55 $2.19 164K 72.2 76.5 79.2 68 72.9
54 DeepSeek Chat 0324 DeepSeek 73.4 72–75 $0.2 $0.6 164K 71.5 76.9 74.7 70.8 73.9
55 GLM 4.7 Flash Zhipu AI / Z.ai 73.13 70–76 $0.04 $0.3 203K 74.6 74.2 70.8 72.7 72.3
56 Kimi K2 0711 Moonshot AI (Kimi) 72.87 71–75 $0.6 $2.5 131K 73.5 74.3 73.1 69.7 73.7
57 INTELLECT 3 Prime Intellect 71.44 67–76 $0.2 $1.1 128K 71.5 75.2 76.8 64.6 70.4
58 Qwen3 32B Alibaba (Qwen) 71.33 66–76 $0.7 $2.8 131K 69.2 69.8 80.6 72.8 67.1
59 Trinity Large Preview Arcee AI 70.5 69–72 131K 74.1 69.5 65.5 70.7 70.6
60 Trinity Large Thinking Arcee AI 70.37 68–72 $0.22 $0.85 262K 69.4 69.2 73 72.5 68.7
61 MiniMax M2 MiniMax 70.28 66–75 $0.3 $1.2 205K 72.6 71.2 69.2 65.1 72.4
62 Llama 3.3 Nemotron Super 49B v1.5 Meta AI 69.29 64–75 $0.05 $0.25 131K 68.9 69.3 79.6 64.7 66.6
63 GLM 4.5V Zhipu AI / Z.ai 69.1 64–74 $0.29 $0.86 128K 67.3 70.2 69.2 71.7 67.6
64 MiniMax M1 MiniMax 68.67 67–71 $0.13 $1.25 1,000K 69.4 70.7 71.8 63.8 68.2
65 Mistral Small 3.2 Mistral AI 67.19 65–70 $0.1 $0.3 128K 70.2 70.7 66.9 60 67.4
66 Qwen: QwQ 32B Alibaba (Qwen) 66.37 64–69 $0.18 $0.2 131K 64.1 67.4 71.2 65.6 65.3
67 Gemma 3 27B Google DeepMind 65 63–67 $0.03 $0.11 131K 61.5 73.8 59.6 61.4 68.2
68 Qwen3 30B A3B Alibaba (Qwen) 64.79 63–67 $0.09 $0.2 131K 64.9 63.9 70.1 63.6 62.8
69 Llama 3.1 Nemotron Ultra 253B Meta AI 64.19 58–70 $0.6 $1.8 128K 59.3 66.2 71.3 63
70 DeepSeek V3 DeepSeek 62.58 60–65 $0.27 $1.1 164K 62.1 66.6 59.5 61.9 62.2
71 Command A Cohere 62.44 61–64 $2.5 $10 256K 63.3 65.9 56.7 59.4 65.3
72 Olmo 3 32B Think Allen Institute for AI 61.18 57–66 $0.15 $0.5 66K 60.9 65.9 60.7 58.3 60
73 Granite 4.1 8B IBM 60.02 55–65 $0.05 $0.1 131K 59 59.3 60.8 63.1 58.4
74 Gemma 3 12B IT Google DeepMind 57.85 53–63 $0.05 $0.1 131K 52.6 67.1 58.6 50.8 61.6
75 GPT OSS 20B OpenAI 56.33 53–60 $0.01 $0.07 131K 58.4 55.5 61.2 53.3 53.9
76 Llama 4 Maverick 17b 128e Instruct Meta AI 55.7 54–58 $0.05 $0.1 1,049K 57.2 55.7 56.9 53.2 55.5
77 Llama 4 Scout 17B 16E Instruct Meta AI 52.86 51–55 $0.05 $0.1 10,000K 53.6 55.7 53.6 49 52.4
78 Gemma 3n E4b It Google DeepMind 52.52 50–55 $0.02 $0.04 33K 49.9 60 45 50.4 56.1
79 Qwen2.5 72B Instruct Alibaba (Qwen) 52.46 51–54 $1.4 $5.6 131K 55 50.6 52.7 50.3 53.1
80 Llama 3.1 Nemotron 70B Instruct Meta AI 52.25 49–56 $0.6 $0.6 128K 50.5 58.8 49.9 49.7 52.2
81 Llama 3.1 405B Instruct Meta AI 52.16 51–54 $0.12 $0.3 128K 53.1 57 51.7 47.1 51.6
82 Llama 3.3 70B Instruct Meta AI 50.35 49–52 $0.05 $0.23 131K 49.8 56 49.1 46.4 50.3
83 Mistral Large 2407 Mistral AI 50.04 48–52 $3 $9 131K 51.6 52.2 47.6 47.8 50
84 Mistral Large Mistral AI 48.77 47–51 $2 $6 131K 51.3 50.9 47.5 43.1 50.1
85 Llama 3.1 70B Meta AI 47.43 46–49 $0.12 $0.3 131K 47.9 53.4 45.3 43.2 46.7
86 Qwen2.5 Coder 32b Instruct Alibaba (Qwen) 46.84 43–51 $0.06 $0.2 128K 51.4 41.8 45 45.5 48.9
87 Gemma 3 4B IT Google DeepMind 46.73 42–52 $0.04 $0.08 131K 41.5 55 42.2 46 49.1
88 Mistral: Mistral Small 3 Mistral AI 43.71 41–46 $0.05 $0.08 33K 44.9 43.8 42.5 41.9 44.9
89 Phi 4 Microsoft 41.31 39–44 $0.06 $0.14 128K 41.7 37.5 43.8 42.1 41.9
90 Llama 3 70B Instruct Meta AI 38.14 37–40 $0.12 $0.3 8K 36.3 49.3 37.1 31.8 36.5
91 Google: Gemma 2 27B Google DeepMind 37.41 36–39 $0.65 $0.65 8K 37.3 40.3 35.7 36.1 37.3
92 Aya Expanse 32B Cohere 35.7 34–38 128K 34.2 37 32.8 38.1 36.2
93 Command R+ Cohere 34.54 31–38 $2.5 $10 128K 32.1 39.2 29.9 36.5 34.5
94 Ministral 8B (latest) Mistral AI 34.19 30–38 $0.1 $0.1 128K 35.3 32.8 30 35.8 35.7
95 Llama 3.1 8B Meta AI 32.02 30–34 $0.02 $0.03 131K 33.8 34.1 27.8 30.8 32.1
96 Command R Cohere 28.05 25–31 $0.15 $0.6 128K 28.3 29.1 21.9 29.6 29.8
97 Aya Expanse 8B Cohere 27.91 25–31 8K 26.3 27.3 25.1 32.4 28
98 Meta: Llama 3 8B Instruct Meta AI 25.44 24–27 $0.03 $0.04 8K 24.4 33 21.1 24.9 23
99 Phi-3-medium instruct (4k) Microsoft 22.36 20–25 $0.17 $0.68 4K 19.7 21.9 26.4 23.9 21.6
100 Mixtral 8x7B Mistral AI 19.81 18–22 $0.15 $0.15 33K 18.9 21.4 20.1 20 19

The table scrolls sideways: not all columns fit.

The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero.

The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.

The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 99 models — their order among themselves is not determined by the available data.

The confidence interval for half the models is wider than 4.4 points — that is about 5 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places.