The overall ranking: five abilities folded into one score so that models can be compared as a whole rather than on a single task. It is built from blind comparisons by real people, so it answers whose replies get preferred more often, not how many tasks were solved. If you need one specific skill, the dedicated collection is the more precise place to look.
| # | Model | Developer | Overall score | Sign in$ / 1M | Output$ / 1M | Contexttokens | code25 % | language20 % | math15 % | knowledge20 % | reasoning20 % |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 99.31 94–100 | $5 | $25 | 1,000K | 99.9 | 98.8 | 100 | 97.9 | 100 |
| 2 | Claude Opus 4.6 ≈ | Anthropic | 98.59 97–99 | $5 | $25 | 1,000K | 100 | 100 | 93.2 | 100 | 98.1 |
| 3 | Claude Fable 5 ≈ | Anthropic | 97.58 95–99 | $10 | $50 | 1,000K | 97 | 98.5 | 98.1 | 98.2 | 96.4 |
| 4 | Claude Opus 4.7 ≈ | Anthropic | 95.03 94–96 | $5 | $25 | 1,000K | 96.5 | 96.1 | 89.6 | 97.5 | 93.7 |
| 5 | Gemini 3.5 Flash ≈ | Google DeepMind | 93.62 91–96 | $1.5 | $9 | 1,049K | 91.9 | 94.4 | 96.2 | 95.3 | 91.3 |
| 6 | Gemini 3.6 Flash ≈ | Google DeepMind | 93.35 90–97 | $1.5 | $7.5 | 1,049K | 92.6 | 95.3 | 93.3 | 94.2 | 91.6 |
| 7 | Kimi K3 ≈ | Moonshot AI (Kimi) | 92.89 89–97 | $2 | $8 | 1,049K | 93.4 | 94 | — | 92.9 | 91.2 |
| 8 | MiMo V2.5 Pro ≈ | Xiaomi | 92.34 91–94 | $0.44 | $0.87 | 1,050K | 93.1 | 94 | 87.6 | 94.3 | 91.3 |
| 9 | Claude Sonnet 4.6 ≈ | Anthropic | 91.79 91–93 | $3 | $15 | 1,000K | 93.9 | 93.1 | 84.5 | 93.9 | 91.2 |
| 10 | Muse Spark 1.1 ≈ | Meta AI | 91.53 89–94 | $1.25 | $4.25 | 1,049K | 92.6 | 94.6 | 86.3 | 91.3 | 91.4 |
| 11 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 91.38 90–93 | $2 | $12 | 1,049K | 90.4 | 94.1 | 89.7 | 91.4 | 91.1 |
| 12 | ERNIE 5.1 ≈ | Baidu (Ernie) | 91.13 90–93 | $0.75 | $3 | 119K | 91.3 | 93.9 | 87.8 | 91.4 | 90.4 |
| 13 | GLM 5.1 ≈ | Zhipu AI / Z.ai | 91.07 90–93 | $0.06 | $0.22 | 205K | 92.1 | 94.1 | 87.8 | 90.9 | 89.5 |
| 14 | Gemini 3 Pro ≈ | Google DeepMind | 90.72 89–92 | $1.6 | $9.6 | 1,049K | 90.2 | 94.5 | 87.3 | 90.3 | 90.7 |
| 15 | Claude 4.5 Opus ≈ | Anthropic | 90.1 89–91 | $5 | $25 | 200K | 93.1 | 90.3 | 84.2 | 91.4 | 89.4 |
| 16 | GPT-5.6 Sol ≈ | OpenAI | 90.08 87–93 | $5 | $30 | 1,050K | 90.8 | 90.4 | 83 | 94.6 | 89.6 |
| 17 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 89.84 88–91 | $0.22 | $1.14 | 262K | 91 | 89.9 | 87 | 92 | 88.4 |
| 18 | GPT-5.6 Terra ≈ | OpenAI | 89.8 87–93 | $2 | $12 | 1,050K | 91.3 | 86.8 | 86.6 | 94.7 | 88.4 |
| 19 | Claude Opus 4.8 ≈ | Anthropic | 89.66 88–91 | $5 | $25 | 1,000K | 90.7 | 89.8 | 85.5 | 91.8 | 89.1 |
| 20 | GLM 5.2 ≈ | Zhipu AI / Z.ai | 89.16 87–91 | $0.42 | $1.32 | 1,049K | 89.2 | 92.2 | 86.4 | 88.5 | 88.9 |
| 21 | Grok 4.5 ≈ | xAI | 89.02 86–92 | $2 | $6 | 500K | 89.5 | 89.5 | 88.8 | 89.9 | 87.3 |
| 22 | Claude Sonnet 5 ≈ | Anthropic | 88.96 87–91 | $2 | $10 | 1,000K | 90.7 | 87.8 | 84.8 | 93 | 87 |
| 23 | Qwen3.7 Plus ≈ | Alibaba (Qwen) | 88.49 87–90 | $0.5 | $3 | 1,000K | 90 | 89.8 | 86 | 88.3 | 87.3 |
| 24 | Gemini 3 Flash ≈ | Google DeepMind | 87.89 86–90 | $0.4 | $2.4 | 1,049K | 86.4 | 91.4 | 86.9 | 87.4 | 87.5 |
| 25 | MiMo V2 Pro ≈ | Xiaomi | 87.62 86–89 | $0.44 | $0.87 | 1,049K | 89 | 89.1 | 82 | 90 | 86.3 |
| 26 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 87.57 86–89 | $0.3 | $1.9 | 256K | 88.5 | 89 | 86.3 | 87.9 | 85.6 |
| 27 | Qwen3.6 Max Preview ≈ | Alibaba (Qwen) | 87.18 84–91 | $1.3 | $7.8 | 262K | 87.3 | 88.2 | 84.7 | 89.2 | 85.9 |
| 28 | Claude Sonnet 4.5 ≈ | Anthropic | 87.11 86–88 | $3 | $15 | 1,000K | 90.7 | 88.5 | 77.1 | 89 | 87 |
| 29 | DeepSeek V4 Pro ≈ | DeepSeek | 86.98 86–88 | $0.44 | $0.87 | 1,049K | 87.9 | 89.9 | 81 | 87.6 | 86.8 |
| 30 | GPT-5.6 Luna ≈ | OpenAI | 86.97 84–90 | $0.2 | $1.2 | 1,050K | 88.1 | 87.2 | 82.4 | 90.3 | 85.4 |
| 31 | Hy3 ≈ | Tencent (Hunyuan) | 86.92 83–91 | $0.13 | $0.53 | 262K | 87.3 | 88.5 | — | 87.1 | 84.6 |
| 32 | Inkling ≈ | Thinking Machines Lab | 86.83 84–90 | $1.87 | $4.68 | 1,049K | 87.3 | 85.9 | 87.7 | 88.3 | 85.1 |
| 33 | MiniMax M3 ≈ | MiniMax | 86.44 85–88 | $0.3 | $1.2 | 1,049K | 88.3 | 88.1 | 80.9 | 88.4 | 84.6 |
| 34 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 86.02 85–87 | $0.6 | $3.6 | 262K | 87 | 87 | 82.2 | 88.1 | 84.7 |
| 35 | MiMo V2.5 ≈ | Xiaomi | 85.93 84–87 | $0.14 | $0.28 | 1,050K | 87.7 | 87.7 | 80.5 | 87.2 | 84.8 |
| 36 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 85.89 83–89 | $0.7 | $3.1 | 203K | 87.3 | 89.1 | 83.2 | 85.1 | 83.7 |
| 37 | Gemini 2.5 Pro ≈ | Google DeepMind | 85.79 85–87 | $1.25 | $10 | 1,049K | 84.5 | 89.8 | 82.5 | 86 | 85.7 |
| 38 | Qwen3.6 Plus ≈ | Alibaba (Qwen) | 85.78 84–87 | $0.5 | $3 | 1,000K | 87 | 87.3 | 82.6 | 86.4 | 84.5 |
| 39 | GLM 5 ≈ | Zhipu AI / Z.ai | 85.73 84–87 | $0.48 | $1.9 | 205K | 86 | 89.6 | 80.2 | 86.3 | 85.1 |
| 40 | Grok 4.20 (Reasoning) ≈ | xAI | 85.62 84–87 | $2 | $6 | 2,000K | 85.7 | 89.7 | 83.7 | 83.3 | 85.2 |
| 41 | Gemma 4 31B ≈ | Google DeepMind | 85.5 82–89 | $0.14 | $0.4 | 262K | 85 | 87.8 | 85.4 | 85.4 | 84 |
| 42 | Seed 2.0 Pro ≈ | ByteDance (Doubao) | 85.5 84–87 | $0.5 | $3 | 131K | 88 | 88.2 | 80.3 | 83.6 | 85.6 |
| 43 | Qwen3 Max Preview ≈ | Alibaba (Qwen) | 85.23 83–87 | $1.2 | $6 | 256K | 85.4 | 86.1 | 82.4 | 86.9 | 84.6 |
| 44 | Grok 4.20 Multi Agent Beta 0309 ≈ | xAI | 84.98 84–86 | $2 | $6 | 2,000K | 85.5 | 88.7 | 80.4 | 84.5 | 84.6 |
| 45 | Gemma 4 26B-A4B ≈ | Google DeepMind | 84.42 81–88 | $0.05 | $0.29 | 262K | 82.9 | 86.6 | 85.5 | 84.9 | 82.9 |
| 46 | MiMo V2 Omni ≈ | Xiaomi | 84.25 82–86 | $0.14 | $0.28 | 262K | 86.9 | 85.4 | 79 | 84.3 | 83.6 |
| 47 | Longcat Flash Chat ≈ | Meituan | 84.17 81–87 | — | — | 131K | 87.4 | 86.2 | 80.7 | 83.2 | 81.7 |
| 48 | DeepSeek V4 Flash ≈ | DeepSeek | 83.91 83–85 | $0.14 | $0.28 | 1,049K | 85.2 | 86.6 | 78.2 | 84.1 | 83.7 |
| 49 | GPT-5.2 Chat ≈ | OpenAI | 83.62 82–85 | $1.75 | $14 | 128K | 83.4 | 87 | 79.4 | 83.3 | 84 |
| 50 | GLM 4.6 ≈ | Zhipu AI / Z.ai | 83.52 82–85 | $0.29 | $1.14 | 205K | 83.9 | 87.8 | 79.1 | 82.5 | 83.1 |
| 51 | GLM 4.7 ≈ | Zhipu AI / Z.ai | 83.44 81–86 | $0.15 | $0.8 | 205K | 85.1 | 89.1 | 77.4 | 80 | 83.7 |
| 52 | Mistral Medium 3.5 ≈ | Mistral AI | 83.32 81–86 | $1.5 | $7.5 | 262K | 86 | 84.8 | 79.3 | 82.5 | 82.4 |
| 53 | Claude 4.1 Opus ≈ | Anthropic | 83.23 82–84 | $15 | $75 | 200K | 88.4 | 83.7 | 77.1 | 80.9 | 83.2 |
| 54 | DeepSeek V3.2 ≈ | DeepSeek | 82.83 81–84 | $0.28 | $0.4 | 164K | 83.8 | 85.1 | 79.8 | 82.6 | 81.8 |
| 55 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 82.76 79–86 | $0.3 | $2.5 | 1,049K | 85.1 | 84.7 | 75.8 | 82.6 | 83.3 |
| 56 | Kimi K2 Thinking Turbo ≈ | Moonshot AI (Kimi) | 82.41 81–84 | $1.15 | $8 | 262K | 84.7 | 83.6 | 78.5 | 83.1 | 80.7 |
| 57 | Qwen3 235B A22B Instruct 2507 ≈ | Alibaba (Qwen) | 82.38 81–83 | $0.1 | $0.1 | 262K | 83.1 | 82.9 | 78.9 | 84.2 | 81.8 |
| 58 | Qwen3 VL 235B A22B Instruct ≈ | Alibaba (Qwen) | 82.19 79–85 | $0.4 | $1.6 | 262K | 82.2 | 84.1 | 77.9 | 84.6 | 81.2 |
| 59 | Grok 4.1 ≈ | xAI | 81.94 81–83 | $2 | $10 | 200K | 83 | 87.4 | 76.3 | 79 | 82.3 |
| 60 | MiniMax M2.7 ≈ | MiniMax | 81.85 81–83 | $0.3 | $1.2 | 205K | 84.3 | 83.5 | 76.8 | 83.1 | 79.7 |
| 61 | Mistral Large 3 ≈ | Mistral AI | 81.55 80–83 | $0.5 | $1.5 | 256K | 82.9 | 86 | 75.9 | 80 | 81.2 |
| 62 | GLM 4.5 ≈ | Zhipu AI / Z.ai | 81.35 79–83 | $0.2 | $0.8 | 131K | 80.9 | 84.7 | 78.1 | 81.4 | 80.9 |
| 63 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 81.26 80–83 | $0.4 | $3.2 | 262K | 81.4 | 83.7 | 78.4 | 82.5 | 79.6 |
| 64 | Qwen3 Next 80B A3B Instruct ≈ | Alibaba (Qwen) | 81.16 79–83 | $0.15 | $1.2 | 262K | 82.4 | 82.9 | 80.9 | 78.4 | 80.8 |
| 65 | Qwen3 235B A22B Thinking-2507 ≈ | Alibaba (Qwen) | 80.72 78–84 | $0.1 | $0.1 | 262K | 79.1 | 83.4 | 75.5 | 86.2 | 78.5 |
| 66 | DeepSeek V3.2 Exp ≈ | DeepSeek | 80.62 78–83 | $0.22 | $0.33 | 164K | 80.9 | 86 | 77.6 | 76.7 | 81.1 |
| 67 | Claude Haiku 4.5 ≈ | Anthropic | 80.51 80–81 | $1 | $5 | 200K | 84.2 | 80.6 | 71.7 | 84.1 | 78.9 |
| 68 | Mistral Medium 3.1 ≈ | Mistral AI | 80.39 79–81 | $0.4 | $2 | 262K | 81.2 | 85.4 | 75 | 78.2 | 80.6 |
| 69 | MiMo V2 Flash ≈ | Xiaomi | 80.37 79–82 | $0.14 | $0.28 | 262K | 82.4 | 84.2 | 72.3 | 81.2 | 79.3 |
| 70 | Qwen3 Max 2025-09-23 ≈ | Alibaba (Qwen) | 80.21 77–83 | $0.86 | $3.43 | 258K | 82 | 81.9 | 80.3 | 76.2 | 80.2 |
| 71 | GPT-5.5 Instant ≈ | OpenAI | 80.14 78–82 | $5 | $30 | 400K | 81.1 | 81.5 | 77.4 | 79.2 | 80.6 |
| 72 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 80.01 78–82 | $0.3 | $2.4 | 262K | 79.3 | 82.4 | 79.3 | 80.9 | 78.1 |
| 73 | Step 3.5 Flash ≈ | StepFun | 79.5 78–81 | $0.1 | $0.3 | 262K | 81.6 | 81.5 | 74.3 | 80.6 | 77.8 |
| 74 | Gemini 2.5 Flash ≈ | Google DeepMind | 79.42 79–80 | $0.3 | $2.5 | 1,049K | 79.1 | 81.5 | 75 | 80.9 | 79.8 |
| 75 | Qwen3 VL 235B A22B Thinking ≈ | Alibaba (Qwen) | 79.04 76–82 | $0.4 | $4 | 262K | 79.9 | 81.4 | 75.8 | 80.1 | 77.1 |
| 76 | DeepSeek R1 0528 ≈ | DeepSeek | 78.91 77–81 | $0.25 | $0.25 | 164K | 79.7 | 84.9 | 73.8 | 76.1 | 78.5 |
| 77 | DeepSeek V3.1 ≈ | DeepSeek | 78.83 76–81 | $0.2 | $0.7 | 164K | 77.4 | 84.1 | 77 | 77.1 | 78.4 |
| 78 | ChatGPT 4o Latest ≈ | OpenAI | 78.67 78–80 | $5 | $15 | 128K | 77.4 | 84.4 | 74.1 | 76.5 | 80.2 |
| 79 | Grok 4 ≈ | xAI | 78.38 77–80 | $3 | $15 | 256K | 76.3 | 81.5 | 77.6 | 79.6 | 77.2 |
| 80 | 2025) ≈ | Google DeepMind | 77.79 76–79 | $0.3 | $2.5 | 1,049K | 75 | 80 | 76.1 | 80.9 | 77.3 |
| 81 | Grok 4.1 Fast Reasoning ≈ | xAI | 77.39 76–79 | $0.2 | $0.5 | 2,000K | 76.9 | 81.4 | 74 | 76.8 | 77.1 |
| 82 | o3 2025-04-16 ≈ | OpenAI | 77.34 76–79 | $2 | $8 | 200K | 76.2 | 80.7 | 77.6 | 76.6 | 76 |
| 83 | Qwen3.5 Flash ≈ | Alibaba (Qwen) | 77.26 76–79 | $0.09 | $0.36 | 1,000K | 77.2 | 78.8 | 74.8 | 78.5 | 76.4 |
| 84 | Grok 4 Fast Reasoning ≈ | xAI | 77.22 75–79 | $0.2 | $0.5 | 2,000K | 78.1 | 79.1 | 73.7 | 77.9 | 76.2 |
| 85 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 77.17 76–78 | $0.25 | $1.5 | 1,049K | 74.6 | 80.3 | 78.9 | 76.3 | 76.9 |
| 86 | DeepSeek V3.1 Terminus ≈ | DeepSeek | 77.17 73–81 | $0.21 | $0.79 | 164K | 75.9 | 82.5 | 72.8 | — | 76.7 |
| 87 | Qwen3.5 35B A3B ≈ | Alibaba (Qwen) | 76.76 75–78 | $0.25 | $2 | 262K | 76.4 | 79.3 | 74 | 77.8 | 75.7 |
| 88 | GPT 5 Chat ≈ | OpenAI | 76.59 75–78 | $1.25 | $10 | 128K | 74.7 | 79.1 | 74.7 | 77.8 | 76.6 |
| 89 | GPT 4.5 Preview ≈ | OpenAI | 76.53 74–79 | $75 | $150 | 128K | 74.1 | 81.8 | 75.3 | 75.5 | 76.3 |
| 90 | MiMo V2 Flash (Thinking) ≈ | Xiaomi | 76.53 74–79 | $0.1 | $0.31 | 256K | 77.9 | 79.9 | 69.1 | 77.1 | 76.5 |
| 91 | Grok 4.3 ≈ | xAI | 75.87 74–77 | $1.25 | $2.5 | 1,000K | 77.6 | 78.8 | 71.4 | 74.4 | 75.6 |
| 92 | Qwen3 30B A3B Instruct 2507 ≈ | Alibaba (Qwen) | 75.5 74–77 | $0.05 | $0.19 | 262K | 77.9 | 75.8 | 72.1 | 75.1 | 75.2 |
| 93 | Nemotron 3 Super ≈ | NVIDIA | 74.85 72–78 | $0.2 | $0.8 | 1,000K | 75.8 | 77.6 | 69.8 | 76.1 | 73.5 |
| 94 | GPT-5.3 Chat (latest) ≈ | OpenAI | 74.8 73–76 | $1.75 | $14 | 128K | 75.8 | 75.5 | 70.2 | 76 | 75.1 |
| 95 | Hunyuan T1 ≈ | Tencent (Hunyuan) | 74.59 70–79 | — | — | 131K | 72.3 | 76.6 | 74.9 | 74.5 | 75.4 |
| 96 | NVIDIA Nemotron 3 Super 120B A12B ≈ | NVIDIA | 74.16 71–77 | $0.15 | $0.65 | 262K | 75.8 | — | 69.8 | 76.1 | 73.5 |
| 97 | GLM 4.5 Air ≈ | Zhipu AI / Z.ai | 73.28 72–75 | $0.11 | $0.29 | 131K | 74 | 76.6 | 72.6 | 71.3 | 71.6 |
| 98 | GLM 4.6V ≈ | Zhipu AI / Z.ai | 72.93 69–77 | $0.14 | $0.42 | 131K | 73.2 | 76 | — | — | 69.5 |
| 99 | Kimi K2 0905 Preview ≈ | Moonshot AI (Kimi) | 72.93 70–76 | $0.6 | $2.5 | 262K | 74.6 | 73.7 | 72.6 | 70.9 | 72.4 |
| 100 | Qwen3 Next 80B A3B Thinking ≈ | Alibaba (Qwen) | 72.61 70–75 | $0.15 | $1.2 | 262K | 73.1 | 76.2 | 72.2 | 71.2 | 70.2 |
The table scrolls sideways: not all columns fit.
The columns on the right are the components of the score, brought to a common 0–100 scale by the actual spread among the measured models. Added together with the weights shown, they produce the number in the main column: they show exactly where one model beat another. A dash means "not measured", not zero.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 99 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 3.4 points — that is about 12 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places.