Resources

AI model rankings

AI models, their providers' prices and rankings from independent measurements. Under every table it says where the figure comes from, who measured it, when, and how far it can be trusted.

2219models
144developers
257API providers
9data sources

Last recalculated: 7 August 2026

Best AI models

Overall score

  1. 1 Claude Opus 5 99.31
  2. 2 Claude Opus 4.6 98.59
  3. 3 Claude Fable 5 97.58

The overall ranking: five abilities folded into one score so that models can be compared as a whole rather than on a single task. It is built from blind comparisons by real people, so it answers whose replies get preferred more often, not how many tasks were solved. If you need one specific skill, the dedicated collection is the more precise place to look.

of 236 models 99 with no significant difference

AI models for programming

Task score

  1. 1 Claude Opus 4.7 59.79
  2. 2 Claude Opus 4.6 59.43
  3. 3 GPT 5.4 2026-03-05 57.49

How many programming tasks the model actually solved on fixed benchmark sets — a verifiable result, not an opinion. The collection has a second view, based on blind human ratings: the order there is different, and the gap between the two lists says more than either list on its own.

of 41 models

AI models for reasoning

Task score

  1. 1 GPT-5.6 Sol 71.53
  2. 2 Claude Opus 5 65.91
  3. 3 Gemini 3.1 Pro Preview 64.89

Reasoning tasks where the answer is either right or wrong: what counts is the share of solved problems, not of pleasing answers. Next to it there is a second view, from blind human ratings. Logic is where preference and a solved task part ways the most: a well-worded chain of reasoning pleases people even when its answer is wrong.

of 91 models

AI models by knowledge

Task score

  1. 1 Gemini 3 Flash Preview 55.84
  2. 2 Qwen3.6 Max Preview 55.18
  3. 3 GPT-5.6 Sol 54.72

A test of factual knowledge on questions that have a single correct answer. It differs from the neighbouring collections in that the model does not reason or write code here — it recalls. The second view on the same topic comes from blind human ratings.

of 79 models

AI models for mathematics

Task score

  1. 1 Claude Fable 5 74.77
  2. 2 GPT-5.6 Sol 72.1
  3. 3 Claude Opus 5 71.39

The share of solved mathematical problems, from olympiad to research level. The percentages are lower than in other collections, and that is a property of the problems, not of the models: comparing them with the numbers from the knowledge collection is meaningless. The second view comes from blind human ratings.

of 88 models

AI models for agent tasks

Task score

  1. 1 Claude Opus 4.6 40.67
  2. 2 Claude 4.5 Opus 40.11
  3. 3 Claude Sonnet 4.5 39.89

Here the model does not answer a question but does the work: edits files in an operating system, runs searches, completes tasks in a working environment. That is why the percentages are low even for strong models, and they cannot be compared with scores from the knowledge collection — it is different work, not different strength.

of 32 models

AI models for English

Arena Score, English

  1. 1 Claude Opus 4.6 1,505.8
  2. 2 Claude Opus 5 1,500.16
  3. 3 Claude Fable 5 1,498.82

How models answer specifically in English: the score is built from blind comparisons in which people picked the better English reply. This is a collection about the language, not about ability — a model strong in the overall ranking can lose here to one that handles English phrasing better.

of 233 models 99 with no significant difference

AI models for French

Arena Score, French

  1. 1 Claude Opus 5 1,522.24
  2. 2 Claude Fable 5 1,519.66
  3. 3 GPT-5.6 Terra 1,503.02

Blind comparison of replies in French. It needs a collection of its own because the order of models changes noticeably from language to language — and the overall ranking does not show that difference.

of 185 models 99 with no significant difference

AI models for Spanish

Arena Score, Spanish

  1. 1 Claude Opus 4.6 1,511.71
  2. 2 Claude Fable 5 1,504.09
  3. 3 Gemini 3.1 Pro Preview 1,479.93

Blind comparison of replies in Spanish. As with the other language slices, the interesting part is not the score itself but where a model stands relative to its own position in other languages: consistency across languages says more about a model than one high line.

of 184 models 99 with no significant difference

Cheapest AI models

Price, $ per 1M tokens

  1. 1 Llama 3.2 1b Instruct 0.01
  2. 2 Ling 2.6 Flash 0.02
  3. 3 Google Gemma 2 0.02

The price of API access per million tokens, in dollars, by the cheapest offer across providers. Input and output are folded into one number: output is usually several times more expensive, and ranking models by the input price alone would put something other than the cheapest on top.

of 1508 models

Multimodal AI models

Arena Score, images

  1. 1 Claude Fable 5 1,334.18
  2. 2 Claude Opus 5 1,328.68
  3. 3 Gemini 3.6 Flash 1,315.75

How well the model understands an image together with text: blind comparison of replies to prompts that come with a picture attached. Not to be confused with image generation — there the model draws, here it looks.

of 90 models 84 with no significant difference

AI models for image generation

Arena Score, generation

  1. 1 GPT Image 2 1,384.81
  2. 2 Mai Image 2.5 1,256.56
  3. 3 Nano Banana 2 Lite 1,249.66

Blind comparison of finished images: people pick the better picture without knowing which model drew it. This judges the result by eye rather than checking it against a task, which is why the difference between neighbouring places is most often statistically insignificant.

of 33 models 19 with no significant difference

AI models for video generation

Arena Score, video

  1. 1 Sora 2 Pro 1,365.62
  2. 2 Sora 2 1,337.11
  3. 3 Seedance v1.5 Pro 1,256.99

The same blind comparison as for images, but for video. There are an order of magnitude fewer models here than in the text collections, and the top places stand apart from the rest more visibly — simply because there are few participants.

of 7 models 2 with no significant difference

AI models for web development

Arena Score, web

  1. 1 Claude Opus 5 1,704.67
  2. 2 Kimi K3 1,675.56
  3. 3 Claude Fable 5 1,630.02

Prompts about layout and web applications, rated blindly by people. It differs from the programming collection in what is measured: there, solved tasks are counted; here, whose result people preferred. For the web that is closer to the point, because how the result looks is part of what gets judged.

of 77 models 70 with no significant difference

AI models for writing

Arena Score, writing

  1. 1 Claude Opus 5 1,510.88
  2. 2 Claude Fable 5 1,498.96
  3. 3 Gemini 3 Pro 1,481.7

Fiction and free-form writing, rated blindly by people: work of this kind has no correct answer, so preference is the most honest measure available. There is also a second view, from a scored run of set tasks; it measures something else and yields a different order.

of 236 models 99 with no significant difference

Best open-weight AI models

Overall score

  1. 1 MiMo V2.5 Pro 99.89
  2. 2 Kimi K3 99.59
  3. 3 GLM 5.1 98.44

The same overall score as in the main ranking, but only among models with open weights — the ones you can download and run on your own hardware. The collection shows how far open models trail the closed ones, and in which tasks they do not trail at all.

of 107 models 99 with no significant difference

Data sources

The catalogue, the prices and the scores are collected from open sources; the right to republish is checked against the licence before a figure reaches the site. Under every table it says who measured what it shows and when. Figures we are not allowed to publish are used only for cross-checking and are not shown on the site.