How well the model understands an image together with text: blind comparison of replies to prompts that come with a picture attached. Not to be confused with image generation — there the model draws, here it looks.
| # | Model | Developer | Arena Score, images | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1,334.18 1,325–1,343 | $10 | $50 | 1,000K |
| 2 | Claude Opus 5 ≈ | Anthropic | 1,328.68 1,314–1,343 | $5 | $25 | 1,000K |
| 3 | Gemini 3.6 Flash ≈ | Google DeepMind | 1,315.75 1,278–1,354 | $1.5 | $7.5 | 1,049K |
| 4 | Claude Opus 4.7 ≈ | Anthropic | 1,314.84 1,308–1,322 | $5 | $25 | 1,000K |
| 5 | Claude Opus 4.6 ≈ | Anthropic | 1,312.19 1,306–1,319 | $5 | $25 | 1,000K |
| 6 | Gemini 3 Pro ≈ | Google DeepMind | 1,304.71 1,297–1,312 | $1.6 | $9.6 | 1,049K |
| 7 | Gemini 3.5 Flash ≈ | Google DeepMind | 1,301.48 1,289–1,314 | $1.5 | $9 | 1,049K |
| 8 | Muse Spark 1.1 ≈ | Meta AI | 1,295.72 1,285–1,306 | $1.25 | $4.25 | 1,049K |
| 9 | Gemini 3.1 Pro Preview ≈ | Google DeepMind | 1,295.15 1,289–1,301 | $2 | $12 | 1,049K |
| 10 | Grok 4.5 ≈ | xAI | 1,293.37 1,281–1,306 | $2 | $6 | 500K |
| 11 | Claude Opus 4.8 ≈ | Anthropic | 1,287.87 1,280–1,296 | $5 | $25 | 1,000K |
| 12 | Gemini 3 Flash ≈ | Google DeepMind | 1,283.64 1,278–1,289 | $0.4 | $2.4 | 1,049K |
| 13 | Kimi K2.6 ≈ | Moonshot AI (Kimi) | 1,282.13 1,275–1,289 | $0.22 | $1.14 | 262K |
| 14 | Claude Sonnet 5 ≈ | Anthropic | 1,281.86 1,272–1,292 | $2 | $10 | 1,000K |
| 15 | Claude Sonnet 4.6 ≈ | Anthropic | 1,281.47 1,275–1,288 | $3 | $15 | 1,000K |
| 16 | Qwen3.7 Plus ≈ | Alibaba (Qwen) | 1,277.36 1,268–1,286 | $0.5 | $3 | 1,000K |
| 17 | GPT-5.6 Terra ≈ | OpenAI | 1,276.97 1,263–1,291 | $2 | $12 | 1,050K |
| 18 | Seed 2.0 Pro ≈ | ByteDance (Doubao) | 1,275.5 1,268–1,283 | $0.5 | $3 | 131K |
| 19 | GPT-5.6 Sol ≈ | OpenAI | 1,273.53 1,260–1,287 | $5 | $30 | 1,050K |
| 20 | Gemma 4 31B ≈ | Google DeepMind | 1,271.6 1,265–1,278 | $0.14 | $0.4 | 262K |
| 21 | GPT-5.2 Chat ≈ | OpenAI | 1,266.98 1,260–1,274 | $1.75 | $14 | 128K |
| 22 | Qwen3.5 397B-A17B ≈ | Alibaba (Qwen) | 1,266.03 1,260–1,272 | $0.6 | $3.6 | 262K |
| 23 | Kimi K2.5 Thinking TEE ≈ | Moonshot AI (Kimi) | 1,265.7 1,260–1,272 | $0.3 | $1.9 | 256K |
| 24 | GLM 5V Turbo ≈ | Zhipu AI / Z.ai | 1,264.04 1,258–1,271 | $0.7 | $3.1 | 203K |
| 25 | Gemini 2.5 Pro ≈ | Google DeepMind | 1,261.67 1,257–1,266 | $1.25 | $10 | 1,049K |
| 26 | Grok 4.20 (Reasoning) ≈ | xAI | 1,261.56 1,255–1,268 | $2 | $6 | 2,000K |
| 27 | Grok 4.20 Multi Agent Beta 0309 ≈ | xAI | 1,259.73 1,253–1,266 | $2 | $6 | 2,000K |
| 28 | MiniMax M3 ≈ | MiniMax | 1,258.26 1,250–1,267 | $0.3 | $1.2 | 1,049K |
| 29 | Gemma 4 26B-A4B ≈ | Google DeepMind | 1,256.71 1,250–1,264 | $0.05 | $0.29 | 262K |
| 30 | Gemini 3.5 Flash Lite ≈ | Google DeepMind | 1,256.63 1,220–1,294 | $0.3 | $2.5 | 1,049K |
| 31 | MiMo V2.5 ≈ | Xiaomi | 1,253.53 1,247–1,260 | $0.14 | $0.28 | 1,050K |
| 32 | GPT-5.5 Instant ≈ | OpenAI | 1,250.81 1,242–1,260 | $5 | $30 | 400K |
| 33 | GPT-5.6 Luna ≈ | OpenAI | 1,248.77 1,235–1,262 | $0.2 | $1.2 | 1,050K |
| 34 | 2025) ≈ | Google DeepMind | 1,248.77 1,239–1,259 | $0.3 | $2.5 | 1,049K |
| 35 | Qwen3 VL 235B A22B Instruct ≈ | Alibaba (Qwen) | 1,246.72 1,240–1,254 | $0.4 | $1.6 | 262K |
| 36 | Qwen3.5 122B-A10B ≈ | Alibaba (Qwen) | 1,246.04 1,239–1,253 | $0.4 | $3.2 | 262K |
| 37 | ChatGPT 4o Latest ≈ | OpenAI | 1,244.77 1,239–1,250 | $5 | $15 | 128K |
| 38 | Qwen3.5 27B ≈ | Alibaba (Qwen) | 1,240.78 1,234–1,247 | $0.3 | $2.4 | 262K |
| 39 | Gemini 3.1 Flash Lite Preview ≈ | Google DeepMind | 1,240.09 1,234–1,246 | $0.25 | $1.5 | 1,049K |
| 40 | Gemini 2.5 Flash ≈ | Google DeepMind | 1,235.38 1,231–1,240 | $0.3 | $2.5 | 1,049K |
| 41 | Grok 4.3 ≈ | xAI | 1,232.65 1,225–1,240 | $1.25 | $2.5 | 1,000K |
| 42 | GPT 5 Chat ≈ | OpenAI | 1,231.63 1,225–1,239 | $1.25 | $10 | 128K |
| 43 | MiMo V2 Omni ≈ | Xiaomi | 1,229.11 1,221–1,237 | $0.14 | $0.28 | 262K |
| 44 | Mistral Large 3 ≈ | Mistral AI | 1,224.01 1,210–1,238 | $0.5 | $1.5 | 256K |
| 45 | Mistral Medium 3.5 ≈ | Mistral AI | 1,221.76 1,212–1,232 | $1.5 | $7.5 | 262K |
| 46 | o3 2025-04-16 ≈ | OpenAI | 1,215.41 1,209–1,222 | $2 | $8 | 200K |
| 47 | GPT 4.1 2025-04-14 ≈ | OpenAI | 1,209.86 1,203–1,217 | $2 | $8 | 1,048K |
| 48 | Grok 4 ≈ | xAI | 1,208.77 1,201–1,216 | $3 | $15 | 256K |
| 49 | Qwen3 VL 235B A22B Thinking ≈ | Alibaba (Qwen) | 1,207.77 1,195–1,220 | $0.4 | $4 | 262K |
| 50 | Grok 4.1 Fast Reasoning ≈ | xAI | 1,200.22 1,193–1,208 | $0.2 | $0.5 | 2,000K |
| 51 | GPT 4.5 Preview ≈ | OpenAI | 1,195.13 1,183–1,207 | $75 | $150 | 128K |
| 52 | o4 Mini 2025-04-16 ≈ | OpenAI | 1,194.83 1,188–1,202 | $1.1 | $4.4 | 200K |
| 53 | Gemini 2.5 Flash Lite Preview ≈ | Google DeepMind | 1,187.49 1,180–1,195 | $0.1 | $0.4 | 1,049K |
| 54 | OpenAI GPT-4.1 Mini ≈ | OpenAI | 1,182.65 1,175–1,190 | $0.4 | $1.6 | 1,048K |
| 55 | Step 3 ≈ | StepFun | 1,176.31 1,165–1,188 | $0.21 | $0.57 | 66K |
| 56 | Claude 4 Sonnet ≈ | Anthropic | 1,175.73 1,162–1,190 | $3 | $15 | 1,000K |
| 57 | Claude 4 Opus ≈ | Anthropic | 1,175.62 1,163–1,189 | $15 | $75 | 200K |
| 58 | Mistral Medium 3.1 ≈ | Mistral AI | 1,172.01 1,166–1,178 | $0.4 | $2 | 262K |
| 59 | o1 2024-12-17 ≈ | OpenAI | 1,168.58 1,158–1,179 | $15 | $60 | 200K |
| 60 | GLM 4.6V ≈ | Zhipu AI / Z.ai | 1,166.43 1,152–1,180 | $0.14 | $0.42 | 131K |
| 61 | Gemma 3 27B ≈ | Google DeepMind | 1,165.18 1,157–1,173 | $0.03 | $0.11 | 131K |
| 62 | Mistral Medium 3 ≈ | Mistral AI | 1,157.54 1,149–1,166 | $0.4 | $2 | 131K |
| 63 | GLM 4.5V ≈ | Zhipu AI / Z.ai | 1,156.05 1,144–1,168 | $0.29 | $0.86 | 128K |
| 64 | Qwen2.5 VL 32B Instruct ≈ | Alibaba (Qwen) | 1,152.49 1,137–1,168 | $0.05 | $0.22 | 128K |
| 65 | Claude Sonnet 3.7 ≈ | Anthropic | 1,151 1,142–1,160 | $3 | $15 | 200K |
| 66 | Llama 4 Maverick 17b 128e Instruct ≈ | Meta AI | 1,142.98 1,134–1,152 | $0.05 | $0.1 | 1,049K |
| 67 | Mistral Small 3.2 ≈ | Mistral AI | 1,142.32 1,133–1,151 | $0.1 | $0.3 | 128K |
| 68 | Mistral Small 3.1 24B Instruct 2503 ≈ | Mistral AI | 1,136.87 1,128–1,145 | $0.1 | $0.3 | 32K |
| 69 | GPT-4o (2024-05-13) ≈ | OpenAI | 1,136.75 1,128–1,145 | $5 | $15 | 128K |
| 70 | Claude 3.5 Sonnet 2024-10-22 ≈ | Anthropic | 1,125.48 1,118–1,133 | $3 | $15 | 200K |
| 71 | Claude Sonnet 3.5 ≈ | Anthropic | 1,120.44 1,110–1,130 | $2.6 | $13 | 200K |
| 72 | Llama 4 Scout 17B 16E Instruct ≈ | Meta AI | 1,118.25 1,109–1,128 | $0.05 | $0.1 | 10,000K |
| 73 | Qwen2.5 VL 72B Instruct ≈ | Alibaba (Qwen) | 1,107.58 1,097–1,118 | $2.8 | $8.4 | 131K |
| 74 | Claude 3.5 Haiku ≈ | Anthropic | 1,092.78 1,077–1,109 | $0.25 | $1.25 | 200K |
| 75 | GPT 4 Turbo 2024-04-09 ≈ | OpenAI | 1,090 1,078–1,102 | $10 | $30 | 128K |
| 76 | Mistral: Pixtral Large 2411 ≈ | Mistral AI | 1,089.58 1,080–1,099 | $2 | $6 | 131K |
| 77 | OpenAI: GPT-4o-mini (2024-07-18) | OpenAI | 1,065.76 1,058–1,074 | $0.15 | $0.6 | 128K |
| 78 | GPT-4o (2024-08-06) ≈ | OpenAI | 1,064.68 1,052–1,077 | $2.5 | $10 | 128K |
| 79 | GPT 4.1 Nano 2025-04-14 ≈ | OpenAI | 1,063.53 1,045–1,082 | $0.1 | $0.4 | 1,048K |
| 80 | Qwen-VL Max ≈ | Alibaba (Qwen) | 1,057.73 1,042–1,074 | $0.8 | $3.2 | 131K |
| 81 | Claude 3 Opus 2024-02-29 | Anthropic | 1,023.06 1,013–1,033 | $15 | $75 | 200K |
| 82 | Pixtral 12B ≈ | Mistral AI | 1,008.72 999–1,018 | $0.15 | $0.15 | 128K |
| 83 | Aya Vision 32B ≈ | Cohere | 995.83 974–1,018 | — | — | 16K |
| 84 | Nova Lite ≈ | Amazon | 990.58 976–1,006 | $0.06 | $0.24 | 300K |
| 85 | Qwen2 VL 7B Instruct ≈ | Alibaba (Qwen) | 990.05 980–1,000 | $0.02 | $0.06 | 131K |
| 86 | Claude 3 Sonnet ≈ | Anthropic | 984.05 973–995 | $3 | $15 | 200K |
| 87 | Nova Pro ≈ | Amazon | 980.58 967–994 | $0.56 | $2.13 | 300K |
| 88 | Anthropic: Claude 3 Haiku | Anthropic | 950.29 938–963 | $0.25 | $1.25 | 200K |
| 89 | Phi-3.5-vision instruct (128k) | Microsoft | 851.6 836–867 | $0.13 | $0.52 | 128K |
| 90 | Phi 3 Vision 128K Instruct | Microsoft | 811.86 793–831 | $0.2 | $0.2 | 32K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 84 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is wider than 18.1 points — that is about 3 adjacent rows of the table. Inside such a group the order is set not by the data but by who happened to vote this time; the data supports the difference between the top and the bottom of the list, but not the ordering of neighbouring places. Votes per model here — median 10,507: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)