Blind comparison of finished images: people pick the better picture without knowing which model drew it. This judges the result by eye rather than checking it against a task, which is why the difference between neighbouring places is most often statistically insignificant.
| # | Model | Developer | Arena Score, generation | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | GPT Image 2 | OpenAI | 1,384.81 1,380–1,390 | $5 | $30 | 272K |
| 2 | Mai Image 2.5 | Microsoft | 1,256.56 1,252–1,261 | — | — | — |
| 3 | Nano Banana 2 Lite ≈ | Google DeepMind | 1,249.66 1,242–1,257 | $0.25 | $30 | 66K |
| 4 | Nano Banana Pro | Google DeepMind | 1,231.96 1,227–1,237 | $2 | $12 | 66K |
| 5 | Seedream 5.0 Pro ≈ | ByteDance (Doubao) | 1,231.15 1,220–1,242 | — | — | — |
| 6 | Grok Imagine Image Quality ≈ | xAI | 1,229.3 1,225–1,234 | — | — | 8K |
| 7 | Qwen Image 2.0 Pro | Alibaba (Qwen) | 1,192.68 1,186–1,200 | — | — | 8K |
| 8 | Grok Imagine Image | xAI | 1,173.01 1,170–1,176 | — | — | 8K |
| 9 | Recraft V4.1 Utility Pro ≈ | Recraft AI | 1,169.22 1,158–1,180 | — | — | — |
| 10 | FLUX.2 [max] ≈ | Black Forest Labs | 1,161.94 1,158–1,165 | — | — | 67K |
| 11 | FLUX.2 [flex] ≈ | Black Forest Labs | 1,156.25 1,153–1,160 | — | — | — |
| 12 | FLUX.2 [pro] ≈ | Black Forest Labs | 1,155.22 1,152–1,158 | — | — | 67K |
| 13 | Seedream 4.5 | ByteDance (Doubao) | 1,146.61 1,144–1,150 | — | — | — |
| 14 | Seedream 5.0 Lite | ByteDance (Doubao) | 1,132.5 1,129–1,136 | — | — | — |
| 15 | Recraft V4.1 Pro ≈ | Recraft AI | 1,130.19 1,120–1,141 | — | — | — |
| 16 | Imagen 4 ≈ | Google DeepMind | 1,128.82 1,126–1,132 | — | — | 0K |
| 17 | GPT Image 1 | OpenAI | 1,115.17 1,112–1,118 | $5 | $40 | 128K |
| 18 | Recraft V4 ≈ | Recraft AI | 1,112.75 1,109–1,116 | — | — | — |
| 19 | GPT Image 1 Mini ≈ | OpenAI | 1,109.28 1,106–1,112 | $2 | $8 | — |
| 20 | Wan2.7 Image Pro ≈ | Alibaba (Qwen) | 1,102.15 1,097–1,107 | — | — | 8K |
| 21 | Wan2.7 Image ≈ | Alibaba (Qwen) | 1,099.33 1,094–1,104 | — | — | 8K |
| 22 | FLUX.1 Kontext Max | Black Forest Labs | 1,074.07 1,071–1,077 | $0.08 | $0.08 | 1K |
| 23 | FLUX.2 [klein] 9B ≈ | Black Forest Labs | 1,069.32 1,066–1,073 | — | — | — |
| 24 | Flux.1 Kontext Pro | Black Forest Labs | 1,058.75 1,055–1,062 | — | — | — |
| 25 | Imagen 3.0 Generate 002 ≈ | Google DeepMind | 1,058.22 1,055–1,061 | — | — | — |
| 26 | Qwen Image ≈ | Alibaba (Qwen) | 1,056.67 1,054–1,059 | $0.5 | $2 | 8K |
| 27 | FLUX.2 Klein 4B | Black Forest Labs | 1,029.92 1,027–1,033 | $1 | $1 | 128K |
| 28 | Recraft V3 | Recraft AI | 1,021.35 1,018–1,025 | — | — | 1K |
| 29 | Flux 1.1 Pro ≈ | Black Forest Labs | 1,015.77 1,012–1,019 | — | — | — |
| 30 | Ideogram V2 ≈ | Ideogram | 1,013.48 1,010–1,017 | — | — | 0K |
| 31 | FLUX.1 Dev | Black Forest Labs | 969.54 966–973 | $0 | $0 | 4K |
| 32 | DALL E 3 ≈ | OpenAI | 967.7 964–971 | — | — | 1K |
| 33 | FLUX.1 Kontext Dev | Black Forest Labs | 940.42 937–944 | — | — | 41K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 19 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is 7.3 points, and that is narrower than the gap to the neighbouring row: the order here is set by the data, not by who happened to vote this time. Votes per model here — median 87,149: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)