The same blind comparison as for images, but for video. There are an order of magnitude fewer models here than in the text collections, and the top places stand apart from the rest more visibly — simply because there are few participants.
| # | Model | Developer | Arena Score, video | Sign in$ / 1M | Output$ / 1M | Contexttokens |
|---|---|---|---|---|---|---|
| 1 | Sora 2 Pro | OpenAI | 1,365.62 1,358–1,374 | — | — | — |
| 2 | Sora 2 | OpenAI | 1,337.11 1,330–1,344 | — | — | — |
| 3 | Seedance v1.5 Pro | ByteDance (Doubao) | 1,256.99 1,250–1,264 | — | — | — |
| 4 | Veo 3 ≈ | Google DeepMind | 1,252.4 1,241–1,263 | — | — | 0K |
| 5 | Veo 3 Fast ≈ | Google DeepMind | 1,247.67 1,236–1,259 | — | — | 0K |
| 6 | Veo 2 | Google DeepMind | 1,162.57 1,146–1,179 | — | — | 0K |
| 7 | Ray2 | Luma AI | 1,063.9 1,047–1,081 | — | — | 5K |
The table scrolls sideways: not all columns fit.
The price is the lowest among the model's providers, base tier, without batch or discounted rates. Input and output are shown separately on purpose: for most models the output costs several times more than the input, and the final bill depends on which of the two your task has more of.
The ≈ sign marks models whose confidence interval overlaps that of the row above. Here there are 2 models — their order among themselves is not determined by the available data.
The confidence interval for half the models is 22.1 points, and that is narrower than the gap to the neighbouring row: the order here is set by the data, not by who happened to vote this time. Votes per model here — median 15,230: the fewer there are, the wider the interval.
Arena (LMArena)
— license CC-BY-4.0 · primary source
По данным Arena (LMArena)