openai/gpt-4o-2024-05-13| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Kilo Code | $5 | $15 | 31 Jul 2026 |
| Merge Gateway | $5 | $15 | 31 Jul 2026 |
| Microsoft Foundry | $5 | $15 | 31 Jul 2026 |
| OpenAI | $5 | $15 | 31 Jul 2026 |
| OpenRouter | $5 | $15 | 31 Jul 2026 |
| OrcaRouter | $5 | $15 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in English | 1,312.52 1,308–1,317 | — | — |
| Arena Score in French | 1,302.84 1,282–1,324 | — | — |
| Arena Score, multi-turn dialogue | 1,301.69 1,295–1,308 | — | — |
| Arena Score, overall | 1,300.6 1,297–1,304 | — | — |
| Arena Score, programming | 1,297.63 1,291–1,304 | — | — |
| Arena Score, creative writing | 1,291.85 1,285–1,299 | — | — |
| Arena Score in Spanish | 1,291.38 1,271–1,311 | — | — |
| Arena Score, long queries | 1,288.84 1,282–1,296 | — | — |
| Arena Score in Russian | 1,284.6 1,277–1,292 | — | — |
| Arena Score, mathematics | 1,284.4 1,278–1,291 | — | — |
| Arena Score, hard prompts | 1,281.1 1,276–1,286 | — | — |
| Arena Score, instruction following | 1,278.12 1,273–1,283 | — | — |
| Arena Score, expert questions | 1,249.54 1,239–1,260 | — | — |
| Arena Score, working with images | 1,136.75 1,128–1,145 | — | — |
| MATH, difficulty level five | 51.05 % | · with a tuned harness | Epoch evaluations |
| BALROG — game environments | 32.3 % | — | Balrog Leaderboard |
| GPQA Diamond — graduate-level questions | 31.86 % | · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 6.16 % | · with a tuned harness | Epoch evaluations |