| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Microsoft Foundry | $30 | $60 | 31 Jul 2026 |
| OpenAI | $30 | $60 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, mathematics | 1,217 1,209–1,225 | — | — |
| Arena Score in English | 1,205.43 1,200–1,210 | — | — |
| Arena Score, creative writing | 1,191.84 1,184–1,200 | — | — |
| Arena Score, instruction following | 1,189.04 1,183–1,195 | — | — |
| Arena Score, long queries | 1,188.09 1,179–1,197 | — | — |
| Arena Score, programming | 1,187.88 1,180–1,196 | — | — |
| Arena Score, overall | 1,186.17 1,182–1,190 | — | — |
| Arena Score, multi-turn dialogue | 1,184.73 1,176–1,193 | — | — |
| Arena Score, hard prompts | 1,174.66 1,168–1,181 | — | — |
| Arena Score in Russian | 1,172.16 1,162–1,182 | — | — |
| Arena Score in French | 1,168.83 1,147–1,191 | — | — |
| Arena Score in Spanish | 1,167.17 1,145–1,190 | — | — |
| Arena Score, expert questions | 1,127.46 1,115–1,140 | — | — |
| MATH, difficulty level five | 22.97 % | · with a tuned harness | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 12.44 % | — | WeirdML Leaderboard |
| GPQA Diamond — graduate-level questions | 7.53 % | · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 1.01 % | · with a tuned harness | Epoch evaluations |