azure/gpt-4.5-preview| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Microsoft Foundry | $75 | $150 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, multi-turn dialogue | 1,443.92 1,430–1,457 | — | — |
| Arena Score in English | 1,419.49 1,412–1,427 | — | — |
| Arena Score in Russian | 1,418.24 1,403–1,434 | — | — |
| Arena Score in French | 1,418.01 1,362–1,474 | — | — |
| Arena Score, overall | 1,417.47 1,412–1,423 | — | — |
| Arena Score, mathematics | 1,411.89 1,397–1,427 | — | — |
| Arena Score, long queries | 1,406 1,392–1,420 | — | — |
| Arena Score, instruction following | 1,404.35 1,396–1,413 | — | — |
| Arena Score, hard prompts | 1,403.65 1,394–1,414 | — | — |
| Arena Score, programming | 1,397.05 1,384–1,410 | — | — |
| Arena Score, creative writing | 1,394.29 1,382–1,406 | — | — |
| Arena Score, expert questions | 1,393.91 1,371–1,417 | — | — |
| Arena Score, working with images | 1,195.13 1,183–1,207 | — | — |
| MATH, difficulty level five | 78.63 % | · with a tuned harness | Epoch evaluations |
| Creative writing (Lech Mazur’s evaluation) | 75.6 % | — | lechmazur/writing Github repository |
| Fiction.LiveBench — holding a long context | 63.9 % | — | Fiction.live leaderboard |
| GPQA Diamond — graduate-level questions | 58.25 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 44.9 % | · with a tuned harness | Aider LLM Leaderboards |
| WeirdML — unusual machine learning tasks | 39.37 % | — | WeirdML Leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 37.72 % | · with a tuned harness | Epoch evaluations |
| SimpleBench — trick questions | 21.4 % | — | SimpleBench Leaderboard |
| Cybench — cybersecurity tasks | 17.5 % | · with a tuned harness | Cybench leaderboard |
| ARC-AGI — generalising to unseen patterns | 10.3 % | · with a tuned harness | ARC Prize Leaderboard |
| ARC-AGI-2 | 0.8 % | — | — |
| Humanity’s Last Exam — expert-level questions | 0.67 % | — | — |