azure_ai/gpt-5.4-pro-2026-03-05| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Microsoft Foundry | $30 | $180 | 31 Jul 2026 |
| OpenAI | $30 | $180 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| ARC-AGI — generalising to unseen patterns | 94.5 % | effort: xhigh · with a tuned harness | — |
| GPQA Diamond — graduate-level questions | 92.8 % | effort: xhigh · with a tuned harness | Epoch evaluations |
| ARC-AGI-2 | 83.33 % | effort: xhigh | — |
| FrontierMath, levels 1–3 | 82.46 % | effort: xhigh | Epoch evaluations |
| SimpleBench — trick questions | 68.92 % | — | SimpleBench Leaderboard |
| FrontierMath, level 4 — research-grade problems | 58.54 % | effort: xhigh | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 57.44 % | effort: none | https://htihle.github.io/weirdml.html |
| Chess puzzles | 56.44 % | effort: xhigh | Epoch evaluations |
| SimpleQA Verified — factual accuracy | 47.8 % | effort: xhigh | Epoch evaluations |
| Humanity’s Last Exam — expert-level questions | 41.51 % | — | — |
| CritPt — physics problems | 30 % | effort: xhigh | — |