azure/gpt-5.2-2025-12-11| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Microsoft Foundry | $1.75 | $14 | 31 Jul 2026 |
| OpenAI | $1.75 | $14 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Mock AIME 2024–2025 — olympiad problems | 96.11 % | effort: none · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 88.53 % | effort: none · with a tuned harness | Epoch evaluations |
| ARC-AGI — generalising to unseen patterns | 86.2 % | effort: xhigh · with a tuned harness | ARC Prize Leaderboard |
| SWE-bench Verified — fixing bugs in repositories | 73.76 % | effort: high · with a tuned harness | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 72.2 % | effort: xhigh | WeirdML Leaderboard |
| FrontierMath, levels 1–3 | 67.4 % | effort: xhigh | Epoch evaluations |
| Terminal-Bench — working in the command line | 64.9 % | effort: medium · with a tuned harness | Terminal-Bench v2 Leaderboard |
| ARC-AGI-2 | 52.91 % | effort: xhigh | — |
| GDPval — tasks from real occupations | 49.7 % | effort: none | — |
| Chess puzzles | 46.34 % | effort: none | Epoch evaluations |
| DeepResearch Bench — deep research | 41.12 % | effort: low | https://drb.futuresearch.ai/#drb self-reported |
| SimpleQA Verified — factual accuracy | 38.9 % | effort: medium | Epoch evaluations |
| SimpleBench — trick questions | 34.96 % | effort: high | SimpleBench Leaderboard |
| APEX-Agents | 34.4 % | effort: xhigh | — |
| FrontierMath, level 4 — research-grade problems | 31.7 % | effort: xhigh | Epoch evaluations |
| GSO-Bench — code optimisation | 27.4 % | effort: high | GSO Leaderboard |
| Humanity’s Last Exam — expert-level questions | 24.16 % | — | — |
| Remote Labor Index — jobs from a freelance marketplace | 2.5 % | effort: medium | — |