gpt-3.5-turbo-0125The provider has deprecated this model. It still responds, but it is not a good choice for new projects.
| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Azure Cognitive Services | $0.5 | $1.5 | 31 Jul 2026 |
| Microsoft Foundry | $0.5 | $1.5 | 31 Jul 2026 |
| OpenAI | $0.5 | $1.5 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, mathematics | 1,141.78 1,133–1,150 | — | — |
| Arena Score in English | 1,139.35 1,134–1,145 | — | — |
| Arena Score, programming | 1,136.9 1,129–1,145 | — | — |
| Arena Score, overall | 1,125.08 1,120–1,130 | — | — |
| Arena Score in Russian | 1,122.31 1,112–1,133 | — | — |
| Arena Score, long queries | 1,121.08 1,112–1,130 | — | — |
| Arena Score in Spanish | 1,119.73 1,095–1,145 | — | — |
| Arena Score, instruction following | 1,119.07 1,113–1,125 | — | — |
| Arena Score, multi-turn dialogue | 1,116.86 1,108–1,125 | — | — |
| Arena Score in French | 1,116.55 1,094–1,140 | — | — |
| Arena Score, hard prompts | 1,108.19 1,102–1,115 | — | — |
| Arena Score, creative writing | 1,092.78 1,084–1,102 | — | — |
| Arena Score, expert questions | 1,065.04 1,052–1,078 | — | — |
| MATH, difficulty level five | 11.63 % | · with a tuned harness | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 3.48 % | — | WeirdML Leaderboard |
| GPQA Diamond — graduate-level questions | 2.9 % | · with a tuned harness | Epoch evaluations |