gpt-3.5-turbo-1106The provider has deprecated this model. It still responds, but it is not a good choice for new projects.
| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Azure Cognitive Services | $1 | $2 | 31 Jul 2026 |
| Microsoft Foundry | $1 | $2 | 31 Jul 2026 |
| OpenAI | $1 | $2 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, mathematics | 1,141.27 1,126–1,156 | — | — |
| Arena Score, programming | 1,116.27 1,101–1,132 | — | — |
| Arena Score in English | 1,113.21 1,104–1,123 | — | — |
| Arena Score, hard prompts | 1,098.86 1,086–1,112 | — | — |
| Arena Score, overall | 1,094.28 1,086–1,103 | — | — |
| Arena Score, instruction following | 1,093.11 1,081–1,105 | — | — |
| Arena Score in Spanish | 1,091.94 1,051–1,133 | — | — |
| Arena Score, multi-turn dialogue | 1,077.91 1,062–1,094 | — | — |
| Arena Score in Russian | 1,076.36 1,047–1,105 | — | — |
| Arena Score, long queries | 1,069.93 1,049–1,091 | — | — |
| Arena Score, expert questions | 1,069.17 1,041–1,097 | — | — |
| Arena Score in French | 1,065.08 1,026–1,104 | — | — |
| Arena Score, creative writing | 1,035.47 1,020–1,051 | — | — |
| MATH, difficulty level five | 15.89 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 4.04 % | · with a tuned harness | Epoch evaluations |