openai/gpt-4o-mini-2024-07-18| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Kilo Code | $0.15 | $0.6 | 31 Jul 2026 |
| OpenAI | $0.15 | $0.6 | 31 Jul 2026 |
| OpenRouter | $0.15 | $0.6 | 31 Jul 2026 |
| Microsoft Foundry | $0.165 | $0.66 | 31 Jul 2026 |
| Venice AI | $0.188 | $0.75 | 1 Aug 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in English | 1,300.59 1,296–1,305 | — | — |
| Arena Score in French | 1,295.68 1,271–1,320 | — | — |
| Arena Score, programming | 1,290.12 1,284–1,297 | — | — |
| Arena Score, long queries | 1,288.94 1,282–1,296 | — | — |
| Arena Score, overall | 1,286.57 1,283–1,290 | — | — |
| Arena Score, multi-turn dialogue | 1,285.04 1,278–1,292 | — | — |
| Arena Score in Spanish | 1,278.34 1,255–1,302 | — | — |
| Arena Score in Russian | 1,274.29 1,267–1,282 | — | — |
| Arena Score, creative writing | 1,268.26 1,261–1,275 | — | — |
| Arena Score, mathematics | 1,267.33 1,260–1,274 | — | — |
| Arena Score, hard prompts | 1,267.19 1,262–1,272 | — | — |
| Arena Score, instruction following | 1,258.81 1,254–1,264 | — | — |
| Arena Score, expert questions | 1,233.87 1,223–1,245 | — | — |
| Arena Score, working with images | 1,065.76 1,058–1,074 | — | — |
| Creative writing (Lech Mazur’s evaluation) | 67.2 % | — | lechmazur/writing Github repository |
| GeoBench — locating a place from a photograph | 64 % | — | GeoBench leaderboard |
| MATH, difficulty level five | 52.63 % | · with a tuned harness | Epoch evaluations |
| BALROG — game environments | 17.4 % | — | Balrog Leaderboard |
| GPQA Diamond — graduate-level questions | 16.96 % | · with a tuned harness | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 11.76 % | — | WeirdML Leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 6.85 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 3.6 % | · with a tuned harness | Aider LLM Leaderboards |
| ARC-AGI-2 | 0 % | — | — |
| SimpleBench — trick questions | 0 % | — | SimpleBench Leaderboard |
| Chess puzzles | 0 % | — | Epoch evaluations |