mistral-medium-2505| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Azure Cognitive Services | $0.4 | $2 | 31 Jul 2026 |
| Merge Gateway | $0.4 | $2 | 31 Jul 2026 |
| Microsoft Foundry | $0.4 | $2 | 31 Jul 2026 |
| Mistral AI | $0.4 | $2 | 31 Jul 2026 |
| IBM watsonx.ai | $3 | $10 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, programming | 1,386.17 1,378–1,394 | — | — |
| Arena Score in English | 1,384.32 1,379–1,390 | — | — |
| Arena Score, multi-turn dialogue | 1,383.26 1,375–1,391 | — | — |
| Arena Score in French | 1,370.05 1,342–1,398 | — | — |
| Arena Score, overall | 1,369.43 1,365–1,374 | — | — |
| Arena Score in Spanish | 1,365.42 1,338–1,393 | — | — |
| Arena Score, hard prompts | 1,364.71 1,359–1,371 | — | — |
| Arena Score, long queries | 1,358.37 1,350–1,366 | — | — |
| Arena Score in Russian | 1,357.45 1,346–1,369 | — | — |
| Arena Score, mathematics | 1,351.89 1,340–1,364 | — | — |
| Arena Score, creative writing | 1,343.78 1,334–1,353 | — | — |
| Arena Score, expert questions | 1,343.67 1,330–1,357 | — | — |
| Arena Score, instruction following | 1,338.98 1,332–1,346 | — | — |
| Arena Score, working with images | 1,157.54 1,149–1,166 | — | — |
| Arena Score, understanding diagrams | 1,154.86 1,139–1,171 | — | — |
| Arena Score, text recognition in images | 1,153.86 1,144–1,164 | — | — |
| MATH, difficulty level five | 81.63 % | · with a tuned harness | Epoch evaluations |
| Creative writing (Lech Mazur’s evaluation) | 77.3 % | — | lechmazur/writing Github repository |
| GPQA Diamond — graduate-level questions | 46.04 % | · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 32.15 % | · with a tuned harness | Epoch evaluations |
| Humanity’s Last Exam — expert-level questions | 0 % | — | — |