xai/grok-3-mini| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Microsoft Foundry | $0.25 | $1.27 | 31 Jul 2026 |
| vercel_ai_gateway | $0.3 | $0.5 | 31 Jul 2026 |
| Helicone | $0.3 | $0.5 | 31 Jul 2026 |
| OCI Generative AI | $0.3 | $0.5 | 31 Jul 2026 |
| Poe | $0.3 | $0.5 | 31 Jul 2026 |
| xAI | $0.3 | $0.5 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in English | 1,373.29 1,367–1,380 | — | — |
| Arena Score, mathematics | 1,370.22 1,356–1,384 | — | — |
| Arena Score, programming | 1,366.03 1,357–1,375 | — | — |
| Arena Score, hard prompts | 1,363.32 1,357–1,370 | — | — |
| Arena Score, overall | 1,362.74 1,358–1,368 | — | — |
| Arena Score, expert questions | 1,360.62 1,344–1,378 | — | — |
| Arena Score, long queries | 1,356.9 1,348–1,366 | — | — |
| Arena Score in Spanish | 1,355.7 1,322–1,390 | effort: high | — |
| Arena Score in French | 1,352.74 1,308–1,398 | effort: high | — |
| Arena Score in Russian | 1,350.69 1,335–1,366 | — | — |
| Arena Score, multi-turn dialogue | 1,349.13 1,340–1,359 | — | — |
| Arena Score, instruction following | 1,345.63 1,338–1,354 | — | — |
| Arena Score, creative writing | 1,330.28 1,317–1,343 | effort: high | — |
| MATH, difficulty level five | 90.94 % | effort: low · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 77.76 % | effort: low · with a tuned harness | Epoch evaluations |
| Creative writing (Lech Mazur’s evaluation) | 73.5 % | effort: low | lechmazur/writing Github repository |
| GPQA Diamond — graduate-level questions | 68.35 % | effort: high · with a tuned harness | Epoch evaluations |
| Fiction.LiveBench — holding a long context | 66.7 % | effort: medium | Fiction.live leaderboard |
| Aider Polyglot — code edits in six languages | 49.3 % | effort: low · with a tuned harness | Aider LLM Leaderboards |
| WeirdML — unusual machine learning tasks | 42.58 % | effort: high | WeirdML Leaderboard |
| SimpleQA Verified — factual accuracy | 21.1 % | effort: high | Epoch evaluations |
| ARC-AGI — generalising to unseen patterns | 16.5 % | effort: low · with a tuned harness | ARC Prize Leaderboard |
| ARC-AGI-2 | 0.42 % | effort: low | — |