xai/grok-3| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| vercel_ai_gateway | $3 | $15 | 31 Jul 2026 |
| Helicone | $3 | $15 | 31 Jul 2026 |
| Microsoft Foundry | $3 | $15 | 31 Jul 2026 |
| OCI Generative AI | $3 | $15 | 31 Jul 2026 |
| Poe | $3 | $15 | 31 Jul 2026 |
| xAI | $3 | $15 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| MATH, difficulty level five | 88.75 % | · with a tuned harness | Epoch evaluations |
| Creative writing (Lech Mazur’s evaluation) | 76.4 % | — | lechmazur/writing Github repository |
| GPQA Diamond — graduate-level questions | 67.68 % | · with a tuned harness | Epoch evaluations |
| Fiction.LiveBench — holding a long context | 58.3 % | — | Fiction.live leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 55.51 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 53.3 % | · with a tuned harness | Aider LLM Leaderboards |
| WeirdML — unusual machine learning tasks | 37.24 % | — | WeirdML Leaderboard |
| BALROG — game environments | 29.5 % | — | Balrog Leaderboard |
| SimpleBench — trick questions | 23.32 % | — | SimpleBench Leaderboard |
| ARC-AGI — generalising to unseen patterns | 5.5 % | · with a tuned harness | ARC Prize Leaderboard |
| APEX-Agents | 2.1 % | — | — |
| ARC-AGI-2 | 0 % | — | — |