x-ai/grok-4.5| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| opencode-go | $2 | $6 | 31 Jul 2026 |
| daoxe | $2 | $6 | 31 Jul 2026 |
| AIHubMix | $2 | $6 | 31 Jul 2026 |
| Abacus.AI | $2 | $6 | 31 Jul 2026 |
| CrossModel | $2 | $6 | 31 Jul 2026 |
| LLM Gateway | $2 | $6 | 31 Jul 2026 |
| Merge Gateway | $2 | $6 | 31 Jul 2026 |
| OpenCode Zen | $2 | $6 | 31 Jul 2026 |
| OpenRouter | $2 | $6 | 31 Jul 2026 |
| Vercel | $2 | $6 | 31 Jul 2026 |
| ZenMux | $2 | $6 | 31 Jul 2026 |
| xAI | $2 | $6 | 31 Jul 2026 |
| Venice AI | $2.27 | $6.8 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, web development | 1,549.24 1,539–1,560 | — | — |
| Arena Score, mathematics | 1,484.15 1,456–1,513 | — | — |
| Arena Score, programming | 1,479.12 1,468–1,491 | — | — |
| Arena Score, expert questions | 1,477.5 1,459–1,496 | — | — |
| Arena Score, long queries | 1,465.65 1,456–1,475 | — | — |
| Arena Score, hard prompts | 1,463.7 1,456–1,472 | — | — |
| Arena Score, multi-turn dialogue | 1,457.65 1,443–1,472 | — | — |
| Arena Score in Russian | 1,456.76 1,439–1,475 | — | — |
| Arena Score in English | 1,455.92 1,446–1,465 | — | — |
| Arena Score in French | 1,454.67 1,422–1,488 | — | — |
| Arena Score, overall | 1,451.46 1,445–1,458 | — | — |
| Arena Score in Spanish | 1,449.5 1,412–1,487 | — | — |
| Arena Score, instruction following | 1,442.43 1,432–1,453 | — | — |
| Arena Score, creative writing | 1,439.21 1,424–1,454 | — | — |
| Arena Score, understanding diagrams | 1,325.46 1,304–1,347 | — | — |
| Arena Score, text recognition in images | 1,300.33 1,287–1,314 | — | — |
| Arena Score, working with images | 1,293.37 1,281–1,306 | — | — |
| Mock AIME 2024–2025 — olympiad problems | 97.78 % | effort: high · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 91.25 % | effort: high · with a tuned harness | Epoch evaluations |
| ARC-AGI — generalising to unseen patterns | 87.17 % | effort: medium · with a tuned harness | https://arcprize.org/leaderboard |
| SimpleBench — trick questions | 64 % | — | SimpleBench Leaderboard |
| FrontierMath, levels 1–3 | 57.19 % | effort: high | Epoch evaluations |
| SimpleQA Verified — factual accuracy | 53.5 % | effort: high | Epoch evaluations |
| ARC-AGI-2 | 52.64 % | effort: high | — |
| WeirdML — unusual machine learning tasks | 46.43 % | — | https://htihle.github.io/weirdml.html |
| FrontierCode — patches fit to be merged into a project | 42.4 % | — | https://cognition.com/frontiercode |
| APEX-Agents | 34.2 % | — | — |
| Chess puzzles | 32.66 % | effort: high | Epoch evaluations |
| FrontierMath, level 4 — research-grade problems | 24.39 % | effort: high | Epoch evaluations |
| CritPt — physics problems | 15.43 % | effort: high | — |
| Arena Score, agent tasks | 0.06 0–0 | — | — |