grok-4-0709| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Jiekou.AI | $2.7 | $13.5 | 31 Jul 2026 |
| Abacus.AI | $3 | $15 | 31 Jul 2026 |
| xAI | $3 | $15 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in French | 1,425.25 1,399–1,452 | — | — |
| Arena Score, mathematics | 1,424.24 1,412–1,437 | — | — |
| Arena Score in English | 1,418.4 1,413–1,423 | — | — |
| Arena Score, expert questions | 1,417.67 1,404–1,431 | — | — |
| Arena Score, multi-turn dialogue | 1,415.59 1,408–1,423 | — | — |
| Arena Score in Spanish | 1,413.5 1,394–1,433 | — | — |
| Arena Score in Russian | 1,412.96 1,401–1,425 | — | — |
| Arena Score, long queries | 1,410.19 1,404–1,417 | — | — |
| Arena Score, overall | 1,409.91 1,406–1,414 | — | — |
| Arena Score, programming | 1,408.98 1,402–1,416 | — | — |
| Arena Score, hard prompts | 1,408.55 1,404–1,414 | — | — |
| Arena Score, creative writing | 1,398.43 1,390–1,407 | — | — |
| Arena Score, instruction following | 1,387.31 1,381–1,393 | — | — |
| Arena Score, working with images | 1,208.77 1,201–1,216 | — | — |
| Arena Score, understanding diagrams | 1,203.03 1,190–1,216 | — | — |
| Arena Score, text recognition in images | 1,194.52 1,186–1,203 | — | — |
| Fiction.LiveBench — holding a long context | 94.4 % | — | Fiction.live leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 83.98 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 82.67 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 79.6 % | · with a tuned harness | Aider LLM Leaderboards |
| Creative writing (Lech Mazur’s evaluation) | 76.9 % | — | lechmazur/writing Github repository |
| ARC-AGI — generalising to unseen patterns | 66.67 % | · with a tuned harness | — |
| SimpleBench — trick questions | 52.6 % | — | SimpleBench Leaderboard |
| SimpleQA Verified — factual accuracy | 47.9 % | — | Epoch evaluations |
| DeepResearch Bench — deep research | 47.9 % | — | DeepResearchBench Leaderboard |
| WeirdML — unusual machine learning tasks | 45.73 % | — | WeirdML Leaderboard |
| GeoBench — locating a place from a photograph | 45 % | — | GeoBench leaderboard |
| BALROG — game environments | 43.6 % | — | Balrog Leaderboard |
| Cybench — cybersecurity tasks | 43 % | · with a tuned harness | Grok 4 Model Card self-reported |
| Terminal-Bench — working in the command line | 27.2 % | · with a tuned harness | Terminal-Bench v2 Leaderboard |
| Chess puzzles | 24.24 % | — | Epoch evaluations |
| GDPval — tasks from real occupations | 21.1 % | effort: high | — |
| ARC-AGI-2 | 15.98 % | — | — |
| APEX-Agents | 15.2 % | — | — |