xAI

Grok 4

Type
language model
Context
256K tokens
Max output
16K
Released
9 July 2025
API string
grok-4-0709

Prices by provider

Provider Input, $ per 1M Output, $ per 1M In our data since
Jiekou.AI $2.7 $13.5 31 Jul 2026
Abacus.AI $3 $15 31 Jul 2026
xAI $3 $15 31 Jul 2026

Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.

The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.

Measurement results

Task set Result Run conditions Measured by
Arena Score in French 1,425.25 1,399–1,452
Arena Score, mathematics 1,424.24 1,412–1,437
Arena Score in English 1,418.4 1,413–1,423
Arena Score, expert questions 1,417.67 1,404–1,431
Arena Score, multi-turn dialogue 1,415.59 1,408–1,423
Arena Score in Spanish 1,413.5 1,394–1,433
Arena Score in Russian 1,412.96 1,401–1,425
Arena Score, long queries 1,410.19 1,404–1,417
Arena Score, overall 1,409.91 1,406–1,414
Arena Score, programming 1,408.98 1,402–1,416
Arena Score, hard prompts 1,408.55 1,404–1,414
Arena Score, creative writing 1,398.43 1,390–1,407
Arena Score, instruction following 1,387.31 1,381–1,393
Arena Score, working with images 1,208.77 1,201–1,216
Arena Score, understanding diagrams 1,203.03 1,190–1,216
Arena Score, text recognition in images 1,194.52 1,186–1,203
Fiction.LiveBench — holding a long context 94.4 % Fiction.live leaderboard
Mock AIME 2024–2025 — olympiad problems 83.98 % · with a tuned harness Epoch evaluations
GPQA Diamond — graduate-level questions 82.67 % · with a tuned harness Epoch evaluations
Aider Polyglot — code edits in six languages 79.6 % · with a tuned harness Aider LLM Leaderboards
Creative writing (Lech Mazur’s evaluation) 76.9 % lechmazur/writing Github repository
ARC-AGI — generalising to unseen patterns 66.67 % · with a tuned harness
SimpleBench — trick questions 52.6 % SimpleBench Leaderboard
SimpleQA Verified — factual accuracy 47.9 % Epoch evaluations
DeepResearch Bench — deep research 47.9 % DeepResearchBench Leaderboard
WeirdML — unusual machine learning tasks 45.73 % WeirdML Leaderboard
GeoBench — locating a place from a photograph 45 % GeoBench leaderboard
BALROG — game environments 43.6 % Balrog Leaderboard
Cybench — cybersecurity tasks 43 % · with a tuned harness Grok 4 Model Card self-reported
Terminal-Bench — working in the command line 27.2 % · with a tuned harness Terminal-Bench v2 Leaderboard
Chess puzzles 24.24 % Epoch evaluations
GDPval — tasks from real occupations 21.1 % effort: high
ARC-AGI-2 15.98 %
APEX-Agents 15.2 %