xAI

Grok 3 Mini

Type
language model
Context
131K tokens
Max output
8K
Released
11 April 2025
API string
xai/grok-3-mini

Prices by provider

Provider Input, $ per 1M Output, $ per 1M In our data since
Microsoft Foundry $0.25 $1.27 31 Jul 2026
vercel_ai_gateway $0.3 $0.5 31 Jul 2026
Helicone $0.3 $0.5 31 Jul 2026
OCI Generative AI $0.3 $0.5 31 Jul 2026
Poe $0.3 $0.5 31 Jul 2026
xAI $0.3 $0.5 31 Jul 2026

Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.

The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.

Measurement results

Task set Result Run conditions Measured by
Arena Score in English 1,373.29 1,367–1,380
Arena Score, mathematics 1,370.22 1,356–1,384
Arena Score, programming 1,366.03 1,357–1,375
Arena Score, hard prompts 1,363.32 1,357–1,370
Arena Score, overall 1,362.74 1,358–1,368
Arena Score, expert questions 1,360.62 1,344–1,378
Arena Score, long queries 1,356.9 1,348–1,366
Arena Score in Spanish 1,355.7 1,322–1,390 effort: high
Arena Score in French 1,352.74 1,308–1,398 effort: high
Arena Score in Russian 1,350.69 1,335–1,366
Arena Score, multi-turn dialogue 1,349.13 1,340–1,359
Arena Score, instruction following 1,345.63 1,338–1,354
Arena Score, creative writing 1,330.28 1,317–1,343 effort: high
MATH, difficulty level five 90.94 % effort: low · with a tuned harness Epoch evaluations
Mock AIME 2024–2025 — olympiad problems 77.76 % effort: low · with a tuned harness Epoch evaluations
Creative writing (Lech Mazur’s evaluation) 73.5 % effort: low lechmazur/writing Github repository
GPQA Diamond — graduate-level questions 68.35 % effort: high · with a tuned harness Epoch evaluations
Fiction.LiveBench — holding a long context 66.7 % effort: medium Fiction.live leaderboard
Aider Polyglot — code edits in six languages 49.3 % effort: low · with a tuned harness Aider LLM Leaderboards
WeirdML — unusual machine learning tasks 42.58 % effort: high WeirdML Leaderboard
SimpleQA Verified — factual accuracy 21.1 % effort: high Epoch evaluations
ARC-AGI — generalising to unseen patterns 16.5 % effort: low · with a tuned harness ARC Prize Leaderboard
ARC-AGI-2 0.42 % effort: low