gemini/gemini-3.1-pro-preview| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| vivgrid | $2 | $12 | 31 Jul 2026 |
| daoxe | $2 | $12 | 31 Jul 2026 |
| frogbot | $2 | $12 | 31 Jul 2026 |
| auriko | $2 | $12 | 31 Jul 2026 |
| ofox | $2 | $12 | 31 Jul 2026 |
| perplexity-agent | $2 | $12 | 31 Jul 2026 |
| AIHubMix | $2 | $12 | 31 Jul 2026 |
| Abacus.AI | $2 | $12 | 31 Jul 2026 |
| CrossModel | $2 | $12 | 31 Jul 2026 |
| FastRouter | $2 | $12 | 31 Jul 2026 |
| GitHub Copilot | $2 | $12 | 31 Jul 2026 |
| Google DeepMind | $2 | $12 | 31 Jul 2026 |
| Google Vertex AI | $2 | $12 | 31 Jul 2026 |
| Kilo Code | $2 | $12 | 31 Jul 2026 |
| LLM Gateway | $2 | $12 | 31 Jul 2026 |
| Merge Gateway | $2 | $12 | 31 Jul 2026 |
| NanoGPT | $2 | $12 | 31 Jul 2026 |
| OpenRouter | $2 | $12 | 31 Jul 2026 |
| Vercel | $2 | $12 | 31 Jul 2026 |
| ZenMux | $2 | $12 | 31 Jul 2026 |
| Venice AI | $2.5 | $15 | 31 Jul 2026 |
| OrcaRouter | $4 | $18 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in Russian | 1,495.77 1,488–1,503 | — | — |
| Arena Score, mathematics | 1,489.07 1,480–1,498 | — | — |
| Arena Score in French | 1,488.01 1,474–1,502 | — | — |
| Arena Score, expert questions | 1,486.42 1,479–1,494 | — | — |
| Arena Score, hard prompts | 1,484.41 1,480–1,489 | — | — |
| Arena Score, programming | 1,483.89 1,479–1,489 | — | — |
| Arena Score, multi-turn dialogue | 1,482.41 1,476–1,489 | — | — |
| Arena Score, long queries | 1,482.3 1,477–1,487 | — | — |
| Arena Score, creative writing | 1,481.61 1,475–1,488 | — | — |
| Arena Score in Spanish | 1,479.93 1,467–1,493 | — | — |
| Arena Score, overall | 1,479.46 1,476–1,483 | — | — |
| Arena Score in English | 1,478.03 1,474–1,483 | — | — |
| Arena Score, instruction following | 1,467.43 1,462–1,473 | — | — |
| Arena Score, web development | 1,446.83 1,441–1,452 | — | — |
| Arena Score, understanding diagrams | 1,312.53 1,305–1,320 | — | — |
| Arena Score, text recognition in images | 1,305.79 1,300–1,311 | — | — |
| Arena Score, working with images | 1,295.15 1,289–1,301 | — | — |
| ARC-AGI — generalising to unseen patterns | 98 % | · with a tuned harness | ARC Prize Leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 95.6 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 92.13 % | · with a tuned harness | Epoch evaluations |
| Terminal-Bench — working in the command line | 80.2 % | · with a tuned harness | https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| SimpleQA Verified — factual accuracy | 77.3 % | — | Epoch evaluations |
| ARC-AGI-2 | 77.1 % | — | — |
| SimpleBench — trick questions | 75.52 % | — | SimpleBench Leaderboard |
| WeirdML — unusual machine learning tasks | 72.1 % | — | WeirdML Leaderboard |
| FrontierMath, levels 1–3 | 59.65 % | — | Epoch evaluations |
| BALROG — game environments | 57 % | — | Balrog Leaderboard |
| Chess puzzles | 52.65 % | — | Epoch evaluations |
| Humanity’s Last Exam — expert-level questions | 43.74 % | — | — |
| APEX-Agents | 33.5 % | — | — |
| FrontierMath, level 4 — research-grade problems | 26.83 % | — | Epoch evaluations |
| GSO-Bench — code optimisation | 22.55 % | — | https://gso-bench.github.io/index.html |
| CritPt — physics problems | 17.71 % | — | — |
| Arena Score, agent tasks | -0.01 0–0 | — | — |