gemini-2.0-flash-001| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| DeepInfra | $0.1 | $0.4 | 31 Jul 2026 |
| Google DeepMind | $0.1 | $0.4 | 31 Jul 2026 |
| Kilo Code | $0.1 | $0.4 | 31 Jul 2026 |
| OpenRouter | $0.1 | $0.4 | 31 Jul 2026 |
| NanoGPT | $0.1 | $0.408 | 31 Jul 2026 |
| Google Vertex AI | $0.15 | $0.6 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in French | 1,389.86 1,363–1,416 | — | — |
| Arena Score in English | 1,365.36 1,361–1,370 | — | — |
| Arena Score in Spanish | 1,360.8 1,335–1,386 | — | — |
| Arena Score, overall | 1,354.17 1,350–1,358 | — | — |
| Arena Score in Russian | 1,352.75 1,343–1,363 | — | — |
| Arena Score, programming | 1,351.72 1,345–1,359 | — | — |
| Arena Score, mathematics | 1,351.68 1,343–1,361 | — | — |
| Arena Score, multi-turn dialogue | 1,349.91 1,343–1,357 | — | — |
| Arena Score, hard prompts | 1,346.49 1,341–1,352 | — | — |
| Arena Score, long queries | 1,344.97 1,338–1,352 | — | — |
| Arena Score, creative writing | 1,340.21 1,333–1,348 | — | — |
| Arena Score, expert questions | 1,339.49 1,327–1,352 | — | — |
| Arena Score, instruction following | 1,336.06 1,331–1,342 | — | — |
| Arena Score, working with images | 1,158.46 1,151–1,166 | — | — |
| Arena Score, understanding diagrams | 1,148.95 1,131–1,167 | — | — |
| Arena Score, text recognition in images | 1,148.81 1,137–1,160 | — | — |
| MATH, difficulty level five | 82.17 % | · with a tuned harness | Epoch evaluations |
| GeoBench — locating a place from a photograph | 77 % | — | GeoBench leaderboard |
| Fiction.LiveBench — holding a long context | 61.1 % | — | Fiction.live leaderboard |
| GPQA Diamond — graduate-level questions | 52.19 % | · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 31.04 % | · with a tuned harness | Epoch evaluations |
| CadEval — building CAD models with code | 30 % | — | CadEval Dashboard |
| WeirdML — unusual machine learning tasks | 25.77 % | — | WeirdML Leaderboard |
| The Agent Company — work tasks in an office environment | 11.4 % | — | TheAgentCompany experiment results github |
| ARC-AGI-2 | 1.3 % | — | — |