google/gemini-3-flash-preview| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| qihang-ai | $0.07 | $0.43 | 31 Jul 2026 |
| frogbot | $0.5 | $3 | 31 Jul 2026 |
| perplexity-agent | $0.5 | $3 | 31 Jul 2026 |
| gmi | $0.5 | $3 | 31 Jul 2026 |
| 302.AI | $0.5 | $3 | 31 Jul 2026 |
| AIHubMix | $0.5 | $3 | 31 Jul 2026 |
| Abacus.AI | $0.5 | $3 | 31 Jul 2026 |
| CrossModel | $0.5 | $3 | 31 Jul 2026 |
| GitHub Copilot | $0.5 | $3 | 31 Jul 2026 |
| Google DeepMind | $0.5 | $3 | 31 Jul 2026 |
| Google Vertex AI | $0.5 | $3 | 31 Jul 2026 |
| Jiekou.AI | $0.5 | $3 | 31 Jul 2026 |
| Kilo Code | $0.5 | $3 | 31 Jul 2026 |
| LLM Gateway | $0.5 | $3 | 31 Jul 2026 |
| Merge Gateway | $0.5 | $3 | 31 Jul 2026 |
| NanoGPT | $0.5 | $3 | 31 Jul 2026 |
| OpenRouter | $0.5 | $3 | 31 Jul 2026 |
| OrcaRouter | $0.5 | $3 | 31 Jul 2026 |
| Requesty | $0.5 | $3 | 31 Jul 2026 |
| ZenMux | $0.5 | $3 | 31 Jul 2026 |
| Venice AI | $0.7 | $3.75 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Mock AIME 2024–2025 — olympiad problems | 92.77 % | · with a tuned harness | Epoch evaluations |
| GeoBench — locating a place from a photograph | 88 % | — | GeoBench leaderboard |
| GPQA Diamond — graduate-level questions | 77.61 % | · with a tuned harness | Epoch evaluations |
| SWE-bench Verified — fixing bugs in repositories | 75.41 % | · with a tuned harness | Epoch evaluations |
| SimpleQA Verified — factual accuracy | 67.4 % | — | Epoch evaluations |
| Terminal-Bench — working in the command line | 64.3 % | · with a tuned harness | Terminal-Bench v2 Leaderboard |
| WeirdML — unusual machine learning tasks | 61.6 % | — | WeirdML Leaderboard |
| SimpleBench — trick questions | 53.32 % | — | SimpleBench Leaderboard |
| FrontierMath, levels 1–3 | 51.23 % | — | Epoch evaluations |
| BALROG — game environments | 48.1 % | — | Balrog Leaderboard |
| Chess puzzles | 34.76 % | — | Epoch evaluations |
| ARC-AGI-2 | 33.61 % | — | — |
| APEX-Agents | 24 % | — | — |
| ARC-AGI — generalising to unseen patterns | 21.5 % | · with a tuned harness | — |
| FrontierMath, level 4 — research-grade problems | 17.07 % | — | Epoch evaluations |
| GSO-Bench — code optimisation | 9.8 % | — | GSO Leaderboard |