claude-sonnet-4-20250514| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Jiekou.AI | $2.7 | $13.5 | 31 Jul 2026 |
| NanoGPT | $2.992 | $14.994 | 31 Jul 2026 |
| 302.AI | $3 | $15 | 31 Jul 2026 |
| Abacus.AI | $3 | $15 | 31 Jul 2026 |
| Amazon Bedrock | $3 | $15 | 2 Aug 2026 |
| Anthropic | $3 | $15 | 1 Aug 2026 |
| Merge Gateway | $3 | $15 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, long queries | 1,379.71 1,372–1,387 | — | — |
| Arena Score, programming | 1,379.68 1,372–1,387 | — | — |
| Arena Score in French | 1,363.66 1,337–1,390 | — | — |
| Arena Score, multi-turn dialogue | 1,363.52 1,356–1,371 | — | — |
| Arena Score, mathematics | 1,357.74 1,346–1,369 | — | — |
| Arena Score, hard prompts | 1,354.2 1,349–1,360 | — | — |
| Arena Score in Spanish | 1,350.72 1,329–1,373 | — | — |
| Arena Score, instruction following | 1,350.14 1,343–1,357 | — | — |
| Arena Score in English | 1,345.71 1,340–1,351 | — | — |
| Arena Score in Russian | 1,344.73 1,333–1,356 | — | — |
| Arena Score, creative writing | 1,339.09 1,330–1,348 | — | — |
| Arena Score, overall | 1,337.9 1,334–1,342 | — | — |
| Arena Score, expert questions | 1,333.62 1,320–1,347 | — | — |
| Arena Score, understanding diagrams | 1,178.16 1,150–1,207 | — | — |
| Arena Score, working with images | 1,175.73 1,162–1,190 | — | — |
| Arena Score, text recognition in images | 1,173.12 1,156–1,191 | — | — |
| MATH, difficulty level five | 84.37 % | · with a tuned harness | Epoch evaluations |
| Creative writing (Lech Mazur’s evaluation) | 81.4 % | thinking budget: 16K | lechmazur/writing Github repository |
| GPQA Diamond — graduate-level questions | 72.25 % | thinking budget: 59K · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 71.08 % | thinking budget: 59K · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 61.3 % | · with a tuned harness | Aider LLM Leaderboards |
| DeepResearch Bench — deep research | 47.8 % | thinking budget: 2K | DeepResearchBench Leaderboard |
| Fiction.LiveBench — holding a long context | 46.9 % | — | Fiction.live leaderboard |
| WeirdML — unusual machine learning tasks | 46.11 % | thinking budget: 16K | WeirdML Leaderboard |
| OSWorld — working inside an operating system, first version | 43.9 % | · with a tuned harness | OS World Website |
| ARC-AGI — generalising to unseen patterns | 40 % | thinking budget: 16K · with a tuned harness | ARC Prize Leaderboard |
| GeoBench — locating a place from a photograph | 37 % | — | GeoBench leaderboard |
| Cybench — cybersecurity tasks | 35 % | · with a tuned harness | Cybench leaderboard |
| SimpleBench — trick questions | 34.6 % | thinking budget: 12K | SimpleBench Leaderboard |
| The Agent Company — work tasks in an office environment | 33.1 % | — | TheAgentCompany leaderboard |
| APEX-Agents | 9.3 % | — | — |
| ARC-AGI-2 | 5.93 % | thinking budget: 16K | — |
| GSO-Bench — code optimisation | 4.9 % | — | GSO Leaderboard |
| Humanity’s Last Exam — expert-level questions | 3.11 % | — | — |
| CritPt — physics problems | 0.29 % | — | — |