claude-opus-4-20250514| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Jiekou.AI | $13.5 | $67.5 | 31 Jul 2026 |
| NanoGPT | $14.994 | $75.004 | 31 Jul 2026 |
| 302.AI | $15 | $75 | 31 Jul 2026 |
| Abacus.AI | $15 | $75 | 31 Jul 2026 |
| Amazon Bedrock | $15 | $75 | 2 Aug 2026 |
| Anthropic | $15 | $75 | 1 Aug 2026 |
| Merge Gateway | $15 | $75 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, programming | 1,401.8 1,395–1,409 | — | — |
| Arena Score, long queries | 1,401.79 1,395–1,409 | — | — |
| Arena Score, multi-turn dialogue | 1,385.96 1,378–1,393 | — | — |
| Arena Score in Russian | 1,380.6 1,369–1,392 | — | — |
| Arena Score in English | 1,376.69 1,371–1,382 | — | — |
| Arena Score, hard prompts | 1,375.85 1,370–1,381 | — | — |
| Arena Score, creative writing | 1,374.61 1,366–1,383 | — | — |
| Arena Score, mathematics | 1,373.35 1,362–1,384 | — | — |
| Arena Score, instruction following | 1,372.89 1,366–1,379 | — | — |
| Arena Score, expert questions | 1,371.29 1,359–1,384 | — | — |
| Arena Score, overall | 1,364.57 1,360–1,369 | — | — |
| Arena Score in French | 1,364.46 1,339–1,390 | — | — |
| Arena Score in Spanish | 1,350.89 1,330–1,372 | — | — |
| Arena Score, text recognition in images | 1,179.05 1,163–1,195 | — | — |
| Arena Score, working with images | 1,175.62 1,163–1,189 | — | — |
| Arena Score, understanding diagrams | 1,175.37 1,150–1,201 | — | — |
| MATH, difficulty level five | 85.05 % | · with a tuned harness | Epoch evaluations |
| Creative writing (Lech Mazur’s evaluation) | 83.6 % | thinking budget: 16K | lechmazur/writing Github repository |
| Aider Polyglot — code edits in six languages | 72 % | · with a tuned harness | Aider LLM Leaderboards |
| SWE-bench Verified — fixing bugs in repositories | 70.66 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 68.35 % | thinking budget: 16K · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 64.41 % | thinking budget: 27K · with a tuned harness | Epoch evaluations |
| Fiction.LiveBench — holding a long context | 61.1 % | — | Fiction.live leaderboard |
| SimpleBench — trick questions | 50.56 % | thinking budget: 12K | SimpleBench Leaderboard |
| DeepResearch Bench — deep research | 49 % | thinking budget: 2K | DeepResearchBench Leaderboard |
| GeoBench — locating a place from a photograph | 49 % | thinking budget: 32K | GeoBench leaderboard |
| WeirdML — unusual machine learning tasks | 43.4 % | thinking budget: 16K | WeirdML Leaderboard |
| Cybench — cybersecurity tasks | 38 % | · with a tuned harness | Cybench leaderboard |
| ARC-AGI — generalising to unseen patterns | 35.7 % | thinking budget: 16K · with a tuned harness | ARC Prize Leaderboard |
| ARC-AGI-2 | 8.61 % | thinking budget: 16K | — |
| GSO-Bench — code optimisation | 6.9 % | — | GSO Leaderboard |
| Humanity’s Last Exam — expert-level questions | 6.22 % | — | — |
| CritPt — physics problems | 0.3 % | — | — |