claude-opus-4-1-20250805| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Jiekou.AI | $13.5 | $67.5 | 31 Jul 2026 |
| NanoGPT | $14.994 | $75.004 | 31 Jul 2026 |
| 302.AI | $15 | $75 | 31 Jul 2026 |
| Abacus.AI | $15 | $75 | 31 Jul 2026 |
| Amazon Bedrock | $15 | $75 | 2 Aug 2026 |
| Anthropic | $15 | $75 | 31 Jul 2026 |
| Helicone | $15 | $75 | 31 Jul 2026 |
| LLM Gateway | $15 | $75 | 31 Jul 2026 |
| Merge Gateway | $15 | $75 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, programming | 1,473.09 1,468–1,478 | — | — |
| Arena Score, long queries | 1,445.19 1,440–1,450 | — | — |
| Arena Score in Spanish | 1,444.59 1,430–1,459 | — | — |
| Arena Score, multi-turn dialogue | 1,441.82 1,436–1,448 | — | — |
| Arena Score, hard prompts | 1,441.49 1,438–1,445 | — | — |
| Arena Score, instruction following | 1,433.82 1,429–1,439 | — | — |
| Arena Score in English | 1,428.72 1,425–1,433 | — | — |
| Arena Score in French | 1,427.47 1,410–1,445 | — | — |
| Arena Score, expert questions | 1,425.18 1,415–1,435 | — | — |
| Arena Score in Russian | 1,422.37 1,414–1,430 | — | — |
| Arena Score, mathematics | 1,421.84 1,413–1,430 | — | — |
| Arena Score, overall | 1,417.42 1,414–1,420 | — | — |
| Arena Score, creative writing | 1,410.22 1,404–1,416 | — | — |
| Arena Score, web development | 1,388.74 1,378–1,400 | — | — |
| Creative writing (Lech Mazur’s evaluation) | 84.7 % | — | lechmazur/writing Github repository |
| SWE-bench Verified — fixing bugs in repositories | 73.35 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 69.7 % | thinking budget: 27K · with a tuned harness | Epoch evaluations |
| Mock AIME 2024–2025 — olympiad problems | 68.86 % | thinking budget: 27K · with a tuned harness | Epoch evaluations |
| SimpleBench — trick questions | 52 % | — | SimpleBench Leaderboard |
| DeepResearch Bench — deep research | 49.7 % | — | DeepResearchBench Leaderboard |
| GDPval — tasks from real occupations | 43.6 % | — | — |
| WeirdML — unusual machine learning tasks | 42.76 % | thinking budget: 16K | WeirdML Leaderboard |
| Cybench — cybersecurity tasks | 42 % | · with a tuned harness | Cybench leaderboard |
| Terminal-Bench — working in the command line | 38 % | · with a tuned harness | Terminal-Bench v2 Leaderboard |
| SimpleQA Verified — factual accuracy | 34.8 % | thinking budget: 27K | Epoch evaluations |
| FrontierMath, levels 1–3 | 12.63 % | thinking budget: 32K | Epoch evaluations |
| Humanity’s Last Exam — expert-level questions | 7.06 % | — | — |
| FrontierMath, level 4 — research-grade problems | 2.44 % | thinking budget: 32K | Epoch evaluations |
| Chess puzzles | 2.15 % | — | Epoch evaluations |