vercel_ai_gateway/anthropic/claude-3-5-sonnet-20241022| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| vercel_ai_gateway | $3 | $15 | 1 Aug 2026 |
| Amazon Bedrock | $3 | $15 | 2 Aug 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, programming | 1,342.87 1,337–1,348 | — | — |
| Arena Score, multi-turn dialogue | 1,325.58 1,320–1,331 | — | — |
| Arena Score, long queries | 1,311.8 1,306–1,317 | — | — |
| Arena Score in English | 1,307.74 1,304–1,312 | — | — |
| Arena Score, mathematics | 1,306.76 1,300–1,313 | — | — |
| Arena Score, hard prompts | 1,305.53 1,301–1,310 | — | — |
| Arena Score in Russian | 1,304.29 1,297–1,311 | — | — |
| Arena Score in French | 1,301.79 1,280–1,323 | — | — |
| Arena Score, overall | 1,297.79 1,295–1,301 | — | — |
| Arena Score, instruction following | 1,297.5 1,293–1,302 | — | — |
| Arena Score, creative writing | 1,291.87 1,286–1,298 | — | — |
| Arena Score in Spanish | 1,283.4 1,264–1,303 | — | — |
| Arena Score, expert questions | 1,264.39 1,255–1,274 | — | — |
| Arena Score, text recognition in images | 1,138.75 1,120–1,158 | — | — |
| Arena Score, understanding diagrams | 1,132.28 1,101–1,164 | — | — |
| Arena Score, working with images | 1,125.48 1,118–1,133 | — | — |
| Creative writing (Lech Mazur’s evaluation) | 80.3 % | — | lechmazur/writing Github repository |
| GeoBench — locating a place from a photograph | 62 % | — | GeoBench leaderboard |
| MATH, difficulty level five | 56.95 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 51.6 % | · with a tuned harness | Aider LLM Leaderboards |
| CadEval — building CAD models with code | 48 % | — | CadEval Dashboard |
| GPQA Diamond — graduate-level questions | 40.4 % | · with a tuned harness | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 39.97 % | — | WeirdML Leaderboard |
| BALROG — game environments | 32.6 % | — | Balrog Leaderboard |
| SimpleBench — trick questions | 29.68 % | — | SimpleBench Leaderboard |
| The Agent Company — work tasks in an office environment | 24 % | — | TheAgentCompany experiment results github |
| Mock AIME 2024–2025 — olympiad problems | 8.38 % | · with a tuned harness | Epoch evaluations |
| GSO-Bench — code optimisation | 4.6 % | — | GSO Leaderboard |
| Humanity’s Last Exam — expert-level questions | 0 % | — | — |