llama-3.1-405b-instruct| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Hyperbolic | $0.12 | $0.3 | 1 Aug 2026 |
| Nebius Token Factory | $1 | $3 | 1 Aug 2026 |
| Google Vertex AI | $5 | $16 | 31 Jul 2026 |
| SambaNova | $5 | $10 | 1 Aug 2026 |
| Databricks | $5 | $15 | 1 Aug 2026 |
| Amazon Bedrock | $5.32 | $16 | 2 Aug 2026 |
| Microsoft Foundry | $5.33 | $16 | 1 Aug 2026 |
| OCI Generative AI | $10.68 | $10.68 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score in English | 1,308.06 1,303–1,313 | — | — |
| Arena Score, multi-turn dialogue | 1,287.34 1,280–1,295 | — | — |
| Arena Score, programming | 1,284 1,277–1,291 | — | — |
| Arena Score, overall | 1,281.97 1,278–1,286 | — | — |
| Arena Score, mathematics | 1,278.23 1,270–1,286 | — | — |
| Arena Score in French | 1,269.69 1,242–1,297 | — | — |
| Arena Score, hard prompts | 1,263.75 1,258–1,270 | — | — |
| Arena Score, creative writing | 1,260.14 1,253–1,268 | — | — |
| Arena Score, long queries | 1,259.79 1,252–1,267 | — | — |
| Arena Score, instruction following | 1,259.06 1,254–1,264 | — | — |
| Arena Score in Russian | 1,255.23 1,247–1,264 | — | — |
| Arena Score in Spanish | 1,252.29 1,227–1,277 | — | — |
| Arena Score, expert questions | 1,228.55 1,216–1,241 | — | — |
| MATH, difficulty level five | 49.77 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 34.55 % | · with a tuned harness | Epoch evaluations |
| WeirdML — unusual machine learning tasks | 21.38 % | — | WeirdML Leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 9.63 % | · with a tuned harness | Epoch evaluations |
| SimpleBench — trick questions | 7.6 % | — | SimpleBench Leaderboard |
| Cybench — cybersecurity tasks | 7.5 % | · with a tuned harness | Cybench leaderboard |
| The Agent Company — work tasks in an office environment | 7.4 % | — | TheAgentCompany experiment results github |