Anthropic

Claude 4 Sonnet

Type
language model
Context
1,000K tokens
Max output
64K
Released
29 September 2025
Knowledge cutoff
March 2025
API string
claude-sonnet-4-20250514

Prices by provider

Provider Input, $ per 1M Output, $ per 1M In our data since
Jiekou.AI $2.7 $13.5 31 Jul 2026
NanoGPT $2.992 $14.994 31 Jul 2026
302.AI $3 $15 31 Jul 2026
Abacus.AI $3 $15 31 Jul 2026
Amazon Bedrock $3 $15 2 Aug 2026
Anthropic $3 $15 1 Aug 2026
Merge Gateway $3 $15 31 Jul 2026

Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.

The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.

Measurement results

Task set Result Run conditions Measured by
Arena Score, long queries 1,379.71 1,372–1,387
Arena Score, programming 1,379.68 1,372–1,387
Arena Score in French 1,363.66 1,337–1,390
Arena Score, multi-turn dialogue 1,363.52 1,356–1,371
Arena Score, mathematics 1,357.74 1,346–1,369
Arena Score, hard prompts 1,354.2 1,349–1,360
Arena Score in Spanish 1,350.72 1,329–1,373
Arena Score, instruction following 1,350.14 1,343–1,357
Arena Score in English 1,345.71 1,340–1,351
Arena Score in Russian 1,344.73 1,333–1,356
Arena Score, creative writing 1,339.09 1,330–1,348
Arena Score, overall 1,337.9 1,334–1,342
Arena Score, expert questions 1,333.62 1,320–1,347
Arena Score, understanding diagrams 1,178.16 1,150–1,207
Arena Score, working with images 1,175.73 1,162–1,190
Arena Score, text recognition in images 1,173.12 1,156–1,191
MATH, difficulty level five 84.37 % · with a tuned harness Epoch evaluations
Creative writing (Lech Mazur’s evaluation) 81.4 % thinking budget: 16K lechmazur/writing Github repository
GPQA Diamond — graduate-level questions 72.25 % thinking budget: 59K · with a tuned harness Epoch evaluations
Mock AIME 2024–2025 — olympiad problems 71.08 % thinking budget: 59K · with a tuned harness Epoch evaluations
Aider Polyglot — code edits in six languages 61.3 % · with a tuned harness Aider LLM Leaderboards
DeepResearch Bench — deep research 47.8 % thinking budget: 2K DeepResearchBench Leaderboard
Fiction.LiveBench — holding a long context 46.9 % Fiction.live leaderboard
WeirdML — unusual machine learning tasks 46.11 % thinking budget: 16K WeirdML Leaderboard
OSWorld — working inside an operating system, first version 43.9 % · with a tuned harness OS World Website
ARC-AGI — generalising to unseen patterns 40 % thinking budget: 16K · with a tuned harness ARC Prize Leaderboard
GeoBench — locating a place from a photograph 37 % GeoBench leaderboard
Cybench — cybersecurity tasks 35 % · with a tuned harness Cybench leaderboard
SimpleBench — trick questions 34.6 % thinking budget: 12K SimpleBench Leaderboard
The Agent Company — work tasks in an office environment 33.1 % TheAgentCompany leaderboard
APEX-Agents 9.3 %
ARC-AGI-2 5.93 % thinking budget: 16K
GSO-Bench — code optimisation 4.9 % GSO Leaderboard
Humanity’s Last Exam — expert-level questions 3.11 %
CritPt — physics problems 0.29 %