Anthropic

Claude Sonnet 3.7

Type
language model
Context
200K tokens
Max output
64K
Released
19 February 2025
Knowledge cutoff
October 2024
API string
claude-3-7-sonnet-20250219

Prices by provider

Provider Input, $ per 1M Output, $ per 1M In our data since
Abacus.AI $3 $15 1 Aug 2026
Amazon Bedrock anthropic.claude-3-7-sonnet-20250219-v1:0 $3 $15 2 Aug 2026
Anthropic $3 $15 1 Aug 2026
Merge Gateway $3 $15 1 Aug 2026
Amazon Bedrock bedrock/us-gov-west-1/anthropic.claude-3-7-sonnet-20250219-v1:0 $3.6 $18 2 Aug 2026

Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.

A single provider appears several times if it sells this model under different identifiers and at different prices — for example, at different weight precision. The identifier is shown next to the name. Identical offers that arrived from two sources under different spellings are merged into one row.

The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.

Measurement results

Task set Result Run conditions Measured by
Arena Score, long queries 1,354.73 1,347–1,362
Arena Score, programming 1,338.88 1,331–1,346
Arena Score, multi-turn dialogue 1,338.32 1,331–1,346
Arena Score, instruction following 1,325.19 1,319–1,331
Arena Score, mathematics 1,318.41 1,308–1,328
Arena Score, creative writing 1,316.76 1,309–1,325
Arena Score, hard prompts 1,315.07 1,309–1,321
Arena Score in English 1,311.02 1,306–1,316
Arena Score in Russian 1,309.67 1,299–1,320
Arena Score in French 1,302.75 1,275–1,330
Arena Score, expert questions 1,300.32 1,287–1,313
Arena Score, overall 1,299.32 1,295–1,303
Arena Score in Spanish 1,289.11 1,263–1,315
Arena Score, text recognition in images 1,158.07 1,139–1,177
Arena Score, understanding diagrams 1,151.01 1,119–1,183
Arena Score, working with images 1,151 1,142–1,160
MATH, difficulty level five 91.16 % thinking budget: 64K · with a tuned harness Epoch evaluations
Fiction.LiveBench — holding a long context 83.3 % Fiction.live leaderboard
Creative writing (Lech Mazur’s evaluation) 81.1 % thinking budget: 16K lechmazur/writing Github repository
GPQA Diamond — graduate-level questions 72.98 % thinking budget: 64K · with a tuned harness Epoch evaluations
GeoBench — locating a place from a photograph 68 % thinking budget: 15K GeoBench leaderboard
Aider Polyglot — code edits in six languages 64.9 % · with a tuned harness Aider LLM Leaderboards
SWE-bench Verified — fixing bugs in repositories 60.95 % · with a tuned harness Epoch evaluations
Mock AIME 2024–2025 — olympiad problems 57.74 % thinking budget: 64K · with a tuned harness Epoch evaluations
CadEval — building CAD models with code 54 % CadEval Dashboard
DeepResearch Bench — deep research 43.6 % thinking budget: 2K DeepResearchBench Leaderboard
OSWorld — working inside an operating system, first version 35.8 % · with a tuned harness OS World Website
SimpleBench — trick questions 35.68 % thinking budget: 12K SimpleBench Leaderboard
The Agent Company — work tasks in an office environment 30.9 % TheAgentCompany leaderboard
ARC-AGI — generalising to unseen patterns 28.6 % thinking budget: 16K · with a tuned harness ARC Prize Leaderboard
Cybench — cybersecurity tasks 20 % · with a tuned harness Cybench leaderboard
GSO-Bench — code optimisation 3.8 % GSO Leaderboard
Humanity’s Last Exam — expert-level questions 3.4 %
ARC-AGI-2 0.9 % thinking budget: 8K