API provider

Replicate

What this provider sells and at what price. A price belongs to the offer, not the model: at another provider the same model may cost several times more or less.

40 models with a price
СШАcountry
resold models
pricing page price source

Models & prices

Model Developer Input, $/1M Output, $/1M Blended, $/1M
Mixtral 8x7B Mistral AI $0.3 $1 $0.48
DeepSeek Reasoner DeepSeek $3.75 $10 $5.31
Gemini 3 Pro deprecated Google DeepMind $2 $12 $4.5
Claude Haiku 4.5 (latest) Anthropic $1 $5 $2
Claude Sonnet 4.5 (latest) Anthropic $3 $15 $6
DeepSeek V3.1 DeepSeek $0.67 $2.02 $1.01
GPT 5 OpenAI $1.25 $10 $3.44
GPT-5 Mini OpenAI $0.25 $2 $0.69
GPT 5 Nano OpenAI $0.05 $0.4 $0.14
GPT OSS 120B OpenAI $0.18 $0.72 $0.32
GPT OSS 20B OpenAI $0.09 $0.36 $0.16
Qwen3 235B A22B Instruct 2507 Alibaba (Qwen) $0.26 $1.06 $0.46
Grok 4 xAI $7.2 $36 $14.4
Gemini 2.5 Flash Google DeepMind $2.5 $2.5 $2.5
Claude Sonnet 4 Anthropic $3 $15 $6
o4 Mini OpenAI $1 $4 $1.75
GPT 4.1 Nano OpenAI $0.1 $0.4 $0.18
GPT 4.1 OpenAI $2 $8 $3.5
GPT-4.1 mini OpenAI $0.4 $1.6 $0.7
Llama 2 7B Chat Meta AI $0.05 $0.25 $0.1
Claude Sonnet 3.7 Anthropic $3 $15 $6
OpenAI: o1-mini OpenAI $1.1 $4.4 $1.93
DeepSeek V3 DeepSeek $1.45 $1.45 $1.45
o1 OpenAI $15 $60 $26.25
Claude Haiku 3.5 Anthropic $1 $5 $2
GPT-4o mini OpenAI $0.15 $0.6 $0.26
Claude Sonnet 3.5 deprecated Anthropic $3.75 $18.75 $7.5
GPT 4o OpenAI $2.5 $10 $4.38
Meta: Llama 3 8B Instruct Meta AI $0.05 $0.25 $0.1
Llama 3 70B Instruct Meta AI $0.65 $2.75 $1.18
Granite 3.3 8B Instruct IBM $0.03 $0.25 $0.09
Llama 2 13B Meta AI $0.1 $0.5 $0.2
Llama 2 13B Chat Meta AI $0.1 $0.5 $0.2
Llama 2 70B Meta AI $0.65 $2.75 $1.18
Llama 2 70B Chat Meta AI $0.65 $2.75 $1.18
Llama 2 7B Meta AI $0.05 $0.25 $0.1
Llama 3 70B Meta AI $0.65 $2.75 $1.18
Llama 3 8B Meta AI $0.05 $0.25 $0.1
Mistral 7B Instruct V0.2 Mistral AI $0.05 $0.25 $0.1
Mistral 7B V0.1 Mistral AI $0.05 $0.25 $0.1

These are this provider’s prices, not the market’s best: the same model may cost differently at another provider. The blended price uses a three-input-tokens-to-one-output scheme — the way to compare models whose output costs several times more than input. Only the base per-token rate counts, without batch or discounted tiers; input and output come from one price list, not two minimums from different ones.