API provider

Weights & Biases

What this provider sells and at what price. A price belongs to the offer, not the model: at another provider the same model may cost several times more or less.

32 models with a price
resold models
aggregate databases price source

Models & prices

Model Developer Input, $/1M Output, $/1M Blended, $/1M
Kimi K3 Moonshot AI (Kimi) $3 $15 $6
GLM 5.2 Zhipu AI / Z.ai $0.76 $2.42 $1.18
Kimi K2.7 Code Moonshot AI (Kimi) $0.71 $3.5 $1.41
MiniMax M3 MiniMax $0.23 $0.96 $0.41
Nemotron 3 Ultra 550B A55B NVIDIA $0.75 $2.75 $1.25
Mellum2 12B A2.5B JetBrains $0.05 $0.1 $0.06
Granite 4.1 8B IBM $0.05 $0.1 $0.06
DeepSeek V4 Flash DeepSeek $0.14 $0.28 $0.18
DeepSeek V4 Pro DeepSeek $1.74 $3.48 $2.18
Qwen3.6 27B Alibaba (Qwen) $0.6 $3.6 $1.35
Kimi K2.6 Moonshot AI (Kimi) $0.65 $3.41 $1.34
Qwen3.6 35B-A3B Alibaba (Qwen) $0.25 $1.25 $0.5
Gemma 4 31B Google DeepMind $0.1 $0.34 $0.16
GLM 5.1 Zhipu AI / Z.ai $1.4 $4.4 $2.15
Nemotron 3 Super NVIDIA $0.2 $0.8 $0.35
Qwen3.5 27B deprecated Alibaba (Qwen) $0.39 $3.12 $1.07
Qwen3.5 35B A3B Alibaba (Qwen) $0.25 $1.25 $0.5
MiniMax M2.5 and another 1 MiniMax $0.3 $1.2 $0.53
Kimi K2.5 and another 1 Moonshot AI (Kimi) $0.6 $3 $1.2
DeepSeek V3.1 DeepSeek $0.55 $1.65 $0.83
GPT OSS 120B OpenAI $0.03 $0.17 $0.07
GPT OSS 20B OpenAI $0.03 $0.13 $0.06
Qwen3 30B A3B Instruct 2507 Alibaba (Qwen) $0.1 $0.3 $0.15
Qwen3 235B A22B Thinking-2507 deprecated Alibaba (Qwen) $0.1 $0.1 $0.1
Qwen3 Coder 480B A35B Alibaba (Qwen) $1 $1.5 $1.13
Qwen3 235B A22B Instruct 2507 Alibaba (Qwen) $0.1 $0.1 $0.1
Kimi K2 Instruct Moonshot AI (Kimi) $0.6 $2.5 $1.08
Qwen3 14B Instruct Alibaba (Qwen) $0.05 $0.22 $0.09
Phi 4 Mini 3.8B deprecated Microsoft $0.08 $0.35 $0.15
Llama 3.3 70B Instruct Meta AI $0.71 $0.71 $0.71
Llama 3.1 70B Meta AI $0.8 $0.8 $0.8
Llama 3.1 8B Meta AI $0.22 $0.22 $0.22

These are this provider’s prices, not the market’s best: the same model may cost differently at another provider. The blended price uses a three-input-tokens-to-one-output scheme — the way to compare models whose output costs several times more than input. Only the base per-token rate counts, without batch or discounted tiers; input and output come from one price list, not two minimums from different ones.