Qwen/Qwen2.5-72B-Instruct| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| Abacus.AI | $0.11 | $0.38 | 31 Jul 2026 |
| DeepInfra | $0.12 | $0.39 | 31 Jul 2026 |
| Hyperbolic | $0.12 | $0.3 | 31 Jul 2026 |
| Nebius Token Factory | $0.13 | $0.4 | 31 Jul 2026 |
| alibaba-cn | $0.574 | $1.721 | 31 Jul 2026 |
| siliconflow-cn | $0.59 | $0.59 | 31 Jul 2026 |
| SiliconFlow | $0.59 | $0.59 | 31 Jul 2026 |
| Fireworks AI | $0.9 | $0.9 | 1 Aug 2026 |
| Alibaba (Qwen) | $1.4 | $5.6 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, programming | 1,292.93 1,285–1,301 | — | — |
| Arena Score in English | 1,282.77 1,278–1,288 | — | — |
| Arena Score, mathematics | 1,282.63 1,274–1,291 | — | — |
| Arena Score, long queries | 1,281.92 1,274–1,290 | — | — |
| Arena Score in French | 1,278.91 1,246–1,312 | — | — |
| Arena Score, multi-turn dialogue | 1,271.85 1,264–1,280 | — | — |
| Arena Score, hard prompts | 1,270.92 1,265–1,277 | — | — |
| Arena Score, overall | 1,269.18 1,265–1,273 | — | — |
| Arena Score in Russian | 1,262.99 1,254–1,272 | — | — |
| Arena Score, instruction following | 1,254.53 1,249–1,260 | — | — |
| Arena Score in Spanish | 1,254.22 1,223–1,285 | — | — |
| Arena Score, expert questions | 1,244.85 1,233–1,257 | — | — |
| Arena Score, creative writing | 1,221.67 1,213–1,230 | — | — |
| MATH, difficulty level five | 63.17 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 32.2 % | · with a tuned harness | Epoch evaluations |
| BALROG — game environments | 16.2 % | — | Balrog Leaderboard |
| WeirdML — unusual machine learning tasks | 15.97 % | — | WeirdML Leaderboard |
| Mock AIME 2024–2025 — olympiad problems | 7.96 % | · with a tuned harness | Epoch evaluations |
| The Agent Company — work tasks in an office environment | 5.7 % | — | TheAgentCompany experiment results github |