qwen/qwen3-235b-a22b| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| DeepInfra | $0.18 | $0.54 | 31 Jul 2026 |
| novita-ai | $0.2 | $0.8 | 31 Jul 2026 |
| Hugging Face | $0.2 | $0.8 | 31 Jul 2026 |
| Jiekou.AI | $0.2 | $0.8 | 31 Jul 2026 |
| LLM Gateway | $0.2 | $0.8 | 31 Jul 2026 |
| Nebius Token Factory | $0.2 | $0.6 | 31 Jul 2026 |
| Novita AI | $0.2 | $0.8 | 31 Jul 2026 |
| Fireworks AI | $0.22 | $0.88 | 31 Jul 2026 |
| alibaba-cn | $0.287 | $1.147 | 31 Jul 2026 |
| Merge Gateway | $0.287 | $1.147 | 31 Jul 2026 |
| 302.AI | $0.29 | $2.86 | 31 Jul 2026 |
| NanoGPT | $0.3 | $0.5 | 31 Jul 2026 |
| Kilo Code | $0.455 | $1.82 | 31 Jul 2026 |
| OpenRouter | $0.455 | $1.82 | 31 Jul 2026 |
| Alibaba (Qwen) | $0.7 | $2.8 | 31 Jul 2026 |
| Hyperbolic | $2 | $2 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, mathematics | 1,395.48 1,382–1,409 | — | — |
| Arena Score, programming | 1,385.71 1,377–1,395 | — | — |
| Arena Score in English | 1,376.7 1,371–1,383 | — | — |
| Arena Score in French | 1,366.73 1,334–1,399 | — | — |
| Arena Score, overall | 1,365.94 1,361–1,371 | — | — |
| Arena Score in Spanish | 1,363.03 1,334–1,392 | — | — |
| Arena Score, hard prompts | 1,362.34 1,356–1,369 | — | — |
| Arena Score, multi-turn dialogue | 1,362.16 1,353–1,371 | — | — |
| Arena Score, long queries | 1,351.96 1,343–1,361 | — | — |
| Arena Score, expert questions | 1,348.1 1,332–1,364 | — | — |
| Arena Score in Russian | 1,338.23 1,325–1,351 | — | — |
| Arena Score, instruction following | 1,333.27 1,326–1,341 | — | — |
| Arena Score, creative writing | 1,316.5 1,306–1,327 | — | — |
| Creative writing (Lech Mazur’s evaluation) | 83 % | — | lechmazur/writing Github repository |
| MATH, difficulty level five | 68.86 % | · with a tuned harness | Epoch evaluations |
| Fiction.LiveBench — holding a long context | 67.7 % | — | Fiction.live leaderboard |
| GPQA Diamond — graduate-level questions | 60.94 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 59.6 % | · with a tuned harness | Aider LLM Leaderboards |
| WeirdML — unusual machine learning tasks | 37.28 % | — | WeirdML Leaderboard |
| SimpleBench — trick questions | 17.2 % | — | SimpleBench Leaderboard |