qwen3.6-max-preview| Provider | Input, $ per 1M | Output, $ per 1M | In our data since |
|---|---|---|---|
| OpenRouter | $1.027 | $6.162 | 31 Jul 2026 |
| Kilo Code | $1.04 | $6.24 | 31 Jul 2026 |
| Pioneer | $1.04 | $6.24 | 31 Jul 2026 |
| AIHubMix | $1.27 | $7.61 | 31 Jul 2026 |
| Alibaba (Qwen) | $1.3 | $7.8 | 31 Jul 2026 |
| LLM Gateway | $1.3 | $7.8 | 31 Jul 2026 |
| NanoGPT | $1.3 | $7.8 | 31 Jul 2026 |
| EmpirioLabs AI | $1.31 | $7.88 | 31 Jul 2026 |
| Merge Gateway | $1.31 | $7.88 | 31 Jul 2026 |
| alibaba-cn | $1.32 | $7.9 | 31 Jul 2026 |
Only the base tier and only per-token prices are shown. Batch, discounted and cached rates, as well as prices per image or per second of video, do not go into this table: they cannot stand in the same column as a price per million tokens.
The date in the last column is the day this price first entered our collection. The price may well be older: before that day we simply were not recording it. It has not changed since — otherwise a new row with a new date would stand in its place.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| Arena Score, web development | 1,478.18 1,465–1,491 | — | — |
| Arena Score, expert questions | 1,473.55 1,448–1,500 | — | — |
| Arena Score, programming | 1,467.25 1,452–1,483 | — | — |
| Arena Score, mathematics | 1,462.55 1,433–1,492 | — | — |
| Arena Score, long queries | 1,457.94 1,444–1,472 | — | — |
| Arena Score, multi-turn dialogue | 1,455.97 1,436–1,476 | — | — |
| Arena Score, hard prompts | 1,455.91 1,445–1,466 | — | — |
| Arena Score in French | 1,455.35 1,413–1,498 | — | — |
| Arena Score in Spanish | 1,452.39 1,415–1,490 | — | — |
| Arena Score in English | 1,449.79 1,438–1,462 | — | — |
| Arena Score, overall | 1,446.35 1,438–1,455 | — | — |
| Arena Score in Russian | 1,444.75 1,419–1,471 | — | — |
| Arena Score, instruction following | 1,436.55 1,422–1,451 | — | — |
| Arena Score, creative writing | 1,433.85 1,411–1,457 | — | — |
| Mock AIME 2024–2025 — olympiad problems | 91.1 % | · with a tuned harness | Epoch evaluations |
| GPQA Diamond — graduate-level questions | 85.44 % | · with a tuned harness | Epoch evaluations |
| SWE-bench Verified — fixing bugs in repositories | 76.65 % | · with a tuned harness | Epoch evaluations |
| SimpleQA Verified — factual accuracy | 56.93 % | — | Epoch evaluations |
| SimpleBench — trick questions | 55.6 % | — | SimpleBench Leaderboard |
| Chess puzzles | 12.67 % | — | Epoch evaluations |