qwen-max-2025-01-25Providers do have offers, but we have not collected a current price yet. That is a gap on our side, not a model missing from the market.
| Task set | Result | Run conditions | Measured by |
|---|---|---|---|
| MATH, difficulty level five | 67.18 % | · with a tuned harness | Epoch evaluations |
| Fiction.LiveBench — holding a long context | 66.7 % | — | Fiction.live leaderboard |
| GPQA Diamond — graduate-level questions | 41.5 % | · with a tuned harness | Epoch evaluations |
| Aider Polyglot — code edits in six languages | 21.8 % | · with a tuned harness | Aider LLM Leaderboards |
| Mock AIME 2024–2025 — olympiad problems | 16.03 % | · with a tuned harness | Epoch evaluations |