Model
Qwen3-235B-A22B-Thinking-2507 NO.33
Open-weight reasoning model with strong math and science.
Alibaba Qwen · Qwen · Reasoning · Open weights
Specification
13 fields
Fig.01 — Qwen — Alibaba Qwen
CHINA
- Context window
- 262Ktokens
- Max output
- 82Ktokens
- Input price
- $0.15/ 1M
- Output price
- $1.50/ 1M
- Cached input
- $0.11/ 1M
- Throughput
- —
- Class
- Reasoning
- Modalities
- Text
- Weights
- Open
- Parameters
- 235B
- License
- Apache-2.0
- Knowledge cutoff
- Apr 2025
- Released
- Jul 2025
Overview
Qwen3-235B-A22B-Thinking-2507 is Alibaba's open-weight reasoning model, a 235B MoE with 22B active params and long chain-of-thought training. It is excellent at math, competition problems, and science QA, and is fully self-hostable under Apache-2.0. Its verbose reasoning traces and weaker tool-use reliability make it less dependable than the frontier for autonomous agentic coding.
Strengths
- Best-in-class open-weight math reasoning
- Apache-2.0 license, fully self-hostable
- Strong GPQA and LiveCodeBench results
- Efficient 22B active params keep inference tractable
Trade-offs
- Verbose, slow chains of thought
- Weaker tool-use for agentic coding
- Text-only, no multimodal input
- Below the frontier on long-horizon tasks
Fit
Best for
- Advanced math and scientific reasoning
- Competitive coding and algorithm work
- Self-hosted reasoning deployments
- Multilingual analytical tasks
Not ideal for
- Latency-sensitive applications
- Autonomous agentic coding
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 82
- Coding
- 72
- Math
- 90
- Writing
- 78
- Knowledge
- 84
- Speed
- 55
- Agentic
- 66
- Vision
- 5
- Multilingual
- 88
- Long Context
- 80
Benchmarks
07 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 21.4%
- Aider PolyglotIndependentCoding · Aider
- 57.3%
- MMLU-ProKnowledge · Vendor-reported
- 84.4%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 81.1%
- AIME 2025Math · Vendor-reported
- 92.3%
- LiveCodeBenchIndependentCoding · LiveCodeBench
- 70.7%
- LiveBenchIndependentGeneral · LiveBench
- 77.1%
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 33 of 41