Model

Qwen3-235B-A22B-Thinking-2507 NO.33

Open-weight reasoning model with strong math and science.

Alibaba Qwen · Qwen · Reasoning · Open weights

Specification

Fig.01Qwen — Alibaba Qwen

CHINA

Context window
262Ktokens
Max output
82Ktokens
Input price
$0.15/ 1M
Output price
$1.50/ 1M
Cached input
$0.11/ 1M
Throughput
Class
Reasoning
Modalities
Text
Weights
Open
Parameters
235B
License
Apache-2.0
Knowledge cutoff
Apr 2025
Released
Jul 2025

Overview

Qwen3-235B-A22B-Thinking-2507 is Alibaba's open-weight reasoning model, a 235B MoE with 22B active params and long chain-of-thought training. It is excellent at math, competition problems, and science QA, and is fully self-hostable under Apache-2.0. Its verbose reasoning traces and weaker tool-use reliability make it less dependable than the frontier for autonomous agentic coding.

Strengths

  • Best-in-class open-weight math reasoning
  • Apache-2.0 license, fully self-hostable
  • Strong GPQA and LiveCodeBench results
  • Efficient 22B active params keep inference tractable

Trade-offs

  • Verbose, slow chains of thought
  • Weaker tool-use for agentic coding
  • Text-only, no multimodal input
  • Below the frontier on long-horizon tasks

Fit

Best for

  • Advanced math and scientific reasoning
  • Competitive coding and algorithm work
  • Self-hosted reasoning deployments
  • Multilingual analytical tasks

Not ideal for

  • Latency-sensitive applications
  • Autonomous agentic coding

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
82
Coding
72
Math
90
Writing
78
Knowledge
84
Speed
55
Agentic
66
Vision
5
Multilingual
88
Long Context
80

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

BenchmarkResult
SWE-bench ProIndependentCoding · SWE-bench Pro
21.4%
Aider PolyglotIndependentCoding · Aider
57.3%
MMLU-ProKnowledge · Vendor-reported
84.4%
GPQA DiamondIndependentReasoning · Independent aggregators
81.1%
AIME 2025Math · Vendor-reported
92.3%
LiveCodeBenchIndependentCoding · LiveCodeBench
70.7%
LiveBenchIndependentGeneral · LiveBench
77.1%

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION