Model
Qwen3.7-Max NO.32
Alibaba's proprietary agentic reasoning flagship with a 1M context.
Alibaba Qwen · Qwen · Balanced · Closed weights
Specification
11 fields
Fig.01 — Qwen — Alibaba Qwen
CHINA
- Context window
- 1Mtokens
- Max output
- 66Ktokens
- Input price
- $2.50/ 1M
- Output price
- $7.50/ 1M
- Cached input
- $0.25/ 1M
- Throughput
- —
- Class
- Balanced
- Modalities
- Text
- Weights
- Closed
- Knowledge cutoff
- Dec 2025
- Released
- May 2026
Overview
Qwen3.7-Max is Alibaba's current proprietary flagship (May 2026, the successor to Qwen3-Max), an agent-oriented reasoning model with a 1M-token context window and an extended-thinking mode aimed at long-horizon tasks. It posts leaderboard-topping numbers and is exceptional value — list pricing is $2.50/$7.50 per 1M tokens, and Alibaba frequently runs a 50%-off promo — with genuinely strong multilingual and math ability. Those benchmarks overstate day-to-day dependability, though — on long autonomous coding it is clearly a step below the Claude/GPT frontier, and it is text-only.
Strengths
- Outstanding multilingual coverage
- Strong math and reasoning benchmarks
- 1M-token context for long tasks
- Aggressive pricing for a flagship
Trade-offs
- Real-world agentic coding below Claude/GPT
- Benchmark scores overstate reliability
- No native vision in the Max tier
- Can loop or over-plan on long agent runs
Fit
Best for
- Multilingual reasoning and generation
- Math and analytical tasks
- Cost-sensitive agentic workflows
- Long-context document processing
Not ideal for
- Mission-critical autonomous coding
- Vision-dependent tasks
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 85
- Coding
- 76
- Math
- 90
- Writing
- 82
- Knowledge
- 86
- Speed
- 62
- Agentic
- 74
- Vision
- 5
- Multilingual
- 92
- Long Context
- 84
Benchmarks
04 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
24/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 24%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 69.6%
- τ-benchIndependentAgentic · τ-bench
- 74.8%
- LMArena EloIndependentGeneral · LMArena
- 1428Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 32 of 41