Model
Kimi K2 Thinking NO.40
Deep-reasoning open variant of Kimi K2 for hard problems.
Moonshot AI · Kimi · Reasoning · Open weights
Specification
13 fields
Fig.01 — Kimi — Moonshot AI
CHINA
- Context window
- 262Ktokens
- Max output
- 66Ktokens
- Input price
- $0.60/ 1M
- Output price
- $2.50/ 1M
- Cached input
- $0.15/ 1M
- Throughput
- —
- Class
- Reasoning
- Modalities
- Text
- Weights
- Open
- Parameters
- 1000B
- License
- Modified MIT
- Knowledge cutoff
- Apr 2025
- Released
- Nov 2025
Overview
Kimi K2 Thinking is Moonshot's dedicated reasoning checkpoint (1T MoE, 32B active, Modified MIT) that extends test-time thinking and tool use for hard math, logic, and research tasks. It posts strong reasoning scores and long tool-call chains, but is slower and from the earlier K2 generation than K2.6. A capable open reasoner rather than a frontier agentic coder.
Strengths
- Strong deep-reasoning and tool-use scores
- Long tool-call chains for research
- Open Modified MIT weights, INT4-efficient
- Competitive math (AIME) performance
Trade-offs
- Slow due to extended thinking
- Older K2 generation than K2.6
- Agentic-coding reliability below the frontier
- Text-only input
Fit
Best for
- Hard math and logic problems
- Tool-augmented research agents
- Reasoning-heavy analysis
- Self-hosted reasoning
Not ideal for
- Latency-sensitive workloads
- Long autonomous production coding
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 84
- Coding
- 72
- Math
- 88
- Writing
- 76
- Knowledge
- 80
- Speed
- 45
- Agentic
- 68
- Vision
- 5
- Multilingual
- 76
- Long Context
- 80
Benchmarks
02 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
- GPQA DiamondIndependentReasoning · Independent aggregators
- 84.5%
- AIME 2025Math · Vendor-reported
- 94%
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 40 of 41