Model
Claude Sonnet 5 NO.05
The most agentic Sonnet yet — near-Opus coding reliability at Sonnet cost.
Anthropic · Claude · Balanced · Closed weights
Specification
11 fields
Fig.01 — Claude — Anthropic
USA
- Context window
- 1Mtokens
- Max output
- 128Ktokens
- Input price
- $3/ 1M
- Output price
- $15/ 1M
- Cached input
- $0.30/ 1M
- Throughput
- —
- Class
- Balanced
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Jan 2026
- Released
- Jun 2026
Overview
Claude Sonnet 5 is Anthropic's newest balanced-tier model and the most agentic Sonnet to date, closing much of the gap to Opus on real coding and terminal/tool-use work while keeping Sonnet pricing. It leads its predecessor decisively and edges Opus 4.8 on Terminal-Bench, sustaining long agent loops reliably — the new default production workhorse for most agentic and coding workloads. It pairs adaptive thinking with a 1M-token context and high-resolution vision. Introductory pricing of $2/$10 per 1M tokens runs through 2026-08-31.
Strengths
- Near-Opus agentic-coding reliability at Sonnet cost
- Best-in-Sonnet terminal and tool-use performance
- 1M-token context with adaptive thinking
- High-resolution vision
- Excellent price-to-capability for production
Trade-offs
- Opus 4.8 and Fable 5 still lead the hardest long-horizon jobs
- 128K output cap
- No audio or video input
- New tokenizer counts ~30% more tokens than Sonnet 4.6
Fit
Best for
- Production coding agents and Claude Code workflows
- Terminal and tool-use automation
- High-volume developer workflows
- Everyday long-context work
Not ideal for
- The very hardest autonomous coding (prefer Opus 4.8)
- Ultra-low-cost bulk classification
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 92
- Coding
- 92
- Math
- 90
- Writing
- 92
- Knowledge
- 91
- Speed
- 70
- Agentic
- 91
- Vision
- 88
- Multilingual
- 90
- Long Context
- 94
Benchmarks
04 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
53/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 53%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 63.2%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 91.1%
- AIME 2025Math · Vendor-reported
- 100%
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 05 of 41