Model
Claude Sonnet 4.6 NO.04
The production workhorse — best balance of speed and intelligence.
Anthropic · Claude · Balanced · Closed weights
Specification
11 fields
Fig.01 — Claude — Anthropic
USA
- Context window
- 1Mtokens
- Max output
- 64Ktokens
- Input price
- $3/ 1M
- Output price
- $15/ 1M
- Cached input
- $0.30/ 1M
- Throughput
- —
- Class
- Balanced
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Nov 2025
- Released
- Feb 2026
Overview
Claude Sonnet 4.6 offers the best balance of speed and intelligence in the Claude lineup and is the default production workhorse for most agentic and coding workloads. It handles the large majority of real tasks reliably at a fraction of Opus pricing, escalating to Opus or Fable only for the hardest jobs. Its 1M context and strong tool use make it a dependable everyday driver.
Strengths
- Excellent speed-to-intelligence balance
- Reliable everyday agentic coding
- 1M context at a moderate price
- Strong tool use and instruction following
Trade-offs
- Below Opus/Fable on the hardest long-horizon jobs
- 64K max output cap
- No audio/video input
Fit
Best for
- Everyday production coding agents
- High-volume developer workflows
- Chat and assistant backends
- RAG and document analysis
Not ideal for
- The most demanding autonomous coding runs
- Tasks needing very long single outputs
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 89
- Coding
- 89
- Math
- 86
- Writing
- 90
- Knowledge
- 90
- Speed
- 78
- Agentic
- 88
- Vision
- 84
- Multilingual
- 89
- Long Context
- 92
Benchmarks
07 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
47/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 47%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 79.6%
- τ-benchIndependentAgentic · τ-bench
- 87.5%
- LiveBenchIndependentGeneral · LiveBench
- 75.5%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 89.9%
- DeepSWEIndependentCoding · Datacurve
- 32%
- LMArena EloIndependentGeneral · LMArena
- 1467Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 04 of 41