Model
Claude Opus 4.8 NO.02
State-of-the-art autonomous coding reliability at half the price of Fable 5.
Anthropic · Claude · Frontier · Closed weights
Specification
11 fields
Fig.01 — Claude — Anthropic
USA
- Context window
- 1Mtokens
- Max output
- 128Ktokens
- Input price
- $5/ 1M
- Output price
- $25/ 1M
- Cached input
- $0.50/ 1M
- Throughput
- —
- Class
- Frontier
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Jan 2026
- Released
- May 2026
Overview
Claude Opus 4.8 is Anthropic's most capable Opus release and a state-of-the-art model for autonomous long-horizon agentic execution and coding reliability. It pairs frontier-grade planning and tool use with the durability to run extended agent loops without corrupting state, and adds parallel-subagent workflows and long-horizon memory. At half the price of Fable 5 it is the default frontier choice for demanding production coding agents.
Strengths
- State-of-the-art autonomous coding reliability
- Frontier long-horizon agentic execution
- Strong price-to-capability at the frontier
- Robust tool use and self-correction
- Parallel subagent workflows and 1M context
Trade-offs
- Pricier than balanced-tier models for everyday tasks
- Not the fastest option on simple jobs
- No audio/video input
- Fable 5 edges it out on the very hardest tasks
Fit
Best for
- Production coding agents and Claude Code workflows
- Long-horizon autonomous workflows
- Complex debugging and refactoring
- Computer-use and browser automation
- Agentic pipelines needing reliability
Not ideal for
- Ultra-low-cost bulk workloads
- Realtime low-latency chat
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 96
- Coding
- 96
- Math
- 92
- Writing
- 93
- Knowledge
- 92
- Speed
- 55
- Agentic
- 96
- Vision
- 88
- Multilingual
- 90
- Long Context
- 95
Benchmarks
08 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
56/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 56%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 88.6%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 69.2%
- τ-benchIndependentAgentic · τ-bench
- 74.6%
- LiveBenchIndependentGeneral · LiveBench
- 77.2%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 93.6%
- Terminal-Bench HardIndependentAgentic · Artificial Analysis
- 58.3%
- LMArena EloIndependentGeneral · LMArena
- 1512Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 02 of 41