Model
GPT-5.6 Sol NO.07
OpenAI's GPT-5.6 flagship — its best coding-and-agentic model yet, for the hardest work.
OpenAI · GPT · Frontier · Closed weights
Specification
11 fields
Fig.01 — GPT — OpenAI
USA
- Context window
- 1.1Mtokens
- Max output
- 128Ktokens
- Input price
- $5/ 1M
- Output price
- $30/ 1M
- Cached input
- $0.50/ 1M
- Throughput
- —
- Class
- Frontier
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Feb 2026
- Released
- Jul 2026
Overview
GPT-5.6 Sol is the flagship of OpenAI's July-2026 GPT-5.6 family and its most capable model — a genuine reliability-frontier peer to Claude for hard, long-horizon agentic coding, security research, and complex reasoning. Sol leads OpenAI's agentic evals (Artificial Analysis' Coding Agent Index ≈80, Terminal-Bench 2.1 88.8%) and sets a new high on Agents' Last Exam (53.6, ahead of Claude Fable 5). It shares the family's ~1.05M-token context, 128K max output, and February-2026 knowledge, and adds programmatic tool calling in the Responses API plus more predictable prompt caching (explicit cache breakpoints, a 30-minute minimum cache life). One wrinkle worth knowing: its SWE-bench Pro score (64.6%) trails the top coding specialists, so its edge is clearest on long-horizon agentic and terminal work rather than that specific eval.
Strengths
- Reliability-frontier agentic coding and tool use
- Leads OpenAI's agentic evals (Coding Agent Index ≈80, Terminal-Bench 2.1 88.8%)
- New state-of-the-art on Agents' Last Exam (53.6)
- ~1.05M-token context with programmatic tool calling
- More predictable prompt caching with explicit breakpoints
Trade-offs
- Expensive output tokens at $30/M
- SWE-bench Pro (64.6%) trails the top coding specialists
- Slower than the Terra and Luna tiers
- No native audio or video input
Fit
Best for
- Hardest autonomous coding and agentic workflows
- Security research and complex engineering
- Long-horizon tool-using agents
- High-stakes reasoning where correctness dominates
- Full-context analysis across ~1M tokens
Not ideal for
- Ultra-cheap high-volume classification
- Latency-critical realtime chat
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 96
- Coding
- 96
- Math
- 95
- Writing
- 92
- Knowledge
- 94
- Speed
- 55
- Agentic
- 96
- Vision
- 89
- Multilingual
- 91
- Long Context
- 93
Benchmarks
03 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
58/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 58%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 64.6%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 94.1%
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 07 of 41