Model
GPT-5.6 Sol Ultra NO.10
Sol's maximum-compute mode — cooperating sub-agents for the hardest multi-step work.
OpenAI · GPT · Reasoning · Closed weights
Specification
11 fields
Fig.01 — GPT — OpenAI
USA
- Context window
- 1.1Mtokens
- Max output
- 128Ktokens
- Input price
- $5/ 1M
- Output price
- $30/ 1M
- Cached input
- $0.50/ 1M
- Throughput
- —
- Class
- Reasoning
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Feb 2026
- Released
- Jul 2026
Overview
GPT-5.6 Sol Ultra is not a separate model but Sol's heaviest reasoning mode. Where Sol Pro spawns independent parallel agents and merges the best result, Sol Ultra runs four cooperating sub-agents that communicate mid-task and synthesize a joint answer, pushing Sol's ceiling higher — reaching 91.9% on Terminal-Bench 2.1 versus base Sol's 88.8% — for the hardest agentic, coding, and research problems. It runs at Sol's per-token price, but because every sub-agent generates and consumes its own tokens, a single Ultra run costs several times a normal Sol call (OpenAI's own estimate is roughly 3×). Reserve it for the narrow set of tasks where maximum accuracy justifies the spend and the latency.
Strengths
- Highest OpenAI accuracy via cooperating sub-agents
- Tops base Sol on Terminal-Bench 2.1 (91.9% vs 88.8%)
- Frontier reasoning, coding, and agentic execution
- Same ~1.05M-token context and Responses-API tooling
Trade-offs
- Effective cost is ~3× a single Sol call (multi-agent tokens)
- High latency from heavy multi-agent compute
- Overkill for routine or interactive work
- No native audio or video input
Fit
Best for
- The hardest agentic and research problems
- Maximum-accuracy single-shot answers
- Complex multi-step engineering
- High-stakes analysis where cost is secondary
Not ideal for
- Cost- or latency-sensitive workloads
- Routine tasks base Sol already handles
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 98
- Coding
- 96
- Math
- 97
- Writing
- 92
- Knowledge
- 94
- Speed
- 25
- Agentic
- 95
- Vision
- 89
- Multilingual
- 91
- Long Context
- 93
Benchmarks
02 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
- GPQA DiamondIndependentReasoning · Independent aggregators
- 94.1%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 64.6%
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 10 of 41