Model

GPT-5.6 Sol Ultra NO.10

Sol's maximum-compute mode — cooperating sub-agents for the hardest multi-step work.

OpenAI · GPT · Reasoning · Closed weights

Specification

Fig.01GPT — OpenAI

USA

Context window
1.1Mtokens
Max output
128Ktokens
Input price
$5/ 1M
Output price
$30/ 1M
Cached input
$0.50/ 1M
Throughput
Class
Reasoning
Modalities
Text, Image, Pdf
Weights
Closed
Knowledge cutoff
Feb 2026
Released
Jul 2026

Overview

GPT-5.6 Sol Ultra is not a separate model but Sol's heaviest reasoning mode. Where Sol Pro spawns independent parallel agents and merges the best result, Sol Ultra runs four cooperating sub-agents that communicate mid-task and synthesize a joint answer, pushing Sol's ceiling higher — reaching 91.9% on Terminal-Bench 2.1 versus base Sol's 88.8% — for the hardest agentic, coding, and research problems. It runs at Sol's per-token price, but because every sub-agent generates and consumes its own tokens, a single Ultra run costs several times a normal Sol call (OpenAI's own estimate is roughly 3×). Reserve it for the narrow set of tasks where maximum accuracy justifies the spend and the latency.

Strengths

  • Highest OpenAI accuracy via cooperating sub-agents
  • Tops base Sol on Terminal-Bench 2.1 (91.9% vs 88.8%)
  • Frontier reasoning, coding, and agentic execution
  • Same ~1.05M-token context and Responses-API tooling

Trade-offs

  • Effective cost is ~3× a single Sol call (multi-agent tokens)
  • High latency from heavy multi-agent compute
  • Overkill for routine or interactive work
  • No native audio or video input

Fit

Best for

  • The hardest agentic and research problems
  • Maximum-accuracy single-shot answers
  • Complex multi-step engineering
  • High-stakes analysis where cost is secondary

Not ideal for

  • Cost- or latency-sensitive workloads
  • Routine tasks base Sol already handles

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
98
Coding
96
Math
97
Writing
92
Knowledge
94
Speed
25
Agentic
95
Vision
89
Multilingual
91
Long Context
93

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

BenchmarkResult
GPQA DiamondIndependentReasoning · Independent aggregators
94.1%
SWE-bench ProIndependentCoding · SWE-bench Pro
64.6%

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION