Model

GPT-5.6 Sol NO.07

OpenAI's GPT-5.6 flagship — its best coding-and-agentic model yet, for the hardest work.

OpenAI · GPT · Frontier · Closed weights

Specification

Fig.01GPT — OpenAI

USA

Context window
1.1Mtokens
Max output
128Ktokens
Input price
$5/ 1M
Output price
$30/ 1M
Cached input
$0.50/ 1M
Throughput
Class
Frontier
Modalities
Text, Image, Pdf
Weights
Closed
Knowledge cutoff
Feb 2026
Released
Jul 2026

Overview

GPT-5.6 Sol is the flagship of OpenAI's July-2026 GPT-5.6 family and its most capable model — a genuine reliability-frontier peer to Claude for hard, long-horizon agentic coding, security research, and complex reasoning. Sol leads OpenAI's agentic evals (Artificial Analysis' Coding Agent Index ≈80, Terminal-Bench 2.1 88.8%) and sets a new high on Agents' Last Exam (53.6, ahead of Claude Fable 5). It shares the family's ~1.05M-token context, 128K max output, and February-2026 knowledge, and adds programmatic tool calling in the Responses API plus more predictable prompt caching (explicit cache breakpoints, a 30-minute minimum cache life). One wrinkle worth knowing: its SWE-bench Pro score (64.6%) trails the top coding specialists, so its edge is clearest on long-horizon agentic and terminal work rather than that specific eval.

Strengths

  • Reliability-frontier agentic coding and tool use
  • Leads OpenAI's agentic evals (Coding Agent Index ≈80, Terminal-Bench 2.1 88.8%)
  • New state-of-the-art on Agents' Last Exam (53.6)
  • ~1.05M-token context with programmatic tool calling
  • More predictable prompt caching with explicit breakpoints

Trade-offs

  • Expensive output tokens at $30/M
  • SWE-bench Pro (64.6%) trails the top coding specialists
  • Slower than the Terra and Luna tiers
  • No native audio or video input

Fit

Best for

  • Hardest autonomous coding and agentic workflows
  • Security research and complex engineering
  • Long-horizon tool-using agents
  • High-stakes reasoning where correctness dominates
  • Full-context analysis across ~1M tokens

Not ideal for

  • Ultra-cheap high-volume classification
  • Latency-critical realtime chat

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
96
Coding
96
Math
95
Writing
92
Knowledge
94
Speed
55
Agentic
96
Vision
89
Multilingual
91
Long Context
93

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

Independent index

Artificial Analysis Intelligence Index

Composite of ~9–10 independent evals · Artificial Analysis

58/ 100

BenchmarkResult
Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
58%
SWE-bench ProIndependentCoding · SWE-bench Pro
64.6%
GPQA DiamondIndependentReasoning · Independent aggregators
94.1%

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION