Model
GPT-5.5 NO.11
Previous-generation OpenAI flagship — frontier-reliable, now succeeded by GPT-5.6 Sol.
OpenAI · GPT · Frontier · Closed weights
Specification
11 fields
Fig.01 — GPT — OpenAI
USA
- Context window
- 1.1Mtokens
- Max output
- 128Ktokens
- Input price
- $5/ 1M
- Output price
- $30/ 1M
- Cached input
- $0.50/ 1M
- Throughput
- —
- Class
- Frontier
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Dec 2025
- Released
- Apr 2026
Overview
GPT-5.5 was OpenAI's flagship before the July-2026 GPT-5.6 family and remains a genuinely frontier-reliable model for hard, long-horizon agentic coding — now superseded by GPT-5.6 Sol at the same price. It combines strong reasoning, elite tool use, and durable multi-step execution with a roughly 1M-token context, and stays a solid choice for teams standardized on it. New deployments should prefer GPT-5.6 Sol.
Strengths
- Frontier long-horizon agentic reliability
- Elite real-world coding and tool use
- Strong reasoning across domains
- Roughly 1M-token context window
- Deep prompt-cache discount lowers repeat-prompt cost
Trade-offs
- Expensive output tokens at $30/M
- Slower than mini/nano tiers
- >272K-token prompts incur premium long-context pricing
- No native audio/video input
Fit
Best for
- Autonomous coding agents and tool orchestration
- Hard professional and engineering work
- Complex multi-step reasoning
- High-stakes production pipelines
- Long-document analysis across full context
Not ideal for
- Ultra-cheap high-volume classification
- Latency-critical realtime use
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 96
- Coding
- 95
- Math
- 95
- Writing
- 92
- Knowledge
- 94
- Speed
- 55
- Agentic
- 95
- Vision
- 89
- Multilingual
- 91
- Long Context
- 92
Benchmarks
09 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
55/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 55%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 80.6%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 58.6%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 94%
- AIME 2025Math · Vendor-reported
- 95.2%
- LiveBenchIndependentGeneral · LiveBench
- 80.7%
- Terminal-Bench HardIndependentAgentic · Artificial Analysis
- 60.6%
- DeepSWEIndependentCoding · Datacurve
- 70%
- LMArena EloIndependentGeneral · LMArena
- 1506Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 11 of 41