Model
GPT-5.5 Pro NO.12
Previous-generation maximum-accuracy OpenAI tier, succeeded by Sol's Ultra mode.
OpenAI · GPT · Reasoning · Closed weights
Specification
10 fields
Fig.01 — GPT — OpenAI
USA
- Context window
- 1.1Mtokens
- Max output
- 128Ktokens
- Input price
- $30/ 1M
- Output price
- $180/ 1M
- Throughput
- —
- Class
- Reasoning
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Dec 2025
- Released
- Apr 2026
Overview
GPT-5.5 Pro is the mandatory-reasoning, maximum-accuracy variant of OpenAI's previous flagship, tuned for the hardest problems where correctness outweighs cost and latency — now succeeded by GPT-5.6 Sol's Pro and Ultra heavy-compute modes. It spends far more compute per query, pushing deeper deliberate reasoning for elite math, science, and analysis. Reserve it for the small slice of tasks that justify its steep pricing; new work should prefer Sol Ultra.
Strengths
- Top-tier deliberate reasoning accuracy
- Elite math and science problem solving
- Frontier coding on hard problems
- Reliable high-stakes single-shot answers
- Million-token context with multimodal input
Trade-offs
- Very expensive at $30/$180 per million tokens
- High latency from heavy reasoning
- Overkill for routine work
- No prompt-cache discount advantage at this tier
Fit
Best for
- Hardest math and science problems
- High-stakes analysis where accuracy is paramount
- Research-grade reasoning
- Complex algorithmic coding
Not ideal for
- Cost-sensitive or high-throughput workloads
- Latency-sensitive interactive apps
- Routine tasks GPT-5.5 already handles
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 98
- Coding
- 95
- Math
- 97
- Writing
- 91
- Knowledge
- 95
- Speed
- 30
- Agentic
- 94
- Vision
- 89
- Multilingual
- 90
- Long Context
- 92
Benchmarks
03 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
- GPQA DiamondIndependentReasoning · Independent aggregators
- 94.5%
- AIME 2025Math · Vendor-reported
- 96.5%
- LMArena EloIndependentGeneral · LMArena
- 1510Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 12 of 41