Model
GPT-5.4 NO.13
Previous-generation balanced workhorse, succeeded by GPT-5.6 Terra.
OpenAI · GPT · Balanced · Closed weights
Specification
11 fields
Fig.01 — GPT — OpenAI
USA
- Context window
- 1.1Mtokens
- Max output
- 128Ktokens
- Input price
- $2.50/ 1M
- Output price
- $15/ 1M
- Cached input
- $0.25/ 1M
- Throughput
- —
- Class
- Balanced
- Modalities
- Text, Image, Pdf
- Weights
- Closed
- Knowledge cutoff
- Oct 2025
- Released
- Mar 2026
Overview
GPT-5.4 was OpenAI's balanced workhorse before GPT-5.6 Terra, delivering near-flagship reasoning and coding at half the per-token cost of the GPT-5.5 flagship. It retains genuinely reliable agentic behavior and a roughly 1M-token context with full multimodal input, and remains a fine choice for teams already on it — though GPT-5.6 Terra now offers GPT-5.5-class quality at the same price point.
Strengths
- Strong, reliable all-round reasoning and coding
- Half the price of the flagship
- Roughly 1M-token context window
- Multimodal text, image, and PDF input
- Good balance of speed and quality
Trade-offs
- Trails GPT-5.5 on the hardest tasks
- Pricier than mini/nano for bulk work
- Long-context premium applies above 272K tokens
Fit
Best for
- General-purpose production assistants
- Everyday coding and refactoring
- Document analysis and summarization
- Agentic tasks on a budget
- RAG over large corpora
Not ideal for
- Absolute frontier reasoning (use GPT-5.5 Pro)
- Ultra-cheap high-volume classification (use nano)
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 90
- Coding
- 90
- Math
- 90
- Writing
- 88
- Knowledge
- 90
- Speed
- 62
- Agentic
- 88
- Vision
- 87
- Multilingual
- 88
- Long Context
- 88
Benchmarks
09 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
51/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 51%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 76.9%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 59.1%
- τ-benchIndependentAgentic · τ-bench
- 78.3%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 88.4%
- LiveBenchIndependentGeneral · LiveBench
- 80.3%
- Terminal-Bench HardIndependentAgentic · Artificial Analysis
- 57.6%
- DeepSWEIndependentCoding · Datacurve
- 56%
- LMArena EloIndependentGeneral · LMArena
- 1495Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 13 of 41