Model
Grok 4.3 NO.20
xAI's flagship: strong math reasoning and real-time knowledge at aggressive pricing.
xAI · Grok · Balanced · Closed weights
Specification
11 fields
Fig.01 — Grok — xAI
USA
- Context window
- 1Mtokens
- Max output
- 66Ktokens
- Input price
- $1.25/ 1M
- Output price
- $2.50/ 1M
- Cached input
- $0.20/ 1M
- Throughput
- 127tok/s
- Class
- Balanced
- Modalities
- Text, Image, Video, Pdf
- Weights
- Closed
- Knowledge cutoff
- Dec 2025
- Released
- Apr 2026
Overview
Grok 4.3 is xAI's current flagship, a large reasoning model with a 1M-token context window and native access to real-time X/web data. It benchmarks near the top on math and science and is a genuinely strong general model. In practice its long-horizon agentic coding is less consistent than the Claude/GPT frontier — it can be verbose, over-confident, and drift on multi-step tool use.
Strengths
- Excellent math and scientific reasoning
- Real-time knowledge via X/web integration
- Native video input (new in 4.3)
- Large 1M-token context window
- Competitive pricing for a flagship
Trade-offs
- Long agentic coding less reliable than Claude/GPT
- Verbose and occasionally over-confident output
- Tool-use consistency lags the frontier
- Tone and moderation can be polarizing
Fit
Best for
- Math and STEM reasoning
- Real-time research and current events
- General chat and analysis
- Long-context document work
Not ideal for
- Unsupervised long-horizon agentic coding
- Safety-critical automated workflows
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 84
- Coding
- 74
- Math
- 90
- Writing
- 82
- Knowledge
- 85
- Speed
- 68
- Agentic
- 73
- Vision
- 80
- Multilingual
- 82
- Long Context
- 84
Benchmarks
07 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
38/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 38%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 76.7%
- τ-benchIndependentAgentic · τ-bench
- 78.9%
- AIME 2025Math · Vendor-reported
- 100%
- MMLU-ProKnowledge · Vendor-reported
- 91.2%
- Terminal-Bench HardIndependentAgentic · Artificial Analysis
- 37.9%
- LMArena EloIndependentGeneral · LMArena
- 1496Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 20 of 41