Model

Grok 4.3 NO.20

xAI's flagship: strong math reasoning and real-time knowledge at aggressive pricing.

xAI · Grok · Balanced · Closed weights

Specification

Fig.01Grok — xAI

USA

Context window
1Mtokens
Max output
66Ktokens
Input price
$1.25/ 1M
Output price
$2.50/ 1M
Cached input
$0.20/ 1M
Throughput
127tok/s
Class
Balanced
Modalities
Text, Image, Video, Pdf
Weights
Closed
Knowledge cutoff
Dec 2025
Released
Apr 2026

Overview

Grok 4.3 is xAI's current flagship, a large reasoning model with a 1M-token context window and native access to real-time X/web data. It benchmarks near the top on math and science and is a genuinely strong general model. In practice its long-horizon agentic coding is less consistent than the Claude/GPT frontier — it can be verbose, over-confident, and drift on multi-step tool use.

Strengths

  • Excellent math and scientific reasoning
  • Real-time knowledge via X/web integration
  • Native video input (new in 4.3)
  • Large 1M-token context window
  • Competitive pricing for a flagship

Trade-offs

  • Long agentic coding less reliable than Claude/GPT
  • Verbose and occasionally over-confident output
  • Tool-use consistency lags the frontier
  • Tone and moderation can be polarizing

Fit

Best for

  • Math and STEM reasoning
  • Real-time research and current events
  • General chat and analysis
  • Long-context document work

Not ideal for

  • Unsupervised long-horizon agentic coding
  • Safety-critical automated workflows

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
84
Coding
74
Math
90
Writing
82
Knowledge
85
Speed
68
Agentic
73
Vision
80
Multilingual
82
Long Context
84

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

Independent index

Artificial Analysis Intelligence Index

Composite of ~9–10 independent evals · Artificial Analysis

38/ 100

BenchmarkResult
Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
38%
SWE-bench VerifiedIndependentCoding · SWE-bench
76.7%
τ-benchIndependentAgentic · τ-bench
78.9%
AIME 2025Math · Vendor-reported
100%
MMLU-ProKnowledge · Vendor-reported
91.2%
Terminal-Bench HardIndependentAgentic · Artificial Analysis
37.9%
LMArena EloIndependentGeneral · LMArena
1496Elo

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION