Model
Gemini 3.1 Pro NO.17
Top-tier reasoning, knowledge and vision with a 1M-token multimodal context.
Google DeepMind · Gemini · Frontier · Closed weights
Specification
11 fields
Fig.01 — Gemini — Google DeepMind
USA
- Context window
- 1.0Mtokens
- Max output
- 66Ktokens
- Input price
- $2/ 1M
- Output price
- $12/ 1M
- Cached input
- $0.20/ 1M
- Throughput
- —
- Class
- Frontier
- Modalities
- Text, Image, Audio, Video, Pdf
- Weights
- Closed
- Knowledge cutoff
- Jan 2025
- Released
- Feb 2026
Overview
Gemini 3.1 Pro is Google DeepMind's flagship, benchmarking at or near the very top on knowledge, math, science and multimodal understanding, with an enormous 1M-token context window. It is a superb research, analysis and vision model — but in long autonomous coding sessions it is materially less reliable than the Claude/GPT frontier, and can derail, loop, or corrupt a codebase over a few prompts. Excellent for knowledge work and multimodal reasoning; use with supervision for extended agentic coding.
Strengths
- Elite math, science and knowledge benchmarks
- Best-in-class multimodal (image/video/PDF) understanding
- 1M-token context for huge documents and codebases
- Strong single-shot reasoning and planning
Trade-offs
- Less reliable than Claude/GPT on long agentic coding
- Can derail, loop, or corrupt a codebase over multiple prompts
- Higher latency than the Flash tiers
- Output capped at ~64K tokens
Fit
Best for
- Deep research and analysis
- Multimodal reasoning over images, video and PDFs
- Long-context document synthesis
- Math and science problem solving
Not ideal for
- Unsupervised long-horizon coding agents
- Latency-sensitive high-volume traffic
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 90
- Coding
- 78
- Math
- 94
- Writing
- 89
- Knowledge
- 94
- Speed
- 60
- Agentic
- 76
- Vision
- 95
- Multilingual
- 92
- Long Context
- 96
Benchmarks
10 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
46/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 46%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 80.6%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 46.1%
- τ-benchIndependentAgentic · τ-bench
- 76.5%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 94.3%
- MMLU-ProKnowledge · Vendor-reported
- 80.5%
- AIME 2025Math · Vendor-reported
- 95%
- MMMUVision · Vendor-reported
- 81%
- LiveBenchIndependentGeneral · LiveBench
- 79.9%
- LMArena EloIndependentGeneral · LMArena
- 1505Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 17 of 41