Model

Gemini 3.1 Pro NO.17

Top-tier reasoning, knowledge and vision with a 1M-token multimodal context.

Google DeepMind · Gemini · Frontier · Closed weights

Specification

Fig.01Gemini — Google DeepMind

USA

Context window
1.0Mtokens
Max output
66Ktokens
Input price
$2/ 1M
Output price
$12/ 1M
Cached input
$0.20/ 1M
Throughput
Class
Frontier
Modalities
Text, Image, Audio, Video, Pdf
Weights
Closed
Knowledge cutoff
Jan 2025
Released
Feb 2026

Overview

Gemini 3.1 Pro is Google DeepMind's flagship, benchmarking at or near the very top on knowledge, math, science and multimodal understanding, with an enormous 1M-token context window. It is a superb research, analysis and vision model — but in long autonomous coding sessions it is materially less reliable than the Claude/GPT frontier, and can derail, loop, or corrupt a codebase over a few prompts. Excellent for knowledge work and multimodal reasoning; use with supervision for extended agentic coding.

Strengths

  • Elite math, science and knowledge benchmarks
  • Best-in-class multimodal (image/video/PDF) understanding
  • 1M-token context for huge documents and codebases
  • Strong single-shot reasoning and planning

Trade-offs

  • Less reliable than Claude/GPT on long agentic coding
  • Can derail, loop, or corrupt a codebase over multiple prompts
  • Higher latency than the Flash tiers
  • Output capped at ~64K tokens

Fit

Best for

  • Deep research and analysis
  • Multimodal reasoning over images, video and PDFs
  • Long-context document synthesis
  • Math and science problem solving

Not ideal for

  • Unsupervised long-horizon coding agents
  • Latency-sensitive high-volume traffic

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
90
Coding
78
Math
94
Writing
89
Knowledge
94
Speed
60
Agentic
76
Vision
95
Multilingual
92
Long Context
96

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

Independent index

Artificial Analysis Intelligence Index

Composite of ~9–10 independent evals · Artificial Analysis

46/ 100

BenchmarkResult
Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
46%
SWE-bench VerifiedIndependentCoding · SWE-bench
80.6%
SWE-bench ProIndependentCoding · SWE-bench Pro
46.1%
τ-benchIndependentAgentic · τ-bench
76.5%
GPQA DiamondIndependentReasoning · Independent aggregators
94.3%
MMLU-ProKnowledge · Vendor-reported
80.5%
AIME 2025Math · Vendor-reported
95%
MMMUVision · Vendor-reported
81%
LiveBenchIndependentGeneral · LiveBench
79.9%
LMArena EloIndependentGeneral · LMArena
1505Elo

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION