Model
Gemini 3.5 Pro NO.16
Google's newest flagship — a 2M-token context and a Deep Think reasoning mode.
Google DeepMind · Gemini · Frontier · Closed weights
Specification
10 fields
Fig.01 — Gemini — Google DeepMind
USA
- Context window
- 2Mtokens
- Max output
- 66Ktokens
- Input price
- $15/ 1M
- Output price
- $60/ 1M
- Throughput
- —
- Class
- Frontier
- Modalities
- Text, Image, Audio, Video, Pdf
- Weights
- Closed
- Knowledge cutoff
- Jan 2026
- Released
- Jul 2026
Overview
Gemini 3.5 Pro is Google DeepMind's newest flagship, unveiled at Google I/O 2026 and reaching general availability in July 2026 with the largest production context window (2M tokens) and a gated 'Deep Think' extended-reasoning mode. It is positioned as the multimodal and ultra-long-context leader. As a brand-new release it is effectively pre-benchmark — Google had not published official evals or final pricing at launch — so the figures here are provisional, and its real-world long-horizon agentic-coding reliability (historically Gemini's weak spot versus the Claude/GPT frontier) remains to be proven independently.
Strengths
- Largest production context window (2M tokens)
- Deep Think extended-reasoning mode
- Natively multimodal (image/audio/video/PDF)
- Frontier knowledge, math, and vision pedigree
Trade-offs
- Pricing and benchmarks unconfirmed at launch (provisional)
- Deep Think gated to top/enterprise tiers
- Gemini agentic-coding reliability historically below Claude/GPT
- Limited-preview rollout at GA
Fit
Best for
- Ultra-long-context document and codebase analysis
- Multimodal reasoning over image, audio, and video
- Deep research and math/science
Not ideal for
- Unsupervised long-horizon coding agents (unproven)
- Cost-sensitive high-volume workloads
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 92
- Coding
- 80
- Math
- 95
- Writing
- 90
- Knowledge
- 95
- Speed
- 55
- Agentic
- 78
- Vision
- 96
- Multilingual
- 92
- Long Context
- 98
Benchmarks
00 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
No public benchmark scores are recorded for this model.
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 16 of 41