Model
DeepSeek-V4 Pro NO.28
Open MIT flagship — 1.6T-parameter MoE with native vision and a 1M-token context at throwaway prices.
DeepSeek · DeepSeek · Open · Open weights
Specification
13 fields
Fig.01 — DeepSeek — DeepSeek
CHINA
- Context window
- 1Mtokens
- Max output
- 384Ktokens
- Input price
- $0.43/ 1M
- Output price
- $0.87/ 1M
- Cached input
- $0.04/ 1M
- Throughput
- —
- Class
- Open
- Modalities
- Text, Image
- Weights
- Open
- Parameters
- 1600B
- License
- MIT
- Knowledge cutoff
- Jan 2026
- Released
- Apr 2026
Overview
DeepSeek-V4 Pro is DeepSeek's open-weight flagship (1.6T-parameter MoE, 49B active, MIT) and the headline of the V4 generation. It expands the context window 8x to 1M tokens, adds native image input, and introduces a refined DeepSeek Sparse Attention (a CSA/HCA hybrid) that cuts long-context compute and KV-cache footprint dramatically. DeepSeek advertises a chart-topping 80.6% on SWE-bench Verified, but independent contamination-resistant runs land far lower (~58%) and it trails Claude Opus on SWE-bench Pro, so it is a spectacular-value workhorse rather than a frontier-reliable autonomous coder. Its API speaks both OpenAI and Anthropic formats, dropping into Claude Code without a proxy.
Strengths
- Among the lowest frontier-class prices in the industry
- 1M-token context via efficient DeepSeek Sparse Attention
- Native image input (new to the V4 generation)
- Very cheap cached input; MIT open weights
- Strong math and general reasoning value
Trade-offs
- Vendor benchmarks overstate real-world coding reliability
- Agentic-coding reliability below the Claude/GPT frontier
- Vision quality trails dedicated multimodal flagships
- Large 1.6T MoE — serving cost and latency
Fit
Best for
- High-volume, cost-sensitive workloads at scale
- Long-context document and repo processing
- Math and reasoning tasks
- Budget Claude Code-compatible backend
- Self-hosted / on-prem deployments
Not ideal for
- Unsupervised long-horizon autonomous coding
- Mission-critical frontier reasoning
Capabilities
Normalized 0—100
Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.
Fig.02 — Capability radar
0 — 100
- Reasoning
- 85
- Coding
- 73
- Math
- 91
- Writing
- 81
- Knowledge
- 85
- Speed
- 60
- Agentic
- 69
- Vision
- 66
- Multilingual
- 83
- Long Context
- 87
Benchmarks
09 results
Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.
Independent index
Artificial Analysis Intelligence Index
Composite of ~9–10 independent evals · Artificial Analysis
54/ 100
- Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
- 54%
- SWE-bench VerifiedIndependentCoding · SWE-bench
- 80.6%
- SWE-bench ProIndependentCoding · SWE-bench Pro
- 55.4%
- GPQA DiamondIndependentReasoning · Independent aggregators
- 83%
- AIME 2025Math · Vendor-reported
- 93%
- LiveCodeBenchIndependentCoding · LiveCodeBench
- 78%
- MMLU-ProKnowledge · Vendor-reported
- 86%
- DeepSWEIndependentCoding · Datacurve
- 8%
- LMArena EloIndependentGeneral · LMArena
- 1462Elo
Alternatives
03 comparable
Models in roughly the same class — the ones worth weighing against this record.
Index
Record 28 of 41