The model guide
Pick the right model, every time.
GPT-5, Claude, Gemini, GLM, Kimi, DeepSeek — the lineup keeps growing and none of them are interchangeable. Side-by-side specs, independent benchmarks, and plain-English guidance on which model fits which job.
41 models · 11 labs · 10M max context

Fig.01 — The model cabinet
Catalogued labs, tracked continuously
Why it matters
Four decisive axes
Picking a model is a real decision.
Coder or writer, analyst or student — the same prompt can cost cents or dollars, fit or overflow, answer instantly or stall. Four axes decide most of it.
01Cost
947×
Price swings are enormous
Output token prices span 947× across the catalog. Pick wrong and a high-volume job costs an order of magnitude more than it should.
02Context
10M
Windows differ wildly
From tight, cheap windows up to 10M tokens. The right size depends on whether you feed in a paragraph, a book, or an entire codebase.
03Openness
17
Open vs. closed weights
17 tracked models ship open weights you can self-host for privacy and control — the rest are API-only. The trade-off shapes your whole setup.
04Latency
250tps
Latency is a feature
Throughput runs from deliberate reasoners to models pushing 250 tokens per second. A live back-and-forth and an overnight batch want opposite ends of that scale.
Featured models
25 headliners
A cross-section of the frontier
Flagships, fast workhorses, and notable open-weight models — one entry each. Open any model for full specs, benchmarks, and pricing.
How to choose
Four questions
Four questions get you 90% of the way
Most model decisions come down to the same handful of trade-offs. The guide walks each one with concrete examples.
Context quiz
14 scenarios · ~5 min
Think you can call the right model?
The context quiz drops you into real, constrained scenarios — budget, latency, volume, privacy — and grades your pick against the trade-offs. You learn the reasoning, not just the answer.
14 scenarios · ~5 min · Graded on trade-offs
Scenario
easy
A near-free pocket assistant for everyday questions
You're launching a free consumer app that answers quick everyday questions — recipe swaps, plain-English explanations, trip ideas, unit conversions. It's plain text, the quality bar is 'feels smart enough,' replies must stream back instantly, and with millions of tiny queries a day the cost per answer has to be almost nothing.
Which would you reach for?
- GPT-5.4 nanoGPT
- Gemini 3.1 Flash-LiteGemini
- DeepSeek-V3.2DeepSeek
- Claude Haiku 4.5Claude
By the numbers
41 models · 11 labs
Computed live from the catalog
- Models tracked
- 41
- Output price / 1M
- $0.19–$180
- Largest context
- 10M
- Open-weight models
- 17
Models tracked
across 11 labs
Output price / 1M
a 947× spread
Largest context
tokens in one call
Open-weight models
self-hostable
The labs
Every catalogued lab
Every major lab, under one roof.
11 labs across 3 countries — from closed frontier models to self-hostable open weights. We track them all, so you can weigh a Claude against a Kimi without opening sixteen tabs.
- OpenAIGPT · 9 models
- AnthropicClaude · 6 models
- Alibaba QwenQwen · 4 models
- DeepSeekDeepSeek · 4 models
- Google DeepMindGemini · 4 models
- Mistral AIMistral · 4 models
- Moonshot AIKimi · 3 models
- Meta AILlama · 2 models
- xAIGrok · 2 models
- Zhipu AIGLM · 2 models
- MiniMaxMiniMax · 1 model
Start matching
Catalog · Compare
Start here
Stop guessing. Start matching.
Compare any two models head-to-head, or browse the full catalog and filter to exactly what your task needs.
41 models · 11 labs · One decision