The model guide

Pick the right model, every time.

GPT-5, Claude, Gemini, GLM, Kimi, DeepSeek — the lineup keeps growing and none of them are interchangeable. Side-by-side specs, independent benchmarks, and plain-English guidance on which model fits which job.

41 models · 11 labs · 10M max context

Flat isometric illustration of a specimen cabinet on a hexagonal plinth: drawers pulled out at staggered depths hold rows of cobalt blocks in graduated shades while a sorting arm lifts one clear and pixel cubes fall away into dither

Fig.01The model cabinet

Catalogued labs, tracked continuously

ClaudeGPTGeminiGLMKimiDeepSeek+ 5 more

Why it matters

Picking a model is a real decision.

Coder or writer, analyst or student — the same prompt can cost cents or dollars, fit or overflow, answer instantly or stall. Four axes decide most of it.

  • 01Cost

    947×

    Price swings are enormous

    Output token prices span 947× across the catalog. Pick wrong and a high-volume job costs an order of magnitude more than it should.

  • 02Context

    10M

    Windows differ wildly

    From tight, cheap windows up to 10M tokens. The right size depends on whether you feed in a paragraph, a book, or an entire codebase.

  • 03Openness

    17

    Open vs. closed weights

    17 tracked models ship open weights you can self-host for privacy and control — the rest are API-only. The trade-off shapes your whole setup.

  • 04Latency

    250tps

    Latency is a feature

    Throughput runs from deliberate reasoners to models pushing 250 tokens per second. A live back-and-forth and an overnight batch want opposite ends of that scale.

How to choose

Context quiz

Think you can call the right model?

The context quiz drops you into real, constrained scenarios — budget, latency, volume, privacy — and grades your pick against the trade-offs. You learn the reasoning, not just the answer.

14 scenarios · ~5 min · Graded on trade-offs

Scenario

easy

A near-free pocket assistant for everyday questions

You're launching a free consumer app that answers quick everyday questions — recipe swaps, plain-English explanations, trip ideas, unit conversions. It's plain text, the quality bar is 'feels smart enough,' replies must stream back instantly, and with millions of tiny queries a day the cost per answer has to be almost nothing.

Which would you reach for?

  • GPT-5.4 nanoGPT
  • Gemini 3.1 Flash-LiteGemini
  • DeepSeek-V3.2DeepSeek
  • Claude Haiku 4.5Claude

By the numbers

Computed live from the catalog

Models tracked
41

Models tracked

across 11 labs

Output price / 1M
$0.19–$180

Output price / 1M

a 947× spread

Largest context
10M

Largest context

tokens in one call

Open-weight models
17

Open-weight models

self-hostable

The labs

Every major lab, under one roof.

11 labs across 3 countries — from closed frontier models to self-hostable open weights. We track them all, so you can weigh a Claude against a Kimi without opening sixteen tabs.

  • OpenAIGPT · 9 models
  • AnthropicClaude · 6 models
  • Alibaba QwenQwen · 4 models
  • DeepSeekDeepSeek · 4 models
  • Google DeepMindGemini · 4 models
  • Mistral AIMistral · 4 models
  • Moonshot AIKimi · 3 models
  • Meta AILlama · 2 models
  • xAIGrok · 2 models
  • Zhipu AIGLM · 2 models
  • MiniMaxMiniMax · 1 model

Start matching

Start here

Stop guessing. Start matching.

Compare any two models head-to-head, or browse the full catalog and filter to exactly what your task needs.

41 models · 11 labs · One decision