Model

Llama 4 Scout NO.23

Ultra-long 10M-token context, efficient single-GPU open model.

Meta AI · Llama · Open · Open weights

Specification

Fig.01Llama — Meta AI

USA

Context window
10Mtokens
Max output
16Ktokens
Input price
$0.10/ 1M
Output price
$0.30/ 1M
Throughput
Class
Open
Modalities
Text, Image
Weights
Open
Parameters
109B
License
Llama 4 Community License
Knowledge cutoff
Aug 2024
Released
Apr 2025

Overview

Llama 4 Scout is the smaller, faster member of the Llama 4 herd (109B-total / 17B-active MoE) that fits on a single high-end GPU and ships a headline 10M-token context window. It is ideal for cheap long-document and retrieval workloads, though effective reasoning quality over the full window is far below the nominal length. A pragmatic open model for scale, not for hard agentic work.

Strengths

  • Industry-leading 10M-token context window
  • Runs on a single GPU / very cheap
  • Open weights and fine-tunable
  • Fast throughput for its class

Trade-offs

  • Reasoning/coding well below the frontier
  • Effective recall degrades long before 10M tokens
  • Weak on autonomous agentic tasks
  • Vision is basic

Fit

Best for

  • Long-document summarization and RAG
  • High-throughput cheap inference
  • On-prem/self-hosted deployments
  • Fine-tuning experiments

Not ideal for

  • Agentic coding
  • Complex multi-step reasoning

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
58
Coding
52
Math
56
Writing
66
Knowledge
70
Speed
88
Agentic
50
Vision
64
Multilingual
76
Long Context
82

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

Independent index

Artificial Analysis Intelligence Index

Composite of ~9–10 independent evals · Artificial Analysis

10/ 100

BenchmarkResult
Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
10%
τ-benchIndependentAgentic · τ-bench
62.3%
MMLU-ProKnowledge · Vendor-reported
74.3%
GPQA DiamondIndependentReasoning · Independent aggregators
57.2%
MMMUVision · Vendor-reported
69.4%

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION