Model

GLM-5.2 NO.36

Open-weight coding flagship at a fraction of frontier cost.

Zhipu AI · GLM · Open · Open weights

Specification

Fig.01GLM — Zhipu AI

CHINA

Context window
1.0Mtokens
Max output
131Ktokens
Input price
$1.40/ 1M
Output price
$4.40/ 1M
Cached input
$0.26/ 1M
Throughput
Class
Open
Modalities
Text
Weights
Open
Parameters
753B
License
MIT
Knowledge cutoff
Mar 2026
Released
Jun 2026

Overview

Zhipu's GLM-5.2 is a ~753B-parameter open-weight MoE (~40B active, MIT) that leads open models on Artificial Analysis' Intelligence Index (51 — the top open-weights score) while its Z.ai API list price ($1.40 in / $4.40 out) runs roughly one-sixth the closed frontier, and open-weight hosts price it lower still. It adds selectable High and Max thinking-effort reasoning modes for long-horizon coding. Its vendor-aggregate SWE-bench Pro (62.1) edges GPT-5.5's, though independent contamination-resistant runs of the GLM family land far lower, so its coding lead is softer than the headline suggests. Its 1M-token context holds an entire mid-sized codebase, and it ships as a drop-in Anthropic-compatible backend for Claude Code workflows — a value challenger rather than a frontier-reliable autonomous agent.

Strengths

  • Best-in-class open-weight coding benchmarks
  • 1M-token context for whole-repo work
  • Roughly one-sixth the cost of frontier closed models
  • Selectable High / Max thinking-effort reasoning modes
  • Permissive MIT weights, self-hostable
  • Strong math and reasoning (AIME ~99%)

Trade-offs

  • Long agentic sessions less reliable than Claude/GPT
  • Text-only flagship (vision is the separate GLM-5V)
  • Benchmark scores overstate real-world autonomy
  • Large model — throughput and latency costs

Fit

Best for

  • Cost-efficient coding assistants
  • Whole-repo refactors within a supervised loop
  • Self-hosted / on-prem deployments
  • Claude Code-compatible budget backend
  • Math and reasoning tasks

Not ideal for

  • Unsupervised long-horizon autonomous coding
  • Vision-heavy multimodal tasks

Capabilities

Ten axes, normalized 0–100 and scored the same way across the whole catalog — so a 78 here means what a 78 means anywhere else on the bench.

Fig.02Capability radar

0 — 100

Reasoning
86
Coding
77
Math
88
Writing
82
Knowledge
85
Speed
58
Agentic
75
Vision
10
Multilingual
84
Long Context
84

Benchmarks

Public results, with independent third-party runs marked. Bars normalize percentages against 100 and Elo ratings against a 1500 ceiling.

Independent index

Artificial Analysis Intelligence Index

Composite of ~9–10 independent evals · Artificial Analysis

51/ 100

BenchmarkResult
Artificial Analysis Intelligence IndexIndependentGeneral · Artificial Analysis
51%
SWE-bench ProIndependentCoding · SWE-bench Pro
62.1%
GPQA DiamondIndependentReasoning · Independent aggregators
91.2%
AIME 2025Math · Vendor-reported
99.2%

Alternatives

Models in roughly the same class — the ones worth weighing against this record.

Index

All models/EDUCATION