Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

Z.AI · OPEN WEIGHTS

GLM-5.2

Benchmarks, pricing and tradeoffs for agent builders.
Tested configuration: source default.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

Add to a custom comparison ↗
Agent Fit68.5General agent profile
Practical Value70.170% capability · 30% cost
Estimated monthly API cost$10.581,000 calls · 10K input + 2K output
Catalog context window1.05MAdvertised capacity

Capability profile

Explain these tests ↗
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.78.6 / 100
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.79.7 / 100
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.51.8 / 100
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.89.8 / 100
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.73.7 / 100
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.76.2 / 100
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.62.3 / 100

What the evidence says

GLM-5.2 scores 73.2 on the equal-weight LiveBench category average. Its highest measured category is mathematics (89.8); its lowest is agentic coding (51.8). Categories differ in difficulty, so these are test results, not universal strengths or weaknesses.

Our General agent profile produces an Agent Fit score of 68.5 by emphasizing reasoning, instruction following, agentic coding, coding and data analysis. See the exact weights or adjust the profile in the interactive comparison.

Independent preference evidence

A matching Arena configuration is not verified in this snapshot. Evidence Consensus is withheld. A result for another reasoning setting or version is not substituted.

API features to verify

The catalog lists tool-calling support and does not list image input. These are declared features, not evaluations of tool accuracy or visual understanding.

Weights are publicly available for this model family. Review the exact model license and serving requirements before deployment. Hosted API costs below are not self-hosting estimates.

GLM-5.2 API pricing

Input / 1M tokens$0.6496
Output / 1M tokens$2.04

At the baseline of 1,000 calls, 10,000 input tokens and 2,000 billed output tokens per call, the text API estimate is $10.58. Reasoning belongs in billed output. Caching, tools, retries, long-context premiums, service tiers, hosting and taxes are excluded.

Trace this model to the source

LiveBench row ID
glm-5.2
Question-set release
2026-06-25 · Source CSV ↗
Arena configuration
No verified matching observation
Reviewed snapshot
2026-09-23

Compare alternatives

For production decisions, test representative tasks with a fixed harness and success criteria. Neither a large context window nor a strong benchmark score guarantees reliability in your own agent.

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗