GLM-5.2

glm-5.2

GLM-5.2 is Z.ai’s flagship foundation model for long-horizon engineering tasks, agent workflows, and complex reasoning. Its usable 1M-token context window supports project-scale codebases, system documentation, and sustained multi-step execution. Use the GLM API through TokenHub to evaluate current pricing, available endpoints, context limits, and benchmark results for workflows spanning requirements analysis, repository understanding, implementation, testing, and deployment.

Context Window

1M tokens

Maximum Output

131.1K tokens

Release Date

Jun 13, 2026

Modalities

GLM-5.2 Pricing

Input PriceOutput PriceCache Read
$1.1429/M$4/M$0.2857/M

GLM-5.2 API Capabilities

Reasoning

Supported

Tool calling

Supported

Temperature parameter

Supported

Attachments

Not supported

Knowledge Base

—

Endpoint Protocols

Messages APIgeminiCompletions API

GLM-5.2 Model Highlights

GLM-5.2 is an open-weight text model designed for million-token context, long-horizon coding agents and stable delivery of complex software projects.

Stable Million-Token Context

Maintains useful performance across a one-million-token context so complete projects and extended task histories can remain in one reasoning chain.

Long-Horizon Agentic Coding

Sustains project decomposition, architecture design, implementation, testing, repair and deployment across extended autonomous workflows.

Production-Grade Coding

Improved frontend, backend, mobile and deep-debugging capabilities help it follow engineering constraints across complete application workflows.

GLM-5.2 Use Cases

GLM-5.2 suits end-to-end application delivery, repository-scale engineering and long-running research or optimization tasks.

End-to-End Application Delivery

Turn product requirements into architecture, frontend and backend code, tests, fixes and deployment-ready application artifacts.

Large-Project Refactoring

Keep an extensive project in context, trace dependencies and implement coordinated changes without losing established engineering constraints.

Long-Running Technical Optimization

Run extended cycles of analysis, implementation, measurement and correction for performance engineering or automated research.

How to Use GLM-5.2 via the TokenHub API

Create API key
curl 'https://us-api.tokenhub.com/v1/messages' \
  -X 'POST' \
  -H "Authorization: Bearer $TOKENHUB_API_KEY"

GLM-5.2 Benchmarks

Index score
Artificial Analysis Intelligence IndexArtificial Analysis broad capability aggregate51.1
Artificial Analysis Coding IndexArtificial Analysis software task aggregate68.8
Knowledge & Reasoning
GPQAAdvanced science problem solving89.5%
HLEBroad expert-level exam set40.1%
Coding & Engineering
SciCodeScientific coding challenges50.5%
Terminal-Bench HardHard terminal task execution50.8%
Instruction Following & Agent Tasks
IFBenchPrompt constraint adherence73.3%
AA-LCRLong-context reasoning71.3%
τ²-BenchAgent workflow tasks99.1%

Metrics sourced from Artificial Analysis

GLM-5.2 Comparisons

Media and Discussions

Selected public videos and posts related to this model.

X (Twitter)

View post on X
View post on X
View post on X

Reddit

YouTube

Watch on YouTube
Watch on YouTube
Watch on YouTube

GLM-5.2 API FAQs

Answers to common questions about the GLM-5.2 API, Z.ai pricing, coding and agent workflows, context limits, and model comparisons on TokenHub.

What is GLM-5.2?+

GLM-5.2 is a Z.ai model designed for long-horizon engineering tasks, complex reasoning, and agent workflows. It is intended for work such as understanding a codebase, planning changes, writing code, testing, and continuing a multi-step task with substantial project context.

How good is GLM-5.2 for coding and agent workflows?+

GLM-5.2 is a strong candidate when a task requires repository-scale context, repeated tool use, or sustained planning across several steps. Before production use, validate its benchmark results, endpoint support, latency, and tool-calling behavior against your own coding or agent workload.

How do I use the GLM-5.2 API through TokenHub?+

Create a TokenHub API key, copy the exact GLM-5.2 model ID shown on this page, and select a supported endpoint in the API section. Use the generated request example as the starting point for your OpenAI-compatible or other supported client integration.

How should I evaluate GLM-5.2 API pricing?+

Review the current input, output, and cache-read prices displayed on this page, then estimate cost using your typical prompt size, response length, request volume, and cache-hit rate. Long-context and agent workflows can have a very different cost profile from short chat requests.

What are the GLM-5.2 context and output limits?+

Check the context length and maximum output values in the model specifications on this page. These limits should be evaluated together: a large context window helps with repositories and documentation, while maximum output determines how much the model can return in one response.

GLM-5.2 vs DeepSeek V4 Flash: which should I choose?+

Start with the workload. Evaluate GLM-5.2 for long-running engineering and agent tasks; evaluate DeepSeek V4 Flash when its current price, output limit, and supported endpoints better fit high-throughput text generation or retrieval workflows. Use the comparison page to verify the live specifications before deciding.

GLM-5.2 vs Qwen3.8 Max: which should I choose?+

There is no universal winner. Compare the two models against the actual task: coding and agent reliability, reasoning quality, context and output limits, pricing, and endpoint compatibility. Run a small evaluation with representative prompts before standardizing a production workflow on either GLM-5.2 or Qwen3.8 Max.

Ready to use GLM-5.2?

Use one API key to access GLM-5.2 and more AI models through TokenHub.

Create API key