Z.ai
GLM-5.3-FlashX
Total Context
1M
Max Output
131.1K
Released
N/A
glm-5.1GLM-5.1 is described as a major upgrade for coding, long-horizon tasks, and agentic engineering. Source materials emphasize the ability to plan, execute, and improve work over extended periods, producing more complete engineering-grade results. It should be positioned as a developer and agent model for sustained software work rather than a general Q&A model.
Context Window
200K tokens
Maximum Output
131.1K tokens
Release Date
Mar 27, 2026
Modalities
| Token Tier | Input Price | Output Price | Cache Read | Cache Create 5m | Cache Read 5m |
|---|---|---|---|---|---|
| <=32K | $0.8571/M | $3.4286/M | $0.1714/M | $1.0714/M | $0.0857/M |
| >32K | $1.1429/M | $4/M | $0.2286/M | $1.4286/M | $0.1143/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
GLM-5.1 is a text-only flagship model built for sustained autonomous execution, agentic engineering and complex repository or terminal tasks.
Can work autonomously on a single task for up to eight hours, covering planning, execution, iterative optimization and final delivery.
Improved coding and sustained execution support autonomous agents that must plan, implement, evaluate and revise engineering work.
Handles repository generation and real-world terminal workloads that require understanding project structure and execution feedback.
GLM-5.1 is suited to extended coding-agent sessions, repository construction and iterative engineering optimization driven by execution feedback.
Plan and implement a substantial feature over a long session, run checks and refine the result until it meets the stated requirements.
Translate a software specification into a multi-file repository with coherent architecture, implementation and supporting configuration.
Run commands, inspect errors or benchmark results and repeatedly revise an implementation to improve correctness or performance.
import OpenAI from "openai"
const client = new OpenAI({
apiKey: process.env.TOKENHUB_API_KEY,
baseURL: "https://us-api.tokenhub.com/v1",
})
const result = await client.chat.completions.create({})
console.log(result.choices[0]?.message?.content)| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 40.2 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 43.4 |
| Knowledge & Reasoning | ||
| GPQA | Advanced science problem solving | 86.8% |
| HLE | Broad expert-level exam set | 28.0% |
| Coding & Engineering | ||
| SciCode | Scientific coding challenges | 43.8% |
| Terminal-Bench Hard | Hard terminal task execution | 43.2% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 76.3% |
| AA-LCR | Long-context reasoning | 62.3% |
| τ²-Bench | Agent workflow tasks | 97.7% |
Metrics sourced from Artificial Analysis
GLM-5.1: capabilities, use cases, limits, and TokenHub guidance.
GLM-5.1 is a Z.AI model for sustained autonomous engineering and iterative delivery.
Best for iterative engineering delivery, complex coding and tool-heavy automation, especially when reliable multi-step execution is the priority.
Key strength: sustained autonomous planning, execution, testing, and refinement.
It belongs to an older generation and may lack newer capabilities. For the latest capabilities matter, consider GLM-5.2.
Use the exact ID shown by TokenHub; follow your account docs and verify current features.
Use one API key to access GLM-5.1 and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube