Production routeThis rate reflects platform-side availability — only confirmed server errors (HTTP 500 / empty response) count as failures. User-side issues (content moderation, invalid params, cancellation) plus rate limits, timeouts and auth errors are excluded. Before real traffic arrives, empty buckets may display as available.Live
Live
Context
1M context · 131K max output
Best For
Coding agents, long task chains, tool-heavy workflows
Z.ai's flagship reasoning model for coding agents and long-horizon engineering work, with a 1M-token context, always-on reasoning, tool calling, and prompt caching over both Chat Completions and Anthropic Messages.
GLM-5.3
Z.ai flagship reasoning model
Selected
Model ID
glm-5.3
Best for
Repository-scale coding agents, long-running task chains, tool-heavy workflows, and stable-prefix agent loops where cache reads carry most of the input volume.
Input
$1.059 / 1M-10%
72 cr / 1M$1.177official price
Cache read
$0.265 / 1M-10%
18 cr / 1M$0.295official price
Output
$3.706 / 1M-10%
252 cr / 1M$4.118official price
All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing.
GLM-5.3 pricing
Estimate what one GLM-5.3 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.
Request calculator
Enter the token mix for one request and the number of successful tool calls.
Only successful server-side calls are billed per call; failed attempts have no tool fee, but tokens still apply.
Web search$0.010/ call·0.68 cr / call
GLM-5.3 API for coding agents and long-horizon engineering
Call Z.ai's flagship reasoning model for $1.06 input and $3.71 output per 1M tokens through EvoLink's unified API. Use model ID glm-5.3 with a 1M-token context, always-on low/high/max reasoning, tool calling, and prompt caching.
GLM-5.3 is served on EvoLink under the model ID glm-5.3 through Chat Completions · Responses · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window and up to 131K output tokens, plus long-context coding, multi-step agents, prompt caching.
GLM-5.3 specs and capabilities
Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.
Context window
1M tokens
Max output
131K tokens
Input
Text
Output
Text · JSON (structured output) · tool calls
Reasoning
Always-on reasoning
Tool use
Function calling with multi-step tool sequences
Prompt caching
Automatic cache reads at a lower rate
Server-side tools
Web search, billed per successful call
Protocols
Chat Completions · Responses · Anthropic Messages
Model ID
glm-5.3
Why GLM-5.3 can handle these workloads
Three properties do most of the work. Each one also implies a way to use the model badly, so they are worth understanding before you route production traffic.
A 1M-token working context at one flat rate
There is no long-context price tier — the rate at token 900,000 equals the rate at token 1,000. That removes a common reason to truncate aggressively, but it does not make filling the window free or accurate.
Byte-faithful request passthrough
Apart from the model name, the request body reaches the upstream unchanged, so GLM-specific fields, prompt-cache prefix hashes, and reasoning signatures are preserved on both endpoints.
Reasoning tokens included in output billing
Reasoning is billed inside completion tokens rather than as a separate line, so a 131,072-token output ceiling covers reasoning and final answer together. Set output budgets by task, not by the ceiling.
Where GLM-5.3 earns a place in a production model stack
GLM-5.3 costs more than GLM-5.2 and much more than Flash, so it should not become the default route just because it is newest. Its strongest fit is work where deeper reasoning removes a retry or a round of human correction.
GLM-5.3 for repository-scale coding
Hold connected modules, tests, and conventions in one 1M-token context instead of feeding a coding agent one file at a time. The gain shows up as fewer broken cross-file edits, not as a better single-file completion.
GLM-5.3 for long task chains
Plan-execute-verify loops that run for many steps benefit most, because a single early reasoning error compounds across the whole chain. Measure completed-chain rate rather than per-step quality.
GLM-5.3 for tool-heavy agents
Large tool schemas and long instruction blocks sit in the cached prefix, so the marginal cost of another step stays close to the cache-read rate rather than the full input rate.
GLM-5.3 through Anthropic-protocol clients
Requests to /v1/messages are passed through byte-for-byte, so prompt-cache prefixes and reasoning signatures survive. Claude-protocol coding CLIs can point at EvoLink and select glm-5.3 without a translation layer.
What GLM-5.3 changes — and what it breaks
GLM-5.3 is not a drop-in replacement for GLM-5.2. One parameter change is a hard break, and the pricing moved. Read these four before you switch a production route.
Breaking: reasoning can no longer be disabled
GLM-5.3 always reasons and rejects thinking.type: "disabled". Any client migrated from GLM-5.2 that still sends the disabled mode will fail outright. The official replacement is thinking.type: "enabled" with reasoning_effort: "low".
Three reasoning effort levels, defaulting to max
reasoning_effort accepts low, high, and max, and defaults to max. Because output tokens include reasoning tokens, the level moves cost directly — leaving the default in place on high-volume routine calls is the most common way to overspend on 5.3.
Cache reads are the cheapest lever on the page
Cache-read tokens bill at a quarter of fresh input. In a long agent loop where the system prompt and tool schemas dominate the input, keeping that prefix byte-stable matters more to the bill than the model choice itself.
Priced above GLM-5.2
Input, output, and cache reads all run 12.5% above the current GLM-5.2 rate: 5.3 is sold at 10% off the official price, while 5.2 carries a deeper 20% discount. Route to 5.3 where the extra reasoning is measurably worth it, and leave the rest on 5.2 or Flash.
Two ways to use GLM-5.3: EvoLink API or Agent
Use the EvoLink API for product backends and batch jobs, or call GLM-5.3 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.
Option 1
Integrate with the EvoLink API
Best for: product backends, batch jobs, automated pipelines
Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.
1Create an EvoLink API key in the console
2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
3Send one representative request and read the usage field for input, cached, and output tokens
4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini
Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls GLM-5.3 through EvoLink, and returns the answer with token usage.
1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
2Describe the task, the inputs to include, and the expected output format
3Ask the Agent to call GLM-5.3 through EvoLink and show the request before sending
4Let the Agent report the answer, token usage, and any error body
GLM-5.3 API code example and error handling
This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.
curl -X POST https://api.evolink.ai/v1/chat/completions \
-H "Authorization: Bearer $EVOLINK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3",
"messages": [
{ "role": "system", "content": "You are a senior engineer. Answer in JSON when asked." },
{ "role": "user", "content": "Review this diff and list risky changes as a JSON array:\n<diff>" }
],
"reasoning_effort": "low",
"max_tokens": 2048,
"stream": true
}'
# Reasoning is always on: reasoning_effort accepts low / high / max
# (default max) and thinking.type "disabled" is rejected.
# The same model ID works on /v1/messages (Anthropic Messages).
# The response includes choices[0].message and a usage object (prompt,
# cached, and completion tokens; completion tokens include reasoning).
Invalid request or unsupported parameter
Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.
Authentication or balance issue
Check the Authorization bearer token and confirm the available balance in the console.
Context length exceeded
Prompt tokens exceed the GLM-5.3 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.
Rate limited (429)
Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.
Content or tool call rejected
Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.
What to verify before routing production traffic to GLM-5.3
Four checks catch nearly every migration problem we see on this model. The first two are hard failures; the last two are silent cost problems.
01
Use glm-5.3 as the model ID
The same ID works on Chat Completions and Anthropic Messages. It matches the upstream name exactly, so there is no alias to translate.
Required
02
Remove any thinking.type: "disabled"
This is the one hard break from GLM-5.2. Replace it with thinking.type: "enabled" plus reasoning_effort: "low" and re-run your regression prompts.
Breaking
03
Set reasoning_effort deliberately
The default is max. Drop routine and latency-sensitive calls to low and confirm the accepted-result rate holds before you keep the saving.
Cost
04
Confirm cache reads are actually hitting
Check that cached tokens appear in the returned usage. If the prefix shifts between calls, every request bills at full input rate and the change is invisible in the response itself.
Cost
Cache reads cost a quarter of fresh input
Cached input is billed on its own lower rate, so a stable system prompt, repository instructions, and tool schemas stay cheap across a long agent loop. There is no separate cache-write charge — the upstream reports no cache-creation tokens.
1M-token context with a 131K output ceiling
Keep connected code, specifications, and agent state in one working context. Treat the ceiling as capacity rather than a target: retrieve what matters, keep the prefix stable so it caches, and set task-appropriate output budgets.
Compare GLM-5.3 cost per accepted task, not token price alone
GLM-5.3 is more expensive per token than both GLM-5.2 and Flash. That only matters alongside how often each model produces a result you actually ship. Track these signals on your own workloads for a week before committing a route.
First-pass success rateAccepted deliverablesRetries per taskReasoning share of output tokensCache-hit ratioValid tool callsTime to accepted resultFallback rate
A route that costs 12.5% more per token but removes one retry in three is cheaper in practice. A route that costs 12.5% more and changes nothing is just a more expensive default. The only way to tell them apart is to measure both on the same tasks.
Compare against GLM-5.2 and Kimi K3 before switching
EvoLink
GLM-5.3 is priced above GLM-5.2, so the question is whether fewer retries and less human correction cover the difference on your own workloads. Measure first, then choose the production route.
Repository-scale coding agents, long-running task chains, tool-heavy workflows, and stable-prefix agent loops where cache reads carry most of the input volume.
The previous GLM flagship, still cheaper per token. Keep it for workloads where 5.3 shows no measurable gain, and for clients that rely on disabling reasoning.
Moonshot's 1M-context reasoning route. A useful cross-vendor baseline for long-context coding and agent work.
GLM-5.3
Input / output$1.059 / $3.706
Context1M
CachingCache reads
Best forRepository-scale coding agents, long-running task chains, tool-heavy workflows, and stable-prefix agent loops where cache reads carry most of the input volume.
Best forThe previous GLM flagship, still cheaper per token. Keep it for workloads where 5.3 shows no measurable gain, and for clients that rely on disabling reasoning.
Best forMoonshot's 1M-context reasoning route. A useful cross-vendor baseline for long-context coding and agent work.
GLM Model Family
Same API key and balance — switch tiers without changing your integration.
GLM-5.2
The previous GLM flagship, still cheaper per token. Keep it for workloads where 5.3 shows no measurable gain, and for clients that rely on disabling reasoning.
Yes. GLM-5.3 is available as a production model, served over both Chat Completions and Anthropic Messages.
What model ID should I use for the GLM-5.3 API?
Use glm-5.3 for both Chat Completions and Anthropic Messages. The EvoLink model name matches the upstream name exactly.
Can I keep using the OpenAI SDK or Anthropic Messages?
Yes. Point either client at EvoLink with the same API key and select glm-5.3. Requests are passed through byte-for-byte apart from the model name, so prompt-cache prefixes and reasoning signatures survive.
Can I still disable reasoning like on GLM-5.2?
No. GLM-5.3 always reasons and rejects thinking.type: "disabled". If you are migrating from GLM-5.2, switch to thinking.type: "enabled" with reasoning_effort: "low" — that is the official replacement for the disabled mode.
How should I choose a reasoning_effort level?
The default is max. Use low for routine edits, classification, and latency-sensitive calls; high or max for architecture, debugging, and long task chains. Output tokens include reasoning tokens, so effort directly moves cost.
Why is GLM-5.3 more expensive than GLM-5.2?
GLM-5.3 is sold at 10% off the official price, while GLM-5.2 carries a deeper 20% discount, so input, output, and cache reads each run 12.5% higher. Route to 5.3 where the extra reasoning removes retries, and keep 5.2 or GLM-5.3 Flash where it does not.
How is prompt caching billed?
Cache reads have their own rate, a quarter of fresh input. There is no cache-write charge because the upstream does not report cache-creation tokens. Keep the prompt prefix byte-stable so it keeps hitting.
What is the context window and output limit?
1,000,000 tokens of context and up to 131,072 output tokens, at one flat rate across the whole window — there is no long-context price tier.
Should I use GLM-5.3 or GLM-5.3 Flash?
Flash costs a tenth as much on input and output, and accepts image, video, and file input. Start on Flash for high-volume, multimodal, and routine work; move to GLM-5.3 when a task needs deeper reasoning and the accepted-result rate justifies it.
Does GLM-5.3 accept images?
No. GLM-5.3 is text-only. Use GLM-5.3 Flash for image, video, and file input.
How should I compare GLM-5.3 with Claude, GPT, or Kimi?
Evaluate GLM-5.3 on cost per accepted task for coding agents and long task chains, treat Claude and GPT as frontier capability baselines, and use Kimi K3 as the cross-vendor long-context comparison.
What should a production GLM-5.3 evaluation measure?
Track first-pass success, accepted deliverables, retries, reasoning-token share of output, cache-hit ratio, valid tool calls, time to accepted result, and fallback rate.