GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

GLM-5.3 API

Access Z.ai GLM-5.3—also searched as Zhipu GLM-5.3 / GLM 5.3— through EvoLink's unified chat API. Test long-context coding, multi-step agents, prompt caching before integrating.

Z.aiText GenerationAvailable
from $1.059 / 1M input tokens$1.177 official price-10%
API docs
Always-on reasoningPrompt cachingTool callingChat + Messages
Production routeLive
Context
1M context · 131K max output
Best For
Coding agents, long task chains, tool-heavy workflows
Input
Text
Output
Text · JSON (structured output) · tool calls

Choose GLM-5.3

Z.ai's flagship reasoning model for coding agents and long-horizon engineering work, with a 1M-token context, always-on reasoning, tool calling, and prompt caching over both Chat Completions and Anthropic Messages.

GLM-5.3

Z.ai flagship reasoning model

Selected
Model ID
glm-5.3
Best for

Repository-scale coding agents, long-running task chains, tool-heavy workflows, and stable-prefix agent loops where cache reads carry most of the input volume.

Input
$1.059 / 1M-10%
72 cr / 1M$1.177official price
Cache read
$0.265 / 1M-10%
18 cr / 1M$0.295official price
Output
$3.706 / 1M-10%
252 cr / 1M$4.118official price

All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing.

GLM-5.3 pricing

Estimate what one GLM-5.3 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.

Request calculator

Enter the token mix for one request and the number of successful tool calls.

Estimated request cost

GLM-5.3
USD$0.0023
Credits0.1512

Official estimate $0.0025 · save $0.0003 (10%)

Input tokens0.072 cr
Cache read tokens0.0036 cr
Output tokens0.0756 cr

Minimum charge: 0.01 credits per request.

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 4497 requests

For quick testing

$50
About 22486 requests

For regular development

$100
About 44973 requests

For production evaluation

Server-side tool rates

Only successful server-side calls are billed per call; failed attempts have no tool fee, but tokens still apply.
  • Web search$0.010/ call0.68 cr / call

GLM-5.3 API for coding agents and long-horizon engineering

Call Z.ai's flagship reasoning model for $1.06 input and $3.71 output per 1M tokens through EvoLink's unified API. Use model ID glm-5.3 with a 1M-token context, always-on low/high/max reasoning, tool calling, and prompt caching.

GLM-5.3 is served on EvoLink under the model ID glm-5.3 through Chat Completions · Responses · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window and up to 131K output tokens, plus long-context coding, multi-step agents, prompt caching.

GLM-5.3

GLM-5.3 specs and capabilities

Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.

Context window
1M tokens
Max output
131K tokens
Input
Text
Output
Text · JSON (structured output) · tool calls
Reasoning
Always-on reasoning
Tool use
Function calling with multi-step tool sequences
Prompt caching
Automatic cache reads at a lower rate
Server-side tools
Web search, billed per successful call
Protocols
Chat Completions · Responses · Anthropic Messages
Model ID
glm-5.3

Why GLM-5.3 can handle these workloads

Three properties do most of the work. Each one also implies a way to use the model badly, so they are worth understanding before you route production traffic.

A 1M-token working context at one flat rate

There is no long-context price tier — the rate at token 900,000 equals the rate at token 1,000. That removes a common reason to truncate aggressively, but it does not make filling the window free or accurate.

Byte-faithful request passthrough

Apart from the model name, the request body reaches the upstream unchanged, so GLM-specific fields, prompt-cache prefix hashes, and reasoning signatures are preserved on both endpoints.

Reasoning tokens included in output billing

Reasoning is billed inside completion tokens rather than as a separate line, so a 131,072-token output ceiling covers reasoning and final answer together. Set output budgets by task, not by the ceiling.

Where GLM-5.3 earns a place in a production model stack

GLM-5.3 costs more than GLM-5.2 and much more than Flash, so it should not become the default route just because it is newest. Its strongest fit is work where deeper reasoning removes a retry or a round of human correction.

GLM-5.3 for repository-scale coding

Hold connected modules, tests, and conventions in one 1M-token context instead of feeding a coding agent one file at a time. The gain shows up as fewer broken cross-file edits, not as a better single-file completion.

GLM-5.3 for long task chains

Plan-execute-verify loops that run for many steps benefit most, because a single early reasoning error compounds across the whole chain. Measure completed-chain rate rather than per-step quality.

GLM-5.3 for tool-heavy agents

Large tool schemas and long instruction blocks sit in the cached prefix, so the marginal cost of another step stays close to the cache-read rate rather than the full input rate.

GLM-5.3 through Anthropic-protocol clients

Requests to /v1/messages are passed through byte-for-byte, so prompt-cache prefixes and reasoning signatures survive. Claude-protocol coding CLIs can point at EvoLink and select glm-5.3 without a translation layer.

What GLM-5.3 changes — and what it breaks

GLM-5.3 is not a drop-in replacement for GLM-5.2. One parameter change is a hard break, and the pricing moved. Read these four before you switch a production route.

Breaking: reasoning can no longer be disabled

GLM-5.3 always reasons and rejects thinking.type: "disabled". Any client migrated from GLM-5.2 that still sends the disabled mode will fail outright. The official replacement is thinking.type: "enabled" with reasoning_effort: "low".

Three reasoning effort levels, defaulting to max

reasoning_effort accepts low, high, and max, and defaults to max. Because output tokens include reasoning tokens, the level moves cost directly — leaving the default in place on high-volume routine calls is the most common way to overspend on 5.3.

Cache reads are the cheapest lever on the page

Cache-read tokens bill at a quarter of fresh input. In a long agent loop where the system prompt and tool schemas dominate the input, keeping that prefix byte-stable matters more to the bill than the model choice itself.

Priced above GLM-5.2

Input, output, and cache reads all run 12.5% above the current GLM-5.2 rate: 5.3 is sold at 10% off the official price, while 5.2 carries a deeper 20% discount. Route to 5.3 where the extra reasoning is measurably worth it, and leave the rest on 5.2 or Flash.

Two ways to use GLM-5.3: EvoLink API or Agent

Use the EvoLink API for product backends and batch jobs, or call GLM-5.3 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated pipelines

Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.

  1. 1Create an EvoLink API key in the console
  2. 2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
  3. 3Send one representative request and read the usage field for input, cached, and output tokens
  4. 4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Option 2

Call it with an Agent

Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini

Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls GLM-5.3 through EvoLink, and returns the answer with token usage.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the task, the inputs to include, and the expected output format
  3. 3Ask the Agent to call GLM-5.3 through EvoLink and show the request before sending
  4. 4Let the Agent report the answer, token usage, and any error body

GLM-5.3 API code example and error handling

This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/chat/completions \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3",
    "messages": [
      { "role": "system", "content": "You are a senior engineer. Answer in JSON when asked." },
      { "role": "user", "content": "Review this diff and list risky changes as a JSON array:\n<diff>" }
    ],
    "reasoning_effort": "low",
    "max_tokens": 2048,
    "stream": true
  }'

# Reasoning is always on: reasoning_effort accepts low / high / max
# (default max) and thinking.type "disabled" is rejected.
# The same model ID works on /v1/messages (Anthropic Messages).
# The response includes choices[0].message and a usage object (prompt,
# cached, and completion tokens; completion tokens include reasoning).

Invalid request or unsupported parameter

Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance in the console.

Context length exceeded

Prompt tokens exceed the GLM-5.3 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.

Rate limited (429)

Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.

Content or tool call rejected

Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.

What to verify before routing production traffic to GLM-5.3

Four checks catch nearly every migration problem we see on this model. The first two are hard failures; the last two are silent cost problems.

01

Use glm-5.3 as the model ID

The same ID works on Chat Completions and Anthropic Messages. It matches the upstream name exactly, so there is no alias to translate.

Required
02

Remove any thinking.type: "disabled"

This is the one hard break from GLM-5.2. Replace it with thinking.type: "enabled" plus reasoning_effort: "low" and re-run your regression prompts.

Breaking
03

Set reasoning_effort deliberately

The default is max. Drop routine and latency-sensitive calls to low and confirm the accepted-result rate holds before you keep the saving.

Cost
04

Confirm cache reads are actually hitting

Check that cached tokens appear in the returned usage. If the prefix shifts between calls, every request bills at full input rate and the change is invisible in the response itself.

Cost

Cache reads cost a quarter of fresh input

Cached input is billed on its own lower rate, so a stable system prompt, repository instructions, and tool schemas stay cheap across a long agent loop. There is no separate cache-write charge — the upstream reports no cache-creation tokens.

1M-token context with a 131K output ceiling

Keep connected code, specifications, and agent state in one working context. Treat the ceiling as capacity rather than a target: retrieve what matters, keep the prefix stable so it caches, and set task-appropriate output budgets.

Compare GLM-5.3 cost per accepted task, not token price alone

GLM-5.3 is more expensive per token than both GLM-5.2 and Flash. That only matters alongside how often each model produces a result you actually ship. Track these signals on your own workloads for a week before committing a route.

First-pass success rateAccepted deliverablesRetries per taskReasoning share of output tokensCache-hit ratioValid tool callsTime to accepted resultFallback rate

A route that costs 12.5% more per token but removes one retry in three is cheaper in practice. A route that costs 12.5% more and changes nothing is just a more expensive default. The only way to tell them apart is to measure both on the same tasks.

Compare against GLM-5.2 and Kimi K3 before switching

EvoLink

GLM-5.3 is priced above GLM-5.2, so the question is whether fewer retries and less human correction cover the difference on your own workloads. Measure first, then choose the production route.

GLM-5.3
Input / output$1.059 / $3.706
Context1M
CachingCache reads
Best forRepository-scale coding agents, long-running task chains, tool-heavy workflows, and stable-prefix agent loops where cache reads carry most of the input volume.
GLM-5.2
Input / output$0.942 / $3.295
Context1M
CachingCache reads
Best forThe previous GLM flagship, still cheaper per token. Keep it for workloads where 5.3 shows no measurable gain, and for clients that rely on disabling reasoning.
Kimi K3
Input / output$2.85 / $14.25
Context1.05M
CachingAutomatic cache reads
Best forMoonshot's 1M-context reasoning route. A useful cross-vendor baseline for long-context coding and agent work.

GLM Model Family

Same API key and balance — switch tiers without changing your integration.

GLM-5.2

GLM-5.2

The previous GLM flagship, still cheaper per token. Keep it for workloads where 5.3 shows no measurable gain, and for clients that rely on disabling reasoning.

View model
GLM-5.3 Flash

GLM-5.3 Flash

Same 5.3 generation at about a tenth of the price, plus native image, video, and file input. The right default for high-volume and multimodal work.

View model
GLM-5.3 FlashX

GLM-5.3 FlashX

Z.ai's speed-focused multimodal model

View model

Other text models on EvoLink besides GLM-5.3

Kimi K3

Kimi K3

Moonshot's 1M-context reasoning route. A useful cross-vendor baseline for long-context coding and agent work.

View model
DeepSeek V4 Pro

DeepSeek V4 Pro

Cost-sensitive long-context baseline for bulk reasoning where per-token price dominates the routing decision.

View model
GPT-5.6

GPT-5.6

OpenAI’s tiered frontier family (Sol/Terra/Luna) for capability, latency, and cost-routing flexibility.

View model
Claude Opus 4.8

Claude Opus 4.8

Anthropic’s premium route for long-running agents, complex reviews, and judgment-heavy work.

View model

GLM-5.3 guides and related reading

GLM-5.3 Flash vs GLM-5.3

GLM-5.3 Flash vs GLM-5.3

Choose the family route by modality, task difficulty, and cost per accepted result — with a Flash-first escalation policy.

Read guide
GLM-5.3 Release: What Changed

GLM-5.3 Release: What Changed

What the August 14 release confirmed, what the live API now costs, and which migration checks matter before production traffic.

Read guide
GLM-5.3 vs GLM-5.2

GLM-5.3 vs GLM-5.2

Same base model, all gains from post-training — plus the breaking change that reasoning can no longer be disabled. When to stay on 5.2.

Read guide
GLM-5.3 vs Claude

GLM-5.3 vs Claude

Benchmarks, API contracts, and access reality compared before you route coding agents to either one.

Read guide
GLM-5.3 Cybersecurity Benchmarks

GLM-5.3 Cybersecurity Benchmarks

What the CyberGym and ExploitBench numbers claim, by Z.ai's own reporting, and how defenders should read them.

Read guide
One Gateway for 3 Coding CLIs

One Gateway for 3 Coding CLIs

Config paths, environment variables, and a troubleshooting checklist for running coding CLIs through a single endpoint.

Read guide

GLM-5.3 API FAQ

Is the GLM-5.3 API available through EvoLink?

Yes. GLM-5.3 is available as a production model, served over both Chat Completions and Anthropic Messages.

What model ID should I use for the GLM-5.3 API?

Use glm-5.3 for both Chat Completions and Anthropic Messages. The EvoLink model name matches the upstream name exactly.

Can I keep using the OpenAI SDK or Anthropic Messages?

Yes. Point either client at EvoLink with the same API key and select glm-5.3. Requests are passed through byte-for-byte apart from the model name, so prompt-cache prefixes and reasoning signatures survive.

Can I still disable reasoning like on GLM-5.2?

No. GLM-5.3 always reasons and rejects thinking.type: "disabled". If you are migrating from GLM-5.2, switch to thinking.type: "enabled" with reasoning_effort: "low" — that is the official replacement for the disabled mode.

How should I choose a reasoning_effort level?

The default is max. Use low for routine edits, classification, and latency-sensitive calls; high or max for architecture, debugging, and long task chains. Output tokens include reasoning tokens, so effort directly moves cost.

Why is GLM-5.3 more expensive than GLM-5.2?

GLM-5.3 is sold at 10% off the official price, while GLM-5.2 carries a deeper 20% discount, so input, output, and cache reads each run 12.5% higher. Route to 5.3 where the extra reasoning removes retries, and keep 5.2 or GLM-5.3 Flash where it does not.

How is prompt caching billed?

Cache reads have their own rate, a quarter of fresh input. There is no cache-write charge because the upstream does not report cache-creation tokens. Keep the prompt prefix byte-stable so it keeps hitting.

What is the context window and output limit?

1,000,000 tokens of context and up to 131,072 output tokens, at one flat rate across the whole window — there is no long-context price tier.

Should I use GLM-5.3 or GLM-5.3 Flash?

Flash costs a tenth as much on input and output, and accepts image, video, and file input. Start on Flash for high-volume, multimodal, and routine work; move to GLM-5.3 when a task needs deeper reasoning and the accepted-result rate justifies it.

Does GLM-5.3 accept images?

No. GLM-5.3 is text-only. Use GLM-5.3 Flash for image, video, and file input.

How should I compare GLM-5.3 with Claude, GPT, or Kimi?

Evaluate GLM-5.3 on cost per accepted task for coding agents and long task chains, treat Claude and GPT as frontier capability baselines, and use Kimi K3 as the cross-vendor long-context comparison.

What should a production GLM-5.3 evaluation measure?

Track first-pass success, accepted deliverables, retries, reasoning-token share of output, cache-hit ratio, valid tool calls, time to accepted result, and fallback rate.