Production routeThis rate reflects platform-side availability — only confirmed server errors (HTTP 500 / empty response) count as failures. User-side issues (content moderation, invalid params, cancellation) plus rate limits, timeouts and auth errors are excluded. Before real traffic arrives, empty buckets may display as available.Live
Live
Context
1M context · 131K max output
Best For
High-volume classification, screenshot and document understanding, routine agent steps
The natively multimodal member of the 5.3 generation: text, image, video, and file input at about a tenth of the GLM-5.3 price, with the same 1M-token context, always-on reasoning, tool calling, and prompt caching.
GLM-5.3 Flash
Z.ai multimodal high-volume model
Selected
Model ID
glm-5.3-flash
Best for
High-volume classification and extraction, screenshot and document understanding, video and file analysis, and routine agent steps where per-token price decides the routing.
Input
$0.106 / 1M-10%
7.2 cr / 1M$0.118official price
Cache read
$0.031 / 1M-9%
2.1 cr / 1M$0.034official price
Output
$0.371 / 1M-10%
25.2 cr / 1M$0.412official price
All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing.
GLM-5.3 Flash pricing
Estimate what one GLM-5.3 Flash request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.
Request calculator
Enter the token mix for one request and the number of successful tool calls.
Only successful server-side calls are billed per call; failed attempts have no tool fee, but tokens still apply.
Web search$0.010/ call·0.68 cr / call
GLM-5.3 Flash API for multimodal and high-volume work
The volume tier of the 5.3 generation on EvoLink's unified API. Text, image, video, and file input at about a tenth of the GLM-5.3 price, with the same 1M-token context, always-on reasoning, tool calling, and prompt caching.
GLM-5.3 Flash is served on EvoLink under the model ID glm-5.3-flash through Chat Completions · Responses · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window and up to 131K output tokens, plus image input, long-document analysis, multi-step agents.
GLM-5.3 Flash specs and capabilities
Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.
Context window
1M tokens
Max output
131K tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Always-on reasoning
Tool use
Function calling with multi-step tool sequences
Prompt caching
Automatic cache reads at a lower rate
Server-side tools
Web search, billed per successful call
Protocols
Chat Completions · Responses · Anthropic Messages
Model ID
glm-5.3-flash
Why GLM-5.3 Flash can carry this much volume
Flash keeps the platform properties of the 5.3 generation and changes only the price and the input modalities. That combination is what makes it a default rather than a fallback.
Mixed text and media in one request
One message can carry instructions, several images, and a file together, so evidence that belongs to one decision stays in one call instead of being stitched across a pipeline.
The same 1M context at one flat rate
Flash is not a short-context model. The full 1M window is available at a single rate with no long-context tier, which is what makes bulk document work practical at this price.
Cache reads at a fraction of input
A stable system prompt and tool schemas bill at the cache-read rate across a long loop. On an already-cheap model this is what pushes the marginal cost of an extra step close to nothing.
Where GLM-5.3 Flash belongs in a production model stack
Flash is priced for volume, so the question is rarely whether you can afford it — it is whether it clears your accuracy bar. Run it first, escalate only what fails, and the price gap against the flagship starts working for you.
High-volume classification and extraction
Ticket routing, tagging, structured extraction, and content triage run at a price where a full re-run costs less than the engineering time to avoid one. Batch size stops being the constraint.
Screenshot and document understanding
Send UI screenshots, scanned pages, or diagrams as image_url blocks and get structured output back. This is the capability GLM-5.3 does not have at all — the flagship is text-only.
Video and file analysis
Video and file input are accepted natively rather than through a separate pipeline, so one request can carry mixed evidence instead of being split across a transcription step and a reasoning step.
Routine steps inside a larger agent
Summarising a tool result, deciding a branch, reformatting output — the steps that dominate call volume but not difficulty. Keep the flagship for the steps where reasoning depth actually decides the outcome.
What GLM-5.3 Flash gives you, and what it asks in return
Flash is not a cut-down 5.3 with a smaller window — the context, protocols, and caching are identical. Two things differ, and one is still worth measuring yourself.
Native multimodal input, which the flagship lacks
Text, image, video, and file input all work through the standard messages structure. Images go in as image_url blocks with a public URL (officially recommended) or a Base64 data URL, one block per image.
About a tenth of the flagship price
Input and output both sit at exactly one tenth of the GLM-5.3 rate, and cache reads are cheaper again. That gap is wide enough to change the shape of a workload, not just its bill.
Same reasoning constraint as the rest of 5.3
Reasoning is always on here too, and thinking.type: "disabled" is rejected. Z.ai additionally recommends clear_thinking: false for Flash; the parameter is passed through unchanged.
Media token accounting is worth verifying yourself
Images and video count into prompt_tokens and bill at the input rate, but the upstream publishes no per-image token formula. Run a representative sample, read the returned usage, and size your batch from that rather than from an estimate.
Two ways to use GLM-5.3 Flash: EvoLink API or Agent
Use the EvoLink API for product backends and batch jobs, or call GLM-5.3 Flash from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.
Option 1
Integrate with the EvoLink API
Best for: product backends, batch jobs, automated pipelines
Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.
1Create an EvoLink API key in the console
2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
3Send one representative request and read the usage field for input, cached, and output tokens
4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini
Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls GLM-5.3 Flash through EvoLink, and returns the answer with token usage.
1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
2Describe the task, the inputs to include, and the expected output format
3Ask the Agent to call GLM-5.3 Flash through EvoLink and show the request before sending
4Let the Agent report the answer, token usage, and any error body
GLM-5.3 Flash API code example and error handling
This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.
curl -X POST https://api.evolink.ai/v1/chat/completions \
-H "Authorization: Bearer $EVOLINK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-flash",
"messages": [
{ "role": "system", "content": "You extract structured data. Answer in JSON only." },
{ "role": "user", "content": [
{ "type": "text", "text": "List every line item on this receipt as JSON." },
{ "type": "image_url", "image_url": { "url": "https://example.com/receipt.jpg" } }
] }
],
"max_tokens": 1024,
"temperature": 0.1
}'
# Reasoning is always on for the 5.3 generation: do not send
# thinking.type "disabled". The same model ID works on /v1/messages
# (Anthropic Messages). The response includes choices[0].message and
# a usage object; image and video input is counted in prompt_tokens.
Invalid request or unsupported parameter
Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.
Authentication or balance issue
Check the Authorization bearer token and confirm the available balance in the console.
Context length exceeded
Prompt tokens exceed the GLM-5.3 Flash context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.
Rate limited (429)
Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.
Content or tool call rejected
Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.
What to verify before routing volume to GLM-5.3 Flash
Two of these are correctness checks and two are cost checks. The media-token check is the one most teams skip and later regret on a large batch.
01
Use glm-5.3-flash as the model ID
The same ID works on Chat Completions and Anthropic Messages and matches the upstream name exactly.
Required
02
Remove any thinking.type: "disabled"
The whole 5.3 generation rejects it. Use thinking.type: "enabled" with reasoning_effort: "low" for the cheapest, fastest path.
Breaking
03
Measure media tokens on a real sample
Send representative images or video, read prompt_tokens from the response, and derive your own per-item cost before committing to a batch size.
Cost
04
Define the escalation rule to GLM-5.3
Decide in advance what counts as a Flash failure and what it triggers. Without a rule, teams either escalate nothing and ship errors, or escalate everything and lose the price advantage.
Routing
Cache reads cost under a third of fresh input
Cached input is billed on its own lower rate, so a stable system prompt and tool schemas stay cheap across a long agent loop. There is no separate cache-write charge — the upstream reports no cache-creation tokens.
Native image, video, and file input
Send images as image_url blocks inside messages[].content[], using a public URL (recommended) or a Base64 data URL, and add one block per image. Media is counted into prompt_tokens and billed at the input rate — measure a representative sample before you size a large batch.
Measure the GLM-5.3 Flash accuracy gap, then price it
The useful comparison is not Flash against a benchmark, it is Flash against GLM-5.3 on your own tasks. If Flash matches the flagship on nine requests in ten, running both and escalating the tenth is far cheaper than running the flagship on all ten.
First-pass success rateAccepted deliverablesEscalation rate to GLM-5.3Media tokens per requestCache-hit ratioReasoning share of output tokensTime to accepted resultCost per accepted task
Price the accuracy gap instead of assuming it. A model at one tenth the cost only has to be right most of the time for a run-then-escalate setup to beat routing everything to the flagship — but "most of the time" is a number you have to measure on your own workload, not inherit from a benchmark.
Compare against GLM-5.3 and DeepSeek V4 Flash
EvoLink
Flash is the volume tier of the 5.3 generation. Check whether it clears your accuracy bar first; if it does, the price gap against the flagship is large enough to change how you route everything routine.
High-volume classification and extraction, screenshot and document understanding, video and file analysis, and routine agent steps where per-token price decides the routing.
The 5.3 flagship: text-only, with Flash at a tenth of its price, and the right escalation target when Flash misses the accuracy bar on hard reasoning.
DeepSeek's high-volume route with a 1M context and very low cache-read pricing. The closest cost comparison for bulk text work.
GLM-5.3 Flash
Input / output$0.106 / $0.371
Context1M
CachingCache reads
Best forHigh-volume classification and extraction, screenshot and document understanding, video and file analysis, and routine agent steps where per-token price decides the routing.
Best forThe 5.3 flagship: text-only, with Flash at a tenth of its price, and the right escalation target when Flash misses the accuracy bar on hard reasoning.
Is the GLM-5.3 Flash API available through EvoLink?
Yes. GLM-5.3 Flash is available as a production model, served over both Chat Completions and Anthropic Messages.
What model ID should I use?
Use glm-5.3-flash for both Chat Completions and Anthropic Messages. The EvoLink model name matches the upstream name exactly.
What input types does Flash accept?
Text, images, video, and files. This is the main difference from GLM-5.3, which is text-only.
How do I send an image?
Put an image_url block in messages[].content[] and pass a public URL (the officially recommended form) or a Base64 data URL. For multiple images, add one image_url block per image.
How is image and video input billed?
Media is counted into prompt_tokens and charged at the input rate — there is no separate media SKU. The upstream does not publish a token conversion formula for images, so run a representative sample and read the returned usage before sizing a large batch.
Can I disable reasoning?
No. The whole 5.3 generation always reasons and rejects thinking.type: "disabled". Use thinking.type: "enabled" with reasoning_effort: "low" when you want the cheapest, fastest path.
What is clear_thinking and should I set it?
It controls whether prior reasoning is cleared between turns. Z.ai recommends clear_thinking: false for Flash. Pass it through as-is; EvoLink does not rewrite it.
How much cheaper is Flash than GLM-5.3?
One tenth of the price on input and output. That gap is usually large enough to justify running Flash first and escalating only the requests that fail your acceptance check.
How is prompt caching billed?
Cache reads have their own rate, under a third of fresh input. There is no cache-write charge because the upstream does not report cache-creation tokens. Keep the prompt prefix byte-stable so it keeps hitting.
What is the context window and output limit?
1,000,000 tokens of context and up to 131,072 output tokens, at one flat rate across the whole window — there is no long-context price tier.
Is Flash a good default for high-volume work?
Yes, that is what it is priced for. Classification, extraction, routine agent steps, and document or screenshot understanding are the natural fit. Escalate to GLM-5.3 only where deeper reasoning measurably changes the outcome.
What should a production evaluation measure?
Track first-pass success, accepted deliverables, retries, reasoning-token share of output, cache-hit ratio, media tokens per request, time to accepted result, and the escalation rate to GLM-5.3.