GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

GLM-5.3 Flash API

Access Z.ai GLM-5.3 Flash—also searched as GLM 5.3 Flash / Zhipu GLM-5.3 Flash— through EvoLink's unified chat API. Test image input, long-document analysis, multi-step agents before integrating.

Z.aiText GenerationAvailable
from $0.106 / 1M input tokens$0.118 official price-10%
API docs
Multimodal inputImage · video · file inputReasoning always onChat + Messages
Production routeLive
Context
1M context · 131K max output
Best For
High-volume classification, screenshot and document understanding, routine agent steps
Input
Text + images
Output
Text · JSON (structured output) · tool calls

Choose GLM-5.3 Flash

The natively multimodal member of the 5.3 generation: text, image, video, and file input at about a tenth of the GLM-5.3 price, with the same 1M-token context, always-on reasoning, tool calling, and prompt caching.

GLM-5.3 Flash

Z.ai multimodal high-volume model

Selected
Model ID
glm-5.3-flash
Best for

High-volume classification and extraction, screenshot and document understanding, video and file analysis, and routine agent steps where per-token price decides the routing.

Input
$0.106 / 1M-10%
7.2 cr / 1M$0.118official price
Cache read
$0.031 / 1M-9%
2.1 cr / 1M$0.034official price
Output
$0.371 / 1M-10%
25.2 cr / 1M$0.412official price

All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing.

GLM-5.3 Flash pricing

Estimate what one GLM-5.3 Flash request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.

Request calculator

Enter the token mix for one request and the number of successful tool calls.

Estimated request cost

GLM-5.3 Flash
USD$0.0003
Credits0.0153

Official estimate $0.0003 · save $0.0001 (9%)

Input tokens0.0072 cr
Cache read tokens0.0005 cr
Output tokens0.0076 cr

Minimum charge: 0.01 credits per request.

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 42483 requests

For quick testing

$50
About 222222 requests

For regular development

$100
About 444444 requests

For production evaluation

Server-side tool rates

Only successful server-side calls are billed per call; failed attempts have no tool fee, but tokens still apply.
  • Web search$0.010/ call0.68 cr / call

GLM-5.3 Flash API for multimodal and high-volume work

The volume tier of the 5.3 generation on EvoLink's unified API. Text, image, video, and file input at about a tenth of the GLM-5.3 price, with the same 1M-token context, always-on reasoning, tool calling, and prompt caching.

GLM-5.3 Flash is served on EvoLink under the model ID glm-5.3-flash through Chat Completions · Responses · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window and up to 131K output tokens, plus image input, long-document analysis, multi-step agents.

GLM-5.3 Flash

GLM-5.3 Flash specs and capabilities

Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.

Context window
1M tokens
Max output
131K tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Always-on reasoning
Tool use
Function calling with multi-step tool sequences
Prompt caching
Automatic cache reads at a lower rate
Server-side tools
Web search, billed per successful call
Protocols
Chat Completions · Responses · Anthropic Messages
Model ID
glm-5.3-flash

Why GLM-5.3 Flash can carry this much volume

Flash keeps the platform properties of the 5.3 generation and changes only the price and the input modalities. That combination is what makes it a default rather than a fallback.

Mixed text and media in one request

One message can carry instructions, several images, and a file together, so evidence that belongs to one decision stays in one call instead of being stitched across a pipeline.

The same 1M context at one flat rate

Flash is not a short-context model. The full 1M window is available at a single rate with no long-context tier, which is what makes bulk document work practical at this price.

Cache reads at a fraction of input

A stable system prompt and tool schemas bill at the cache-read rate across a long loop. On an already-cheap model this is what pushes the marginal cost of an extra step close to nothing.

Where GLM-5.3 Flash belongs in a production model stack

Flash is priced for volume, so the question is rarely whether you can afford it — it is whether it clears your accuracy bar. Run it first, escalate only what fails, and the price gap against the flagship starts working for you.

High-volume classification and extraction

Ticket routing, tagging, structured extraction, and content triage run at a price where a full re-run costs less than the engineering time to avoid one. Batch size stops being the constraint.

Screenshot and document understanding

Send UI screenshots, scanned pages, or diagrams as image_url blocks and get structured output back. This is the capability GLM-5.3 does not have at all — the flagship is text-only.

Video and file analysis

Video and file input are accepted natively rather than through a separate pipeline, so one request can carry mixed evidence instead of being split across a transcription step and a reasoning step.

Routine steps inside a larger agent

Summarising a tool result, deciding a branch, reformatting output — the steps that dominate call volume but not difficulty. Keep the flagship for the steps where reasoning depth actually decides the outcome.

What GLM-5.3 Flash gives you, and what it asks in return

Flash is not a cut-down 5.3 with a smaller window — the context, protocols, and caching are identical. Two things differ, and one is still worth measuring yourself.

Native multimodal input, which the flagship lacks

Text, image, video, and file input all work through the standard messages structure. Images go in as image_url blocks with a public URL (officially recommended) or a Base64 data URL, one block per image.

About a tenth of the flagship price

Input and output both sit at exactly one tenth of the GLM-5.3 rate, and cache reads are cheaper again. That gap is wide enough to change the shape of a workload, not just its bill.

Same reasoning constraint as the rest of 5.3

Reasoning is always on here too, and thinking.type: "disabled" is rejected. Z.ai additionally recommends clear_thinking: false for Flash; the parameter is passed through unchanged.

Media token accounting is worth verifying yourself

Images and video count into prompt_tokens and bill at the input rate, but the upstream publishes no per-image token formula. Run a representative sample, read the returned usage, and size your batch from that rather than from an estimate.

Two ways to use GLM-5.3 Flash: EvoLink API or Agent

Use the EvoLink API for product backends and batch jobs, or call GLM-5.3 Flash from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated pipelines

Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.

  1. 1Create an EvoLink API key in the console
  2. 2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
  3. 3Send one representative request and read the usage field for input, cached, and output tokens
  4. 4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Option 2

Call it with an Agent

Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini

Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls GLM-5.3 Flash through EvoLink, and returns the answer with token usage.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the task, the inputs to include, and the expected output format
  3. 3Ask the Agent to call GLM-5.3 Flash through EvoLink and show the request before sending
  4. 4Let the Agent report the answer, token usage, and any error body

GLM-5.3 Flash API code example and error handling

This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/chat/completions \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      { "role": "system", "content": "You extract structured data. Answer in JSON only." },
      { "role": "user", "content": [
        { "type": "text", "text": "List every line item on this receipt as JSON." },
        { "type": "image_url", "image_url": { "url": "https://example.com/receipt.jpg" } }
      ] }
    ],
    "max_tokens": 1024,
    "temperature": 0.1
  }'

# Reasoning is always on for the 5.3 generation: do not send
# thinking.type "disabled". The same model ID works on /v1/messages
# (Anthropic Messages). The response includes choices[0].message and
# a usage object; image and video input is counted in prompt_tokens.

Invalid request or unsupported parameter

Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance in the console.

Context length exceeded

Prompt tokens exceed the GLM-5.3 Flash context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.

Rate limited (429)

Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.

Content or tool call rejected

Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.

What to verify before routing volume to GLM-5.3 Flash

Two of these are correctness checks and two are cost checks. The media-token check is the one most teams skip and later regret on a large batch.

01

Use glm-5.3-flash as the model ID

The same ID works on Chat Completions and Anthropic Messages and matches the upstream name exactly.

Required
02

Remove any thinking.type: "disabled"

The whole 5.3 generation rejects it. Use thinking.type: "enabled" with reasoning_effort: "low" for the cheapest, fastest path.

Breaking
03

Measure media tokens on a real sample

Send representative images or video, read prompt_tokens from the response, and derive your own per-item cost before committing to a batch size.

Cost
04

Define the escalation rule to GLM-5.3

Decide in advance what counts as a Flash failure and what it triggers. Without a rule, teams either escalate nothing and ship errors, or escalate everything and lose the price advantage.

Routing

Cache reads cost under a third of fresh input

Cached input is billed on its own lower rate, so a stable system prompt and tool schemas stay cheap across a long agent loop. There is no separate cache-write charge — the upstream reports no cache-creation tokens.

Native image, video, and file input

Send images as image_url blocks inside messages[].content[], using a public URL (recommended) or a Base64 data URL, and add one block per image. Media is counted into prompt_tokens and billed at the input rate — measure a representative sample before you size a large batch.

Measure the GLM-5.3 Flash accuracy gap, then price it

The useful comparison is not Flash against a benchmark, it is Flash against GLM-5.3 on your own tasks. If Flash matches the flagship on nine requests in ten, running both and escalating the tenth is far cheaper than running the flagship on all ten.

First-pass success rateAccepted deliverablesEscalation rate to GLM-5.3Media tokens per requestCache-hit ratioReasoning share of output tokensTime to accepted resultCost per accepted task

Price the accuracy gap instead of assuming it. A model at one tenth the cost only has to be right most of the time for a run-then-escalate setup to beat routing everything to the flagship — but "most of the time" is a number you have to measure on your own workload, not inherit from a benchmark.

Compare against GLM-5.3 and DeepSeek V4 Flash

EvoLink

Flash is the volume tier of the 5.3 generation. Check whether it clears your accuracy bar first; if it does, the price gap against the flagship is large enough to change how you route everything routine.

GLM-5.3 Flash
Input / output$0.106 / $0.371
Context1M
CachingCache reads
Best forHigh-volume classification and extraction, screenshot and document understanding, video and file analysis, and routine agent steps where per-token price decides the routing.
GLM-5.3
Input / output$1.059 / $3.706
Context1M
CachingCache reads
Best forThe 5.3 flagship: text-only, with Flash at a tenth of its price, and the right escalation target when Flash misses the accuracy bar on hard reasoning.
DeepSeek V4 Flash
Input / output$0.398 / $1.192
Context1M
CachingCache reads
Best forDeepSeek's high-volume route with a 1M context and very low cache-read pricing. The closest cost comparison for bulk text work.

GLM Model Family

Same API key and balance — switch tiers without changing your integration.

GLM-5.3

GLM-5.3

Z.ai flagship reasoning model

View model
GLM-5.2

GLM-5.2

The previous GLM flagship. Still relevant for clients that depend on disabling reasoning, which the 5.3 generation no longer allows.

View model
GLM-5.3 FlashX

GLM-5.3 FlashX

Z.ai's speed-focused multimodal model

View model

Other text models on EvoLink besides GLM-5.3 Flash

DeepSeek V4 Flash

DeepSeek V4 Flash

DeepSeek's high-volume route with a 1M context and very low cache-read pricing. The closest cost comparison for bulk text work.

View model
Gemini 3.7 Flash

Gemini 3.7 Flash

Google's low-cost multimodal route — a useful cross-vendor baseline when image and document understanding drive the choice.

View model
Kimi K3

Kimi K3

Moonshot’s long-context reasoning route for repository-scale coding and multi-document work.

View model
GPT-5.6

GPT-5.6

OpenAI’s tiered frontier family (Sol/Terra/Luna) for capability, latency, and cost-routing flexibility.

View model

GLM-5.3 Flash guides and related reading

GLM-5.3 Flash vs GLM-5.3

GLM-5.3 Flash vs GLM-5.3

Choose the family route by modality, task difficulty, and cost per accepted result — with a Flash-first escalation policy.

Read guide
GLM-5.3 Is Out: What Shipped

GLM-5.3 Is Out: What Shipped

What the August 14 release confirmed about the 5.3 generation — specs, model IDs, and staged access — and what was still open at launch.

Read guide
GLM-5.3 vs GLM-5.2

GLM-5.3 vs GLM-5.2

Same base model, all gains from post-training, plus the breaking change that reasoning can no longer be disabled across the generation.

Read guide
GLM-5.3 vs Claude

GLM-5.3 vs Claude

Benchmarks, API contracts, and access reality compared before you route agents to either family.

Read guide
GLM-5.3 Cybersecurity Benchmarks

GLM-5.3 Cybersecurity Benchmarks

What the CyberGym and ExploitBench numbers claim, by Z.ai's own reporting, and how defenders should read them.

Read guide
One Gateway for 3 Coding CLIs

One Gateway for 3 Coding CLIs

Config paths, environment variables, and a troubleshooting checklist for running coding CLIs through a single endpoint.

Read guide

GLM-5.3 Flash API FAQ

Is the GLM-5.3 Flash API available through EvoLink?

Yes. GLM-5.3 Flash is available as a production model, served over both Chat Completions and Anthropic Messages.

What model ID should I use?

Use glm-5.3-flash for both Chat Completions and Anthropic Messages. The EvoLink model name matches the upstream name exactly.

What input types does Flash accept?

Text, images, video, and files. This is the main difference from GLM-5.3, which is text-only.

How do I send an image?

Put an image_url block in messages[].content[] and pass a public URL (the officially recommended form) or a Base64 data URL. For multiple images, add one image_url block per image.

How is image and video input billed?

Media is counted into prompt_tokens and charged at the input rate — there is no separate media SKU. The upstream does not publish a token conversion formula for images, so run a representative sample and read the returned usage before sizing a large batch.

Can I disable reasoning?

No. The whole 5.3 generation always reasons and rejects thinking.type: "disabled". Use thinking.type: "enabled" with reasoning_effort: "low" when you want the cheapest, fastest path.

What is clear_thinking and should I set it?

It controls whether prior reasoning is cleared between turns. Z.ai recommends clear_thinking: false for Flash. Pass it through as-is; EvoLink does not rewrite it.

How much cheaper is Flash than GLM-5.3?

One tenth of the price on input and output. That gap is usually large enough to justify running Flash first and escalating only the requests that fail your acceptance check.

How is prompt caching billed?

Cache reads have their own rate, under a third of fresh input. There is no cache-write charge because the upstream does not report cache-creation tokens. Keep the prompt prefix byte-stable so it keeps hitting.

What is the context window and output limit?

1,000,000 tokens of context and up to 131,072 output tokens, at one flat rate across the whole window — there is no long-context price tier.

Is Flash a good default for high-volume work?

Yes, that is what it is priced for. Classification, extraction, routine agent steps, and document or screenshot understanding are the natural fit. Escalate to GLM-5.3 only where deeper reasoning measurably changes the outcome.

What should a production evaluation measure?

Track first-pass success, accepted deliverables, retries, reasoning-token share of output, cache-hit ratio, media tokens per request, time to accepted result, and the escalation rate to GLM-5.3.