Model ID and API Format
Z.ai's official announcement uses the request model ID glm-5.3, with always-on thinking and a new reasoning_effort parameter (low/high/max).
GLM-5.3 launched on August 14, 2026 — but its per-token API is rolling out in stages with no published pricing, so the practical question is when you can actually call it. Join the alert to be notified when EvoLink verifies a callable route.
We will email you when EvoLink verifies a callable GLM-5.3 route, the request model ID, live pricing, and supported API behavior.
We will only email you about GLM-5.3 availability. No spam.
Want faster updates? Join the EvoLink Discord
When access opens we will email you a launch alert, and this page will be updated with the verified model ID, live pricing, API compatibility, regions and limits — plus the path to integrate through EvoLink's unified API while keeping a fallback model.
The August 14 release settled some integration fields and left others open. EvoLink updates each one only after official publication and route-level verification.
Z.ai's official announcement uses the request model ID glm-5.3, with always-on thinking and a new reasoning_effort parameter (low/high/max).
Per-token pricing is unpublished — Z.ai's own pricing page still stops at GLM-5.2. The live EvoLink pricing module will be the source of truth after launch.
Official documentation lists a 1M-token context window and up to 128K output tokens for GLM-5.3.
Regions, concurrency, quotas, rate limits, and 429 behavior require a callable production route; none are published yet.
On release day no aggregator could serve GLM-5.3, and the BigModel API page says "coming soon." Open weights are promised about two weeks after launch.
Only real workloads on a real API route can answer the production question. These six user-demand metrics are what EvoLink will verify after access opens; vendor benchmarks are not a substitute.
Measure whether bug fixes, multi-file refactors, and test tasks actually pass instead of judging one response.
Track argument errors, repeated calls, unproductive loops, and how often a person must take over.
Test whether the model keeps constraints, finds evidence, and completes work deep inside large repositories and documents.
GLM-5.3 removes the option to disable thinking. Measure latency and output-token overhead at each reasoning_effort level against your former non-thinking paths.
Evaluate time to first token, end-to-end speed, capacity, error rate, and peak-hour 429 responses.
Include token use, retries, failed runs, and human correction instead of comparing list prices alone.
GLM-5.3 is not yet callable by the token, and its pricing is unpublished — so teams should not delay a project that can ship on GLM-5.2 today.
If delivery matters now, establish a real baseline on the documented and callable GLM-5.2 route.
Switch only when your workloads show a verified improvement on a real route — and note the breaking change: thinking can no longer be disabled.
Keep the current model behind EvoLink's unified API, then add GLM-5.3 after validation without rebuilding another provider integration.
The release is staged: what shipped, what's pending, and whether to switch.
What shipped on August 14, what's staged, and the open-weight timeline — updated as status changes.
What actually changed, the always-on-thinking breaking change, and when to stay on GLM-5.2.
A separate, still-unannounced model — possibly the release after GLM-5.3 — tracked on its own page.
Not by the token. On release day, GLM-5.3 was available only through the GLM Coding Plan subscription and ZCode; Z.ai's pricing page has no glm-5.3 entry and the BigModel API page says "coming soon."
Z.ai's official announcement uses glm-5.3 in its API example. Verify against official API documentation before production use.
Unpublished. Z.ai's pricing page still lists only GLM-5.2 ($1.40/$0.26/$4.40 per million tokens — GLM-5.2's own rates). This page will use EvoLink's live pricing module only after the route and SKUs are approved.
Not yet verified. EvoLink publishes a callable route, model ID, and pricing only after real request, billing, and behavior checks pass. Join the alert to be notified.
Same base model — all gains come from post-training. The confirmed breaking change: thinking can no longer be disabled, and reasoning_effort (low/high/max) becomes the control surface.
No. GLM-5.3 is text-only; Z.ai's multimodal work remains in the separate GLM-V line.
Z.ai promises the weights about two weeks after launch (around August 28, 2026), after a safety evaluation. The license has not been stated.
No — they are separate models. GLM 5.5 remains unannounced; if it comes, it is expected as a later release, and it is tracked on its own availability page.
GLM-5.2 — live, priced, 1M context, callable through EvoLink today. Keep the model ID configurable so switching to a verified GLM-5.3 route is a configuration change.
A notification after EvoLink verifies the GLM-5.3 route: official release status, model ID, pricing, route behavior, and access state.