curl --request POST \
--url https://direct.evolink.ai/v1/messages \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "glm-5.3",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Hello, world"
}
]
}
'{
"id": "msg_0842a705-9d0b-4eaa-b12d-09a4106326c5",
"type": "message",
"role": "assistant",
"model": "glm-5.3",
"content": [
{
"type": "thinking",
"thinking": "The user asked to greet them with one word, so answering \"Hi\" will do.",
"signature": ""
},
{
"type": "text",
"text": "Hi."
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 18,
"output_tokens": 101,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0,
"prompt_tokens_details": {
"cached_tokens": 0
}
}
}GLM All-Model API - Messages Reference
- Call GLM series models through the Anthropic Messages protocol, choosing the specific model with the
modelparameter - The request and response structures match the Anthropic API
- System prompt: passed through the top-level
systemfield - Thinking mode: thinking is on by default across the series and is returned in
content[type=thinking]blocks; onlyglm-5.2can turn it off withthinking.type=disabled - Streaming: SSE event stream
- Tool calling: compatible with the Anthropic
tool_use/tool_resultflow - Image input: supported only by
glm-5.3-flashandglm-5.3-flashx; see themessagesfield for details
curl --request POST \
--url https://direct.evolink.ai/v1/messages \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "glm-5.3",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Hello, world"
}
]
}
'{
"id": "msg_0842a705-9d0b-4eaa-b12d-09a4106326c5",
"type": "message",
"role": "assistant",
"model": "glm-5.3",
"content": [
{
"type": "thinking",
"thinking": "The user asked to greet them with one word, so answering \"Hi\" will do.",
"signature": ""
},
{
"type": "text",
"text": "Hi."
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 18,
"output_tokens": 101,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0,
"prompt_tokens_details": {
"cached_tokens": 0
}
}
}https://direct.evolink.ai, which has better support for text models and long-lived connections. https://api.evolink.ai is the primary endpoint for multimodal services and serves as a fallback address for text models.signature may be an empty string.glm-5.2:thinking.type=disabledturns thinking off.glm-5.3/glm-5.3-flash/glm-5.3-flashx: thinking is always enabled. Passingdisabledretains the default thinking mode and its output-token charges.
thinking.budget_tokens and thinking.effort do not set GLM thinking levels. See the top-level reasoning_effort field for compatible values. A setting that disables thinking on glm-5.2 does not disable it after switching to the 5.3 series.reasoning_effort: GLM reasoning-level compatibility field at the request top level, outside thinking. The endpoint accepts this field but does not guarantee that it controls actual thinking effort.
Compatibility rules for glm-5.3 / glm-5.3-flash / glm-5.3-flashx
| Supplied value | Compatible value |
|---|---|
low / high / max | Unchanged |
xhigh | max |
medium | high |
minimal / none | low; thinking remains enabled |
minimal and none do not disable thinking on the 5.3 series. Thinking tokens are billed as output. Unrecognized values are unchanged and have no compatibility mapping; use the listed values. These mappings do not apply to glm-5.2.
To control thinking effort, use reasoning_effort in Chat Completions or reasoning.effort in Responses.
glm-5.3-flash and glm-5.3-flashx support image input. Image blocks sent to glm-5.3 or glm-5.2 may return 200 even though the model cannot read them. Choose an image-capable model for image understanding.Authorizations
##All interfaces require authentication using a Bearer Token##
Get an API Key:
Visit the API Key management page to obtain your API Key
Add it to the request header when using:
Authorization: Bearer YOUR_API_KEY
Note: EvoLink uses Bearer Token authentication uniformly for /v1/messages.
Body
Model to call:
| Model ID | Positioning | Thinking can be turned off | Image input |
|---|---|---|---|
glm-5.3 | Flagship model, with across-the-board gains on complex software engineering and agent tasks; 1M context | Cannot be turned off | Not supported |
glm-5.3-flash | Lightweight multimodal model, extremely low cost with native vision; 1M context | Cannot be turned off | Supported |
glm-5.3-flashx | Multimodal model with image input; 1M context | Cannot be turned off | Supported |
glm-5.2 | Previous-generation flagship, complex reasoning and very long context; 1M context | Can be turned off (thinking.type=disabled) | Not supported |
glm-5.3, glm-5.3-flash, glm-5.3-flashx, glm-5.2 "glm-5.3"
Conversation messages, alternating between user and assistant. Include at least one message; the last message is usually from the user. Include earlier messages for multi-turn context.
Image input: only glm-5.3-flash and glm-5.3-flashx support image blocks in the content array. Use text only with glm-5.3 and glm-5.2. An HTTP 200 response does not mean that an unsupported model read the image.
1Show child attributes
Show child attributes
Upper limit on the length of the generated content (in tokens)
Note:
- The GLM series supports up to 131,072 tokens (128K) of output; setting at least
1024is recommended - Tokens produced by thinking also count towards this limit
- When the limit is reached the content is truncated and the response carries
stop_reason=max_tokens
1 <= x <= 1310721024
System prompt, used to set the AI's role and behavior
Notes:
- Supports a string or an array of content blocks
- Passed via the top-level
systemfield (do not place it inside messages) - The model follows the system constraints
- An overly long system may be truncated: For long context, place it in
messagesrather than piling everything intosystem
"You are a helpful assistant."
Sampling temperature
Notes:
- Higher values make output more varied, lower values more deterministic
- Recommended range
[0, 1]
0 <= x <= 11
Nucleus sampling threshold
Notes:
- Range
[0, 1] - It is recommended not to adjust temperature and top_p at the same time
0 <= x <= 10.9
Sample only from the K highest-probability tokens (an Anthropic-specific parameter)
Notes:
- Smaller values make output more deterministic, larger values make candidates more diverse
x >= 010
Custom stop sequences: generation stops when it hits any of these strings
Notes:
- Hitting one truncates output, and content before the hit is returned normally
- Note: When a stop sequence is hit, the GLM series returns
stop_reasonasend_turn(rather than the Anthropic-standardstop_sequence), and the response does not include astop_sequencefield. If your client relies onstop_reason=="stop_sequence"to detect a hit, special handling is required
["\n\n"]
Whether to return via SSE streaming
true: Server-Sent Events streaming (standard Anthropic event sequence: message_start / content_block_start / content_block_delta / message_delta / message_stop)false: Returns the complete response all at once (default)
false
Thinking settings. Thinking is enabled by default for all models. Response thinking blocks are billed as output tokens; signature may be an empty string.
- glm-5.2: thinking.type=disabled turns thinking off.
- glm-5.3 / glm-5.3-flash / glm-5.3-flashx: thinking is always enabled. Passing disabled retains the default thinking mode and its output-token charges.
thinking.budget_tokens and thinking.effort do not set GLM thinking levels. See the top-level reasoning_effort field for compatible values. A setting that disables thinking on glm-5.2 does not disable it after switching to the 5.3 series.
Show child attributes
Show child attributes
GLM reasoning-level compatibility field at the request top level, outside thinking. The endpoint accepts this field but does not guarantee that it controls actual thinking effort.
Compatibility rules for glm-5.3 / glm-5.3-flash / glm-5.3-flashx
| Supplied value | Compatible value |
|---|---|
low / high / max | Unchanged |
xhigh | max |
medium | high |
minimal / none | low; thinking remains enabled |
minimal and none do not disable thinking on the 5.3 series. Thinking tokens are billed as output. Unrecognized values are unchanged and have no compatibility mapping; use the listed values. These mappings do not apply to glm-5.2.
To control thinking effort, use reasoning_effort in Chat Completions or reasoning.effort in Responses.
max, xhigh, high, medium, low, minimal, none The list of tool definitions
Notes:
- Follows the Anthropic tool definition spec
input_schemauses a JSON Schema object- The model returns standard
tool_useblocks withstop_reason=tool_use
Show child attributes
Show child attributes
Tool selection strategy
Show child attributes
Show child attributes
Request metadata
Show child attributes
Show child attributes
Response
Message object
Anthropic-style message response
The message's unique ID (format: msg_<uuid>)
Response object type
message assistant Model actually used
"glm-5.3"
The list of response content blocks
Possible block types:
thinking: the reasoning process (when thinking is enabled, which is the default)text: the final answer texttool_use: a tool call initiated by the model
Show child attributes
Show child attributes
Stop reason
end_turn: natural completion (also returned when stop_sequences is hit)max_tokens: reached the max_tokens limittool_use: the model triggered a tool call
end_turn, max_tokens, tool_use Token usage statistics (Anthropic specification)
Show child attributes
Show child attributes