Skip to main content
POST
BaseURL: The default BaseURL is https://direct.evolink.ai, which has better support for text models and long-lived connections. https://api.evolink.ai is the primary endpoint for multimodal services and serves as a fallback address for text models.
Thinking settings. Thinking is enabled by default for all models. Response thinking blocks are billed as output tokens; signature may be an empty string.
  • glm-5.2: thinking.type=disabled turns thinking off.
  • glm-5.3 / glm-5.3-flash / glm-5.3-flashx: thinking is always enabled. Passing disabled retains the default thinking mode and its output-token charges.
thinking.budget_tokens and thinking.effort do not set GLM thinking levels. See the top-level reasoning_effort field for compatible values. A setting that disables thinking on glm-5.2 does not disable it after switching to the 5.3 series.
reasoning_effort: GLM reasoning-level compatibility field at the request top level, outside thinking. The endpoint accepts this field but does not guarantee that it controls actual thinking effort. Compatibility rules for glm-5.3 / glm-5.3-flash / glm-5.3-flashx minimal and none do not disable thinking on the 5.3 series. Thinking tokens are billed as output. Unrecognized values are unchanged and have no compatibility mapping; use the listed values. These mappings do not apply to glm-5.2. To control thinking effort, use reasoning_effort in Chat Completions or reasoning.effort in Responses.
Only glm-5.3-flash and glm-5.3-flashx support image input. Image blocks sent to glm-5.3 or glm-5.2 may return 200 even though the model cannot read them. Choose an image-capable model for image understanding.

Authorizations

Authorization
string
header
required

##All interfaces require authentication using a Bearer Token##

Get an API Key:

Visit the API Key management page to obtain your API Key

Add it to the request header when using:

Note: EvoLink uses Bearer Token authentication uniformly for /v1/messages.

Body

application/json
model
enum<string>
default:glm-5.3
required

Model to call:

Available options:
glm-5.3,
glm-5.3-flash,
glm-5.3-flashx,
glm-5.2
Example:

"glm-5.3"

messages
object[]
required

Conversation messages, alternating between user and assistant. Include at least one message; the last message is usually from the user. Include earlier messages for multi-turn context.

Image input: only glm-5.3-flash and glm-5.3-flashx support image blocks in the content array. Use text only with glm-5.3 and glm-5.2. An HTTP 200 response does not mean that an unsupported model read the image.

Minimum array length: 1
max_tokens
integer

Upper limit on the length of the generated content (in tokens)

Note:

  • The GLM series supports up to 131,072 tokens (128K) of output; setting at least 1024 is recommended
  • Tokens produced by thinking also count towards this limit
  • When the limit is reached the content is truncated and the response carries stop_reason=max_tokens
Required range: 1 <= x <= 131072
Example:

1024

system

System prompt, used to set the AI's role and behavior

Notes:

  • Supports a string or an array of content blocks
  • Passed via the top-level system field (do not place it inside messages)
  • The model follows the system constraints
  • An overly long system may be truncated: For long context, place it in messages rather than piling everything into system
Example:

"You are a helpful assistant."

temperature
number

Sampling temperature

Notes:

  • Higher values make output more varied, lower values more deterministic
  • Recommended range [0, 1]
Required range: 0 <= x <= 1
Example:

1

top_p
number

Nucleus sampling threshold

Notes:

  • Range [0, 1]
  • It is recommended not to adjust temperature and top_p at the same time
Required range: 0 <= x <= 1
Example:

0.9

top_k
integer

Sample only from the K highest-probability tokens (an Anthropic-specific parameter)

Notes:

  • Smaller values make output more deterministic, larger values make candidates more diverse
Required range: x >= 0
Example:

10

stop_sequences
string[]

Custom stop sequences: generation stops when it hits any of these strings

Notes:

  • Hitting one truncates output, and content before the hit is returned normally
  • Note: When a stop sequence is hit, the GLM series returns stop_reason as end_turn (rather than the Anthropic-standard stop_sequence), and the response does not include a stop_sequence field. If your client relies on stop_reason=="stop_sequence" to detect a hit, special handling is required
Example:
stream
boolean
default:false

Whether to return via SSE streaming

  • true: Server-Sent Events streaming (standard Anthropic event sequence: message_start / content_block_start / content_block_delta / message_delta / message_stop)
  • false: Returns the complete response all at once (default)
Example:

false

thinking
object

Thinking settings. Thinking is enabled by default for all models. Response thinking blocks are billed as output tokens; signature may be an empty string.

  • glm-5.2: thinking.type=disabled turns thinking off.
  • glm-5.3 / glm-5.3-flash / glm-5.3-flashx: thinking is always enabled. Passing disabled retains the default thinking mode and its output-token charges.

thinking.budget_tokens and thinking.effort do not set GLM thinking levels. See the top-level reasoning_effort field for compatible values. A setting that disables thinking on glm-5.2 does not disable it after switching to the 5.3 series.

reasoning_effort
enum<string>

GLM reasoning-level compatibility field at the request top level, outside thinking. The endpoint accepts this field but does not guarantee that it controls actual thinking effort.

Compatibility rules for glm-5.3 / glm-5.3-flash / glm-5.3-flashx

minimal and none do not disable thinking on the 5.3 series. Thinking tokens are billed as output. Unrecognized values are unchanged and have no compatibility mapping; use the listed values. These mappings do not apply to glm-5.2.

To control thinking effort, use reasoning_effort in Chat Completions or reasoning.effort in Responses.

Available options:
max,
xhigh,
high,
medium,
low,
minimal,
none
tools
object[]

The list of tool definitions

Notes:

  • Follows the Anthropic tool definition spec
  • input_schema uses a JSON Schema object
  • The model returns standard tool_use blocks with stop_reason=tool_use
tool_choice
object

Tool selection strategy

metadata
object

Request metadata

Response

Message object

Anthropic-style message response

id
string

The message's unique ID (format: msg_<uuid>)

type
enum<string>

Response object type

Available options:
message
role
enum<string>
Available options:
assistant
model
string

Model actually used

Example:

"glm-5.3"

content
object[]

The list of response content blocks

Possible block types:

  • thinking: the reasoning process (when thinking is enabled, which is the default)
  • text: the final answer text
  • tool_use: a tool call initiated by the model
stop_reason
enum<string>

Stop reason

  • end_turn: natural completion (also returned when stop_sequences is hit)
  • max_tokens: reached the max_tokens limit
  • tool_use: the model triggered a tool call
Available options:
end_turn,
max_tokens,
tool_use
usage
object

Token usage statistics (Anthropic specification)