curl --request POST \
--url https://direct.evolink.ai/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "glm-5.3",
"messages": [
{
"role": "user",
"content": "请介绍一下你自己"
}
]
}
'{
"id": "chatcmpl-a6613b56-c61c-94ba-9a9f-43d4cdc7d77a",
"object": "chat.completion",
"request_id": "req-7f3a2c1e8b9d4f0a",
"created": 1777021417,
"model": "glm-5.3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "你好!我是 GLM-5.3,可以帮你完成对话、推理、写作、代码等多种任务。",
"reasoning_content": "让我先分析这个问题……",
"tool_calls": [
{
"id": "<string>",
"type": "function",
"function": {
"name": "<string>",
"arguments": "<string>"
}
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 346,
"total_tokens": 370,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 321
}
},
"web_search": [
{
"icon": "<string>",
"title": "<string>",
"link": "<string>",
"media": "<string>",
"publish_date": "<string>",
"content": "<string>",
"refer": "<string>"
}
],
"content_filter": [
{
"role": "assistant",
"level": 1
}
]
}{
"error": {
"code": 400,
"message": "Invalid request parameters",
"type": "invalid_request_error"
}
}{
"error": {
"code": 401,
"message": "Invalid or expired token",
"type": "authentication_error"
}
}{
"error": {
"code": 402,
"message": "Insufficient quota",
"type": "insufficient_quota_error",
"fallback_suggestion": "https://evolink.ai/dashboard/credits"
}
}{
"error": {
"code": 403,
"message": "Access denied for this model",
"type": "permission_error",
"param": "model"
}
}{
"error": {
"code": 404,
"message": "Specified model not found",
"type": "not_found_error",
"param": "model",
"fallback_suggestion": "glm-5.3"
}
}{
"error": {
"code": 429,
"message": "Rate limit exceeded",
"type": "rate_limit_error",
"fallback_suggestion": "retry after 60 seconds"
}
}{
"error": {
"code": 500,
"message": "Internal server error",
"type": "internal_server_error",
"fallback_suggestion": "try again later"
}
}{
"error": {
"code": 123,
"message": "<string>",
"type": "<string>",
"param": "<string>",
"fallback_suggestion": "<string>"
}
}{
"error": {
"code": 503,
"message": "Service temporarily unavailable",
"type": "service_unavailable_error",
"fallback_suggestion": "retry after 30 seconds"
}
}GLM 全模型接口 - Chat Completions 完整参数
- 使用 OpenAI Chat Completions 协议调用 GLM 系列模型,通过
model参数选择具体型号 - 同步处理模式,实时返回对话内容
- 文本对话:单轮或多轮上下文对话;
glm-5.3-flash与glm-5.3-flashx额外支持图像输入 - 系统提示词:通过
role=system消息自定义 AI 的角色和行为 - 深度思考:
thinking.type控制思维链,reasoning_effort调节推理强度;推理过程通过reasoning_content返回 - 流式输出:支持 SSE 流式返回(
stream=true) - 工具调用:支持 Function Calling、网络搜索(
web_search,最多 128 个工具) - 结构化输出:通过
response_format启用 JSON 模式
流式响应说明:当 stream=true 时,通过 Server-Sent Events 返回,每条消息格式为 data: {JSON},结束时返回 data: [DONE]。每个数据块(ChatCompletionChunk)包含 id、created、model、choices、可选 usage 与 content_filter;其中 choices[].delta 增量返回 role / content / reasoning_content / tool_calls,choices[].finish_reason 在最后一块给出终止原因。
curl --request POST \
--url https://direct.evolink.ai/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "glm-5.3",
"messages": [
{
"role": "user",
"content": "请介绍一下你自己"
}
]
}
'{
"id": "chatcmpl-a6613b56-c61c-94ba-9a9f-43d4cdc7d77a",
"object": "chat.completion",
"request_id": "req-7f3a2c1e8b9d4f0a",
"created": 1777021417,
"model": "glm-5.3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "你好!我是 GLM-5.3,可以帮你完成对话、推理、写作、代码等多种任务。",
"reasoning_content": "让我先分析这个问题……",
"tool_calls": [
{
"id": "<string>",
"type": "function",
"function": {
"name": "<string>",
"arguments": "<string>"
}
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 346,
"total_tokens": 370,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 321
}
},
"web_search": [
{
"icon": "<string>",
"title": "<string>",
"link": "<string>",
"media": "<string>",
"publish_date": "<string>",
"content": "<string>",
"refer": "<string>"
}
],
"content_filter": [
{
"role": "assistant",
"level": 1
}
]
}{
"error": {
"code": 400,
"message": "Invalid request parameters",
"type": "invalid_request_error"
}
}{
"error": {
"code": 401,
"message": "Invalid or expired token",
"type": "authentication_error"
}
}{
"error": {
"code": 402,
"message": "Insufficient quota",
"type": "insufficient_quota_error",
"fallback_suggestion": "https://evolink.ai/dashboard/credits"
}
}{
"error": {
"code": 403,
"message": "Access denied for this model",
"type": "permission_error",
"param": "model"
}
}{
"error": {
"code": 404,
"message": "Specified model not found",
"type": "not_found_error",
"param": "model",
"fallback_suggestion": "glm-5.3"
}
}{
"error": {
"code": 429,
"message": "Rate limit exceeded",
"type": "rate_limit_error",
"fallback_suggestion": "retry after 60 seconds"
}
}{
"error": {
"code": 500,
"message": "Internal server error",
"type": "internal_server_error",
"fallback_suggestion": "try again later"
}
}{
"error": {
"code": 123,
"message": "<string>",
"type": "<string>",
"param": "<string>",
"fallback_suggestion": "<string>"
}
}{
"error": {
"code": 503,
"message": "Service temporarily unavailable",
"type": "service_unavailable_error",
"fallback_suggestion": "retry after 30 seconds"
}
}https://direct.evolink.ai,对文本模型支持更好,支持长连接;https://api.evolink.ai 是多模态主力地址,对文本模型作为备用地址使用。glm-5.3、glm-5.3-flash 与 glm-5.3-flashx 始终思考、无法关闭,reasoning_effort 的实际思考档位为 low / high / max;兼容 xhigh → max、medium → high、minimal / none → low。传入 minimal / none 仍会思考,思考 token 按输出计费。传入 thinking.type: "disabled" 也不会关闭思考,会继续使用默认思考模式。glm-5.2 不适用这些兼容规则,可通过 thinking.type: "disabled" 关闭思考。详见 thinking 与 reasoning_effort 字段说明。glm-5.3-flash 与 glm-5.3-flashx 支持,通过 messages[].content[] 的 image_url 内容块传入。其余模型传入图像块会报错。授权
##所有接口均需要使用 Bearer Token 进行认证##
获取 API Key:
访问 API Key 管理页面 获取您的 API Key
使用时在请求头中添加:
Authorization: Bearer YOUR_API_KEY
请求体
要调用的模型:
| 模型 ID | 定位 | 思考控制 | 图像输入 |
|---|---|---|---|
glm-5.3 | 旗舰模型,复杂软件工程与 Agent 任务能力全面进阶,编程能力较上代大幅提升;1M 上下文 | 始终思考,不可关闭;实际思考档位为 low / high / max;兼容取值见 reasoning_effort | 不支持 |
glm-5.3-flash | 轻量多模态模型,稀疏注意力与线性注意力混合架构,成本极低且原生支持视觉;1M 上下文 | 同 glm-5.3 | 支持,见 messages 字段 |
glm-5.3-flashx | 多模态模型,支持图像输入;1M 上下文 | 同 glm-5.3 | 支持,见 messages 字段 |
glm-5.2 | 上一代旗舰,复杂推理与超长上下文;1M 上下文 | 可用 thinking.type: "disabled" 关闭;reasoning_effort 支持 7 档 | 不支持 |
glm-5.3, glm-5.3-flash, glm-5.3-flashx, glm-5.2 "glm-5.3"
对话消息列表,包含当前对话的完整上下文信息
支持四种角色:system、user、assistant、tool。不同角色的消息具有不同的字段结构,请选择对应角色查看。至少包含 1 条消息,且不能只包含系统消息或助手消息。
1- System Message
- User Message
- Assistant Message
- Tool Message
Show child attributes
Show child attributes
是否启用流式输出模式
false:模型生成完整响应后一次性返回(默认),适合短文本与批处理true:通过 Server-Sent Events(SSE)实时逐块返回,适合聊天与长文本;流式结束时返回data: [DONE]
false
控制是否开启思维链(Chain of Thought)
Show child attributes
Show child attributes
控制推理强度,默认 max。
glm-5.3 / glm-5.3-flash / glm-5.3-flashx 的兼容规则:
| 传入值 | 实际思考档位 |
|---|---|
low / high / max | 保持对应档位 |
xhigh | max |
medium | high |
minimal / none | low,仍会思考 |
minimal 与 none 不会关闭 5.3 系列的思考,思考 token 仍按输出计费。未识别的值保持原值,不提供兼容映射;请使用表中列出的值。glm-5.2 不适用上述兼容规则。
glm-5.2:支持 max、xhigh、high、medium、low、minimal、none;xhigh 等价于 max,medium / low 等价于 high,minimal / none 可关闭思考。也可通过 thinking.type: "disabled" 关闭思考。
编程等复杂任务可使用 max。
max, xhigh, high, medium, low, minimal, none "max"
是否启用采样策略
true(默认):使用temperature/top_p进行随机采样,输出更多样false:总是选择概率最高的词汇(贪心解码),输出更确定,此时temperature与top_p被忽略
对需要一致性、可重复性的任务(如代码生成、翻译),建议设置为 false
true
采样温度,控制输出的随机性和创造性
说明:
- 取值范围:
[0.0, 1.0],限两位小数 - 较高值(如 0.8):更随机、更有创意,适合创意写作
- 较低值(如 0.2):更稳定、更确定,适合事实问答与代码生成
- 默认值:
1.0
建议:不要同时调整 temperature 和 top_p
0 <= x <= 11
核采样(Nucleus Sampling)参数,是 temperature 采样的替代方法
说明:
- 取值范围:
[0.01, 1.0],限两位小数 - 模型只考虑累积概率达到
top_p的候选词汇,例如 0.1 表示只考虑前 10% 概率的词汇 - 较小值产生更集中、更一致的输出;较大值增加多样性
- 默认值:
0.95
建议:不要同时调整 temperature 和 top_p
0.01 <= x <= 10.95
模型输出的最大 token 数量限制
说明:
- GLM 系列最大支持 131,072 tokens(128K)输出长度,建议设置不小于
1024 - 开启
thinking时,思维链 token 也计入该上限 - 若生成因
length原因被截断,请尝试调高此值
1 <= x <= 1310721024
模型可以调用的工具列表
说明:
- 支持函数调用(
function)与网络搜索(web_search) - 最多支持 128 个函数
- 其中
web_search在实际发生搜索时按次单独计费,其余工具不额外收费
128- Function 工具
- Web Search 工具(网络搜索)
Show child attributes
Show child attributes
控制模型选择调用哪个函数的方式
说明:仅在工具类型为 function 时生效,默认且仅支持 auto(由模型自动决定是否调用工具)
auto "auto"
停止词列表
说明:
- 当模型生成文本遇到指定字符串时立即停止生成(停止词本身不包含在返回文本中)
- 目前仅支持单个停止词,格式为
["stop_word1"],例如["Human:"]
4["Human:"]
指定模型响应输出格式,默认为 text
说明:
{ "type": "json_object" }启用 JSON 模式,模型返回有效的 JSON 格式数据,适用于结构化数据提取等场景- 使用 JSON 模式时,建议在
system或user消息中明确要求输出 JSON
Show child attributes
Show child attributes
请求唯一标识符
说明:
- 由用户端传递,长度 6-64 字符,建议使用 UUID 格式确保唯一性
- 若未提供,平台将自动生成
6 - 64"req-7f3a2c1e8b9d4f0a"
终端用户的唯一标识符
说明:长度 6-128 字符,建议使用不包含敏感信息的唯一标识,可帮助平台监控和检测滥用行为
6 - 128"user-abc123456"
响应
对话生成成功
任务 ID
"chatcmpl-a6613b56-c61c-94ba-9a9f-43d4cdc7d77a"
响应类型
chat.completion "chat.completion"
请求 ID(在请求中提供 request_id 时回传)
"req-7f3a2c1e8b9d4f0a"
请求创建时间,Unix 时间戳(秒)
1777021417
模型名称
"glm-5.3"
模型响应列表
Show child attributes
Show child attributes
调用结束时返回的 Token 使用统计
Show child attributes
Show child attributes
网页搜索相关信息,使用 web_search 工具且命中搜索时返回
Show child attributes
Show child attributes
内容安全相关信息
Show child attributes
Show child attributes