curl --request POST \
--url https://direct.evolink.ai/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "glm-5.3",
"messages": [
{
"role": "user",
"content": "자기소개를 해 주세요"
}
]
}
'{
"id": "chatcmpl-a6613b56-c61c-94ba-9a9f-43d4cdc7d77a",
"object": "chat.completion",
"request_id": "req-7f3a2c1e8b9d4f0a",
"created": 1777021417,
"model": "glm-5.3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "안녕하세요! 저는 GLM-5.3으로, 대화, 추론, 작문, 코딩 등 다양한 작업을 도와드릴 수 있습니다.",
"reasoning_content": "먼저 이 문제를 분석해 보겠습니다……",
"tool_calls": [
{
"id": "<string>",
"type": "function",
"function": {
"name": "<string>",
"arguments": "<string>"
}
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 346,
"total_tokens": 370,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 321
}
},
"web_search": [
{
"icon": "<string>",
"title": "<string>",
"link": "<string>",
"media": "<string>",
"publish_date": "<string>",
"content": "<string>",
"refer": "<string>"
}
],
"content_filter": [
{
"role": "assistant",
"level": 1
}
]
}{
"error": {
"code": 400,
"message": "Invalid request parameters",
"type": "invalid_request_error"
}
}{
"error": {
"code": 401,
"message": "Invalid or expired token",
"type": "authentication_error"
}
}{
"error": {
"code": 402,
"message": "Insufficient quota",
"type": "insufficient_quota_error",
"fallback_suggestion": "https://evolink.ai/dashboard/credits"
}
}{
"error": {
"code": 403,
"message": "Access denied for this model",
"type": "permission_error",
"param": "model"
}
}{
"error": {
"code": 404,
"message": "Specified model not found",
"type": "not_found_error",
"param": "model",
"fallback_suggestion": "glm-5.3"
}
}{
"error": {
"code": 429,
"message": "Rate limit exceeded",
"type": "rate_limit_error",
"fallback_suggestion": "retry after 60 seconds"
}
}{
"error": {
"code": 500,
"message": "Internal server error",
"type": "internal_server_error",
"fallback_suggestion": "try again later"
}
}{
"error": {
"code": 123,
"message": "<string>",
"type": "<string>",
"param": "<string>",
"fallback_suggestion": "<string>"
}
}{
"error": {
"code": 503,
"message": "Service temporarily unavailable",
"type": "service_unavailable_error",
"fallback_suggestion": "retry after 30 seconds"
}
}GLM 전체 모델 인터페이스 - Chat Completions 전체 매개변수
- OpenAI Chat Completions 프로토콜로 GLM 시리즈 모델을 호출하며,
model파라미터로 구체적인 모델을 선택합니다 - 동기 처리 모드로 대화 내용을 실시간으로 반환합니다
- 텍스트 대화: 단일 턴 또는 멀티턴 컨텍스트 대화.
glm-5.3-flash및glm-5.3-flashx는 이미지 입력도 지원합니다 - 시스템 프롬프트:
role=system메시지로 AI의 역할과 동작을 사용자 지정 - 심층 사고:
thinking.type으로 사고 체인을 제어하고reasoning_effort로 추론 강도를 조절하며, 추론 과정은reasoning_content로 반환됩니다 - 스트리밍 출력: SSE 스트리밍 반환 지원(
stream=true) - 도구 호출: Function Calling과 웹 검색 지원(
web_search, 최대 128개 도구) - 구조화 출력:
response_format으로 JSON 모드 활성화
스트리밍 응답 안내: stream=true이면 Server-Sent Events로 반환되며, 각 메시지 형식은 data: {JSON}이고 종료 시 data: [DONE]이 반환됩니다. 각 데이터 청크(ChatCompletionChunk)에는 id, created, model, choices와 선택적으로 usage, content_filter가 포함됩니다. 그중 choices[].delta는 role / content / reasoning_content / tool_calls를 증분으로 반환하고, choices[].finish_reason은 마지막 청크에서 종료 사유를 제공합니다.
curl --request POST \
--url https://direct.evolink.ai/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "glm-5.3",
"messages": [
{
"role": "user",
"content": "자기소개를 해 주세요"
}
]
}
'{
"id": "chatcmpl-a6613b56-c61c-94ba-9a9f-43d4cdc7d77a",
"object": "chat.completion",
"request_id": "req-7f3a2c1e8b9d4f0a",
"created": 1777021417,
"model": "glm-5.3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "안녕하세요! 저는 GLM-5.3으로, 대화, 추론, 작문, 코딩 등 다양한 작업을 도와드릴 수 있습니다.",
"reasoning_content": "먼저 이 문제를 분석해 보겠습니다……",
"tool_calls": [
{
"id": "<string>",
"type": "function",
"function": {
"name": "<string>",
"arguments": "<string>"
}
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 346,
"total_tokens": 370,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 321
}
},
"web_search": [
{
"icon": "<string>",
"title": "<string>",
"link": "<string>",
"media": "<string>",
"publish_date": "<string>",
"content": "<string>",
"refer": "<string>"
}
],
"content_filter": [
{
"role": "assistant",
"level": 1
}
]
}{
"error": {
"code": 400,
"message": "Invalid request parameters",
"type": "invalid_request_error"
}
}{
"error": {
"code": 401,
"message": "Invalid or expired token",
"type": "authentication_error"
}
}{
"error": {
"code": 402,
"message": "Insufficient quota",
"type": "insufficient_quota_error",
"fallback_suggestion": "https://evolink.ai/dashboard/credits"
}
}{
"error": {
"code": 403,
"message": "Access denied for this model",
"type": "permission_error",
"param": "model"
}
}{
"error": {
"code": 404,
"message": "Specified model not found",
"type": "not_found_error",
"param": "model",
"fallback_suggestion": "glm-5.3"
}
}{
"error": {
"code": 429,
"message": "Rate limit exceeded",
"type": "rate_limit_error",
"fallback_suggestion": "retry after 60 seconds"
}
}{
"error": {
"code": 500,
"message": "Internal server error",
"type": "internal_server_error",
"fallback_suggestion": "try again later"
}
}{
"error": {
"code": 123,
"message": "<string>",
"type": "<string>",
"param": "<string>",
"fallback_suggestion": "<string>"
}
}{
"error": {
"code": 503,
"message": "Service temporarily unavailable",
"type": "service_unavailable_error",
"fallback_suggestion": "retry after 30 seconds"
}
}https://direct.evolink.ai이며, 텍스트 모델과 장시간 연결을 더 잘 지원합니다. https://api.evolink.ai는 멀티모달 서비스의 기본 엔드포인트이자 텍스트 모델의 대체 주소 역할을 합니다.enabled는 사고를 켜며 모든 모델의 기본값입니다. disabled는 glm-5.2에서만 사고를 끕니다.glm-5.3, glm-5.3-flash, glm-5.3-flashx는 항상 사고합니다. disabled를 보내면 기본 사고 모드를 유지하며, 사고를 끄거나 강도를 low로 설정하는 것이 아닙니다. 강도를 낮추려면 reasoning_effort=low를 설정하세요. glm-5.2에서 5.3 시리즈로 전환한 뒤 thinking.type=disabled를 유지해도 사고가 계속되며 사고 토큰은 출력으로 과금됩니다.reasoning_effort: 추론 강도이며 기본값은 max입니다.
glm-5.3 / glm-5.3-flash / glm-5.3-flashx 호환 규칙
| 전달 값 | 실제 사고 수준 |
|---|---|
low / high / max | 해당 수준 유지 |
xhigh | max |
medium | high |
minimal / none | low, 사고 유지 |
minimal과 none은 5.3 시리즈의 사고를 끄지 않습니다. 사고 토큰은 출력으로 과금됩니다. 알 수 없는 값은 호환 변환 없이 원래 값을 유지하므로 표에 나온 값을 사용하세요. 이 규칙은 glm-5.2에는 적용되지 않습니다.
glm-5.2는 max, xhigh, high, medium, low, minimal, none을 지원합니다. xhigh는 max, medium / low는 high와 같으며 minimal / none은 사고를 끌 수 있습니다. thinking.type=disabled로도 끌 수 있습니다. 코딩 같은 복잡한 작업에는 max를 사용할 수 있습니다.
- 문자열: 일반 텍스트, 모든 모델이 지원
- 콘텐츠 블록 배열: 텍스트와 이미지 혼합,
glm-5.3-flash및glm-5.3-flashx만 지원
glm-5.3이나 glm-5.2에 이미지 콘텐츠 블록을 보내면 오류가 발생합니다.인증
##모든 인터페이스는 Bearer Token 인증이 필요합니다##
API Key 발급:
API Key 관리 페이지에 방문하여 API Key를 발급받으세요
요청 헤더에 추가:
Authorization: Bearer YOUR_API_KEY
본문
호출할 모델:
| 모델 ID | 포지셔닝 | 사고 제어 | 이미지 입력 |
|---|---|---|---|
glm-5.3 | 플래그십 모델, 복잡한 소프트웨어 엔지니어링과 에이전트 작업 능력이 전반적으로 향상되고 코딩 능력은 이전 세대 대비 큰 폭으로 강화; 1M 컨텍스트 | 항상 활성화됩니다. 실제 수준: low / high / max. 호환 값은 reasoning_effort를 참고하세요. | 미지원 |
glm-5.3-flash | 경량 멀티모달 모델, 희소 어텐션과 선형 어텐션 혼합 구조로 비용이 매우 낮고 비전을 기본 지원; 1M 컨텍스트 | glm-5.3과 동일 | 지원, messages 필드 참조 |
glm-5.3-flashx | 이미지 입력을 지원하는 멀티모달 모델, 1M 컨텍스트 | glm-5.3과 동일 | 지원, messages 필드 참조 |
glm-5.2 | 이전 세대 플래그십, 복잡한 추론과 초장문 컨텍스트; 1M 컨텍스트 | thinking.type: "disabled"로 끌 수 있음. reasoning_effort는 7단계 모두 지원 | 미지원 |
glm-5.3, glm-5.3-flash, glm-5.3-flashx, glm-5.2 "glm-5.3"
대화 메시지 목록으로, 현재 대화의 완전한 컨텍스트 정보를 포함합니다
system, user, assistant, tool 네 가지 역할을 지원합니다. 역할마다 메시지의 필드 구조가 다르므로 해당 역할을 선택하여 확인하세요. 최소 1개의 메시지를 포함해야 하며, 시스템 메시지나 어시스턴트 메시지만 포함할 수는 없습니다.
1- System Message
- User Message
- Assistant Message
- Tool Message
Show child attributes
Show child attributes
스트리밍 출력 모드를 활성화할지 여부
false: 모델이 완전한 응답을 생성한 후 한 번에 반환(기본값), 짧은 텍스트와 일괄 처리에 적합true: Server-Sent Events(SSE)를 통해 청크 단위로 실시간 반환, 채팅과 장문에 적합; 스트리밍 종료 시data: [DONE]을 반환
false
사고 체인(Chain of Thought)을 켤지 여부를 제어합니다
Show child attributes
Show child attributes
추론 강도이며 기본값은 max입니다.
glm-5.3 / glm-5.3-flash / glm-5.3-flashx 호환 규칙
| 전달 값 | 실제 사고 수준 |
|---|---|
low / high / max | 해당 수준 유지 |
xhigh | max |
medium | high |
minimal / none | low, 사고 유지 |
minimal과 none은 5.3 시리즈의 사고를 끄지 않습니다. 사고 토큰은 출력으로 과금됩니다. 알 수 없는 값은 호환 변환 없이 원래 값을 유지하므로 표에 나온 값을 사용하세요. 이 규칙은 glm-5.2에는 적용되지 않습니다.
glm-5.2는 max, xhigh, high, medium, low, minimal, none을 지원합니다. xhigh는 max, medium / low는 high와 같으며 minimal / none은 사고를 끌 수 있습니다. thinking.type=disabled로도 끌 수 있습니다. 코딩 같은 복잡한 작업에는 max를 사용할 수 있습니다.
max, xhigh, high, medium, low, minimal, none "max"
샘플링 전략을 활성화할지 여부
true(기본값):temperature/top_p로 무작위 샘플링을 수행하여 출력이 더 다양해짐false: 항상 확률이 가장 높은 단어를 선택(그리디 디코딩)하여 출력이 더 확정적이며, 이때temperature와top_p는 무시됩니다
일관성과 재현성이 필요한 작업(예: 코드 생성, 번역)에는 false로 설정하는 것을 권장합니다
true
샘플링 온도로, 출력의 무작위성과 창의성을 제어합니다
설명:
- 범위:
[0.0, 1.0], 소수점 둘째 자리까지 - 높은 값(예: 0.8): 더 무작위하고 창의적이며 창작 글쓰기에 적합
- 낮은 값(예: 0.2): 더 안정적이고 확정적이며 사실 기반 질의응답과 코드 생성에 적합
- 기본값:
1.0
권장 사항: temperature와 top_p를 동시에 조정하지 마세요
0 <= x <= 11
핵 샘플링(Nucleus Sampling) 파라미터로, temperature 샘플링의 대체 방법입니다
설명:
- 범위:
[0.01, 1.0], 소수점 둘째 자리까지 - 모델은 누적 확률이
top_p에 도달하는 후보 단어만 고려하며, 예를 들어 0.1은 상위 10% 확률의 단어만 고려함을 의미합니다 - 작은 값은 더 집중되고 일관된 출력을 만들며, 큰 값은 다양성을 높입니다
- 기본값:
0.95
권장 사항: temperature와 top_p를 동시에 조정하지 마세요
0.01 <= x <= 10.95
모델 출력의 최대 토큰 수 제한
설명:
- GLM 시리즈는 최대 131,072 tokens(128K) 출력 길이를 지원하며
1024이상 설정을 권장합니다 thinking이 켜져 있으면 사고 체인 토큰도 이 상한에 포함됩니다length사유로 생성이 잘렸다면 이 값을 높여 보세요
1 <= x <= 1310721024
모델이 호출할 수 있는 도구 목록
설명:
- 함수 호출(
function)과 웹 검색(web_search)을 지원 - 함수는 최대 128개
- 이 중
web_search는 실제로 검색이 발생하면 호출당 별도로 과금되며, 나머지 도구는 추가 비용이 없습니다
128- Function 도구
- Web Search 도구(웹 검색)
Show child attributes
Show child attributes
모델이 어떤 함수를 호출할지 선택하는 방식을 제어합니다
설명: 도구 유형이 function일 때만 유효하며, 기본값이자 auto만 지원합니다(모델이 도구 호출 여부를 자동으로 결정)
auto "auto"
중지 단어 목록
설명:
- 모델이 생성하는 텍스트가 지정한 문자열을 만나면 즉시 생성을 중지합니다(중지 단어 자체는 반환 텍스트에 포함되지 않음)
- 현재는 단일 중지 단어만 지원하며, 형식은
["stop_word1"], 예:["Human:"]
4["Human:"]
모델 응답 출력 형식을 지정하며, 기본값은 text입니다
설명:
{ "type": "json_object" }는 JSON 모드를 활성화하며, 모델이 유효한 JSON 형식 데이터를 반환하여 구조화 데이터 추출 등의 시나리오에 적합합니다- JSON 모드를 사용할 때는
system또는user메시지에서 JSON 출력을 명확히 요구하는 것을 권장합니다
Show child attributes
Show child attributes
요청 고유 식별자
설명:
- 사용자 측에서 전달하며, 길이는 6-64자이고 고유성을 보장하기 위해 UUID 형식을 권장합니다
- 제공하지 않으면 플랫폼이 자동으로 생성합니다
6 - 64"req-7f3a2c1e8b9d4f0a"
최종 사용자의 고유 식별자
설명: 길이는 6-128자이며, 민감한 정보를 포함하지 않는 고유 식별자 사용을 권장합니다. 플랫폼이 남용 행위를 모니터링하고 탐지하는 데 도움이 됩니다
6 - 128"user-abc123456"
응답
채팅 완성이 성공적으로 생성되었습니다
작업 ID
"chatcmpl-a6613b56-c61c-94ba-9a9f-43d4cdc7d77a"
응답 유형
chat.completion "chat.completion"
요청 ID(요청에서 request_id를 제공한 경우 다시 반환)
"req-7f3a2c1e8b9d4f0a"
요청 생성 시각, Unix 타임스탬프(초)
1777021417
모델 이름
"glm-5.3"
모델 응답 목록
Show child attributes
Show child attributes
호출 종료 시 반환되는 Token 사용 통계
Show child attributes
Show child attributes
웹 검색 관련 정보, web_search 도구를 사용하고 검색에 적중했을 때 반환
Show child attributes
Show child attributes
콘텐츠 안전 관련 정보
Show child attributes
Show child attributes