Skip to main content
POST
GPT Chat Completions (All Models, Full Parameters)
BaseURL: The default BaseURL is https://direct.evolink.ai, which has better support for text models and long-lived connections. https://api.evolink.ai is the primary endpoint for multimodal services and serves as a fallback address for text models.
Server-side tools (web search, code execution, file search, MCP) are only available on the Responses API. The Chat Completions endpoint supports regular function tool calling only.
GPT-6: Function calling with Sol / Luna requires reasoning_effort: "none" in this endpoint. Use the Responses API for function calling with Astra or 6.1 Sol, or for GPT-6 max effort.

Authorizations

Authorization
string
header
required

##All APIs require Bearer Token authentication##

Get API Key:

Visit API Key Management Page to get your API Key

Add to request header:

Body

application/json
model
enum<string>
required

Model to call:

Available options:
gpt-6.1-sol,
gpt-6-astra,
gpt-6-sol,
gpt-6-luna,
gpt-5.6-sol,
gpt-5.6-terra,
gpt-5.6-luna,
gpt-5.5,
gpt-5.4,
gpt-5.2,
gpt-5.1
Example:

"gpt-6.1-sol"

messages
object[]
required

List of chat messages, supporting multi-turn context and multimodal input.

role can be system / developer / user / assistant / tool.

content can be a string or an array of content blocks. Two block types are supported: text (text) and image_url (image):

Image

  • Pass the public URL of the image in image_url.url
  • image_url can also be written directly as a string, equivalent to { "url": "..." }
  • detail controls image analysis fidelity: auto (default) / low / high / original
  • The image must be downloadable, otherwise 400 is returned

Note The block types on this API differ from those on the Responses API (which uses input_text / input_image). They cannot be mixed; using the wrong ones returns 400.

Explicit cache breakpoint for GPT-6 / GPT-5.6, placed on a content block. Up to four writes per request; an implicit breakpoint uses one slot. prompt_cache_breakpoint: {"mode": "explicit"}.

Example:
stream
boolean
default:false

Whether to return the result as a stream (an SSE event stream terminated by data: [DONE]). Defaults to false.

Example:

false

max_completion_tokens
integer

Maximum generated tokens, including reasoning. Prefer max_completion_tokens. GPT-6 accepts legacy max_tokens through conversion; if both are supplied, max_completion_tokens takes precedence and max_tokens is removed. GPT-6 Astra / Sol / Luna and GPT-6.1 Sol support up to 128,000 output tokens.

Example:

2048

reasoning_effort
enum<string>

Reasoning depth control. The allowed values vary by model:

Reasoning tokens are billed as output tokens and counted in usage.completion_tokens_details.reasoning_tokens.

GPT-6 defaults to medium. gpt-6-astra and gpt-6.1-sol do not support none; gpt-6-sol / gpt-6-luna do. GPT-6 max effort is available through Responses only.

Available options:
none,
low,
medium,
high,
xhigh
Example:

"medium"

verbosity
enum<string>

How detailed the answer should be: low / medium / high.

GPT-6 Sol / Luna and GPT-6.1 Sol: support for this parameter has not been confirmed; omit it from basic requests.

Rules for existing models below exclude GPT-6 Sol / Luna and GPT-6.1 Sol:

Note Supported by gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, and gpt-5.5; other models do not support this parameter.

Available options:
low,
medium,
high
Example:

"low"

temperature
number

Sampling temperature, ranging from 0 to 2. Lower values make the output more deterministic.

GPT-6: omit this parameter for gpt-6-astra and gpt-6.1-sol. For gpt-6-sol / gpt-6-luna, adjust it only with reasoning effort set to none; omit it at other effort levels. Omitting effort selects medium, not none.

Existing models: gpt-5.5 / gpt-5.4 / gpt-5.2 / gpt-5.1 allow adjustment; the gpt-5.6 family accepts only the default value 1.

Required range: 0 <= x <= 2
Example:

1

top_p
number

Nucleus sampling parameter, ranging from 0 to 1. Adjusting it together with temperature is not recommended.

GPT-6: omit this parameter for gpt-6-astra and gpt-6.1-sol. For gpt-6-sol / gpt-6-luna, adjust it only with reasoning effort set to none; omit it at other effort levels. Omitting effort selects medium, not none.

Existing models: gpt-5.5 / gpt-5.4 / gpt-5.2 / gpt-5.1 allow adjustment; the gpt-5.6 family accepts only the default value 1.

Required range: 0 <= x <= 1
Example:

1

frequency_penalty
number

Frequency penalty, ranging from -2 to 2. Positive values penalize tokens by how often they have appeared, reducing repeated content.

GPT-6 Sol / Luna and GPT-6.1 Sol: support for this parameter has not been confirmed; omit it from basic requests.

Rules for existing models below exclude GPT-6 Sol / Luna and GPT-6.1 Sol:

Note Adjustable only on gpt-5.4 / gpt-5.2 / gpt-5.1; the gpt-5.6 family and gpt-5.5 do not allow adjustment. gpt-6-astra accepts only the default value 0; passing any other value returns 400.

Required range: -2 <= x <= 2
Example:

0

presence_penalty
number

Presence penalty, ranging from -2 to 2. Positive values encourage the model to talk about new topics.

GPT-6 Sol / Luna and GPT-6.1 Sol: support for this parameter has not been confirmed; omit it from basic requests.

Rules for existing models below exclude GPT-6 Sol / Luna and GPT-6.1 Sol:

Note Adjustable only on gpt-5.4 / gpt-5.2 / gpt-5.1; the gpt-5.6 family and gpt-5.5 do not allow adjustment. gpt-6-astra accepts only the default value 0; passing any other value returns 400.

Required range: -2 <= x <= 2
Example:

0

logprobs
boolean
default:false

Whether to return the log probabilities of each output token.

GPT-6: Astra and 6.1 Sol do not support output logprobs. For Sol / Luna, use them only with reasoning effort none. At other levels, remove logprobs, top_logprobs, and message.output_text.logprobs from Responses include.

Rules for existing models below exclude GPT-6 Sol / Luna and GPT-6.1 Sol:

Note Supported only by gpt-5.4 / gpt-5.2 / gpt-5.1; the gpt-5.6 family and gpt-5.5 do not support this parameter.

GPT-6 Astra and GPT-6.1 Sol do not support this parameter.

Example:

true

top_logprobs
integer

Number of candidate tokens returned at each position, ranging from 0 to 20; must be used together with logprobs: true.

Note Same model support as logprobs.

GPT-6: Astra and 6.1 Sol do not support output logprobs. For Sol / Luna, use them only with reasoning effort none. At other levels, remove logprobs, top_logprobs, and message.output_text.logprobs from Responses include.

Required range: 0 <= x <= 20
Example:

2

n
integer
default:1

Number of candidate replies to generate, returned as multiple entries in the choices array. All tokens (including the output of every candidate) are billed.

GPT-6 Sol / Luna and GPT-6.1 Sol: support for this parameter has not been confirmed; omit it from basic requests.

Example:

1

seed
integer

Random seed. With the same seed and parameter combination, the model tries to return consistent results (best effort; full reproducibility is not guaranteed).

GPT-6 Sol / Luna and GPT-6.1 Sol: support for this parameter has not been confirmed; omit it from basic requests.

Example:

42

response_format
object

Output format control:

  • {"type": "text"}: free-form text, the default
  • {"type": "json_object"}: returns valid JSON, and requires the word json to appear in messages, otherwise 400 is returned
  • {"type": "json_schema", "json_schema": {...}}: returns structured output following the given JSON Schema; combine it with "strict": true to enforce conformance to the schema
tools
object[]

Tool list, used for function calling (client-side function calls, no per-call fee).

Server-side tools (web search, code execution, and so on) are not provided on this API; use the Responses API instead.

GPT-6: gpt-6-sol / gpt-6-luna default to medium reasoning. Function calling in Chat Completions requires an explicit reasoning_effort: "none". gpt-6-astra and gpt-6.1-sol do not support none; use the Responses API for function calling. Use Responses for GPT-6 max effort; this endpoint does not support it.

tool_choice

Tool choice control: "auto" (default) / "none" / "required", or an object naming a specific function, such as {"type": "function", "function": {"name": "get_weather"}}.

Available options:
none,
auto,
required
parallel_tool_calls
boolean
default:true

Whether the model may call multiple tools in parallel within one turn. Defaults to true; set it to false to force calls one at a time.

GPT-6 Sol / Luna and GPT-6.1 Sol: support for this parameter has not been confirmed; omit it from basic requests.

Example:

true

prompt_cache_key
string

Cache grouping key. GPT-6 / GPT-5.6 route cache traffic automatically; this field is not needed to optimize routing. Separate keys can distinguish cache reuse and accounting for customers or users. Keep the key stable for requests that should reuse a prefix. Earlier models can use stable keys to help cache routing.

Example:

"app-chat-v1"

user
string

End-user identifier, used to distinguish the source of calls.

Example:

"user-1024"

prompt_cache_options
object

Prompt caching options for GPT-6 and GPT-5.6. Implicit breakpoints are the default. mode: "explicit" uses only explicit breakpoints; without one, no caching occurs.

Example:

Response

Chat completion succeeded (a JSON object; when stream=true, an SSE event stream terminated by data: [DONE])

id
string

Unique identifier for this conversation

Example:

"chatcmpl-CvJ2p8mQxK7nR4wS"

object
enum<string>

Response type

Available options:
chat.completion
Example:

"chat.completion"

created
integer

Creation timestamp

Example:

1786705221

model
string

Actual model name used

Example:

"gpt-6.1-sol"

choices
object[]

List of generated results (length equals n in the request)

usage
object

Token usage statistics. Prompt caching applies automatically, and cached input tokens are billed at the lower cached rate.

GPT-6 bills uncached input, cache reads, cache writes, and output separately. Above 272,000 input tokens, the entire request uses 2× the regular input and cache rates and 1.5× the output rate. Built-in image generation is billed separately. See current model pricing.