GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

GPT-5.6 API

Access OpenAI GPT-5.6—also searched as GPT 5.6 / gpt-5.6-sol / terra / luna— through EvoLink's unified chat API. Test reasoning, tool use, structured output before integrating.

OpenAIText GenerationAvailable
From $0.181 / 1M input tokens$0.200 official price-10%
API docs
Chat CompletionsResponses APIReasoning effortPrompt cachingTool use
Production routeLive
Context
1.05M context · 128K max output
Best For
Coding, production agents, complex reasoning
Input
Text
Output
Text · JSON (structured output) · tool calls

Choose GPT-5.6

All three variants use the same OpenAI-compatible integration. Switch capability and cost with one exact model ID.

GPT-5.6 Sol

Flagship tier for maximum reasoning depth and agentic capability.

Selected
Model ID
gpt-5.6-sol
Best for

Hard reasoning, coding agents, and high-value workflows

< 272K
Input
$3.600 / 1M-10%
244.8 cr / 1M$4.000official price
Cache read
$0.361 / 1M-10%
24.5 cr / 1M$0.400official price
Cache write
$4.500 / 1M-10%
306 cr / 1M$5.000official price
Output
$18.000 / 1M-10%
1224 cr / 1M$20.000official price
≥ 272K
Input
$7.200 / 1M-10%
489.6 cr / 1M$8.000official price
Cache read
$0.721 / 1M-10%
49 cr / 1M$0.800official price
Cache write
$9.000 / 1M-10%
612 cr / 1M$10.000official price
Output
$27.000 / 1M-10%
1836 cr / 1M$30.000official price

GPT-5.6 Terra

Balanced tier for dependable production reasoning and coding.

Model ID
gpt-5.6-terra
Best for

Product features, analysis, chat, and everyday agents

< 272K
Input
$1.800 / 1M-10%
122.4 cr / 1M$2.000official price
Cache read
$0.181 / 1M-10%
12.3 cr / 1M$0.200official price
Cache write
$2.250 / 1M-10%
153 cr / 1M$2.500official price
Output
$10.800 / 1M-10%
734.4 cr / 1M$12.000official price
≥ 272K
Input
$3.600 / 1M-10%
244.8 cr / 1M$4.000official price
Cache read
$0.362 / 1M-10%
24.6 cr / 1M$0.400official price
Cache write
$4.500 / 1M-10%
306 cr / 1M$5.000official price
Output
$16.200 / 1M-10%
1101.6 cr / 1M$18.000official price

GPT-5.6 Luna

Fast, cost-efficient tier for high-volume applications.

Model ID
gpt-5.6-luna
Best for

Latency-sensitive chat, extraction, and scaled automation

< 272K
Input
$0.181 / 1M-10%
12.3 cr / 1M$0.200official price
Cache read
$0.020 / 1M-4%
1.3 cr / 1M$0.020official price
Cache write
$0.225 / 1M-10%
15.3 cr / 1M$0.250official price
Output
$1.081 / 1M-10%
73.5 cr / 1M$1.200official price
≥ 272K
Input
$0.362 / 1M-10%
24.6 cr / 1M$0.400official price
Cache read
$0.039 / 1M-4%
2.6 cr / 1M$0.040official price
Cache write
$0.450 / 1M-10%
30.6 cr / 1M$0.500official price
Output
$1.622 / 1M-10%
110.25 cr / 1M$1.800official price

All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing. Above 272K prompt tokens the long-context tier applies: ×2 for input and cache roles, ×1.5 for output.

GPT-5.6 pricing

Estimate what one GPT-5.6 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.

Request calculator

Enter the token mix for one request.

Estimated request cost

GPT-5.6 Sol
Short-context rate
USD$0.0091
Credits0.6169

Official estimate $0.011 · save $0.0011 (10%)

Input tokens0.2448 cr
Cache read tokens0.0049 cr
Cache write tokens0 cr
Output tokens0.3672 cr

Minimum charge: 0.01 credits per request. At 272K prompt tokens, input and cache roles are billed at ×2 and output at ×1.5. Tool calls are not multiplied.

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 1053 requests

For quick testing

$50
About 5511 requests

For regular development

$100
About 11022 requests

For production evaluation

GPT-5.6 API — Sol, Terra & Luna

Use all three GPT-5.6 tiers through one EvoLink API key: Sol for maximum capability, Terra for balanced production workloads, and Luna for fast, cost-efficient scale.

GPT-5.6 is served on EvoLink under the model ID gpt-5.6-sol through Chat Completions · Responses, with the same API key and balance you use for every other model. It offers a 1.05M context window and up to 128K output tokens, plus reasoning, tool use, structured output.

GPT-5.6

GPT-5.6 specs and capabilities

Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.

Context window
1.05M tokens
Max output
128K tokens
Input
Text
Output
Text · JSON (structured output) · tool calls
Tool use
Function calling with multi-step tool sequences
Prompt caching
Cache write + cache read rates
Long-context tier
×2 in / ×1.5 out above 272K prompt tokens
Protocols
Chat Completions · Responses
Model ID
gpt-5.6-sol

What GPT-5.6 is best for

Pick the tier by workload: Sol for depth, Terra for everyday production, Luna for volume.

Agentic coding and long refactors with GPT-5.6 Sol

Multi-file changes, code review, and tool-heavy agent loops where reasoning depth pays for itself.

Production assistants on GPT-5.6 Terra

Customer-facing chat, document Q&A, and structured extraction with a balanced price-to-quality ratio.

High-volume classification and routing with GPT-5.6 Luna

Tagging, triage, and lightweight summarization where per-token cost dominates.

Long-context analysis up to 1M tokens

Load whole repositories or document sets and use prompt caching to keep repeated context cheap.

Two ways to use GPT-5.6: EvoLink API or Agent

Use the EvoLink API for product backends and batch jobs, or call GPT-5.6 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated pipelines

Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.

  1. 1Create an EvoLink API key in the console
  2. 2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
  3. 3Send one representative request and read the usage field for input, cached, and output tokens
  4. 4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Option 2

Call it with an Agent

Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini

Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls GPT-5.6 through EvoLink, and returns the answer with token usage.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the task, the inputs to include, and the expected output format
  3. 3Ask the Agent to call GPT-5.6 through EvoLink and show the request before sending
  4. 4Let the Agent report the answer, token usage, and any error body

GPT-5.6 API code example and error handling

This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/chat/completions \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "messages": [
      { "role": "system", "content": "You are a senior engineer reviewing a pull request." },
      { "role": "user", "content": "Review this diff and list the blocking issues with file and line:\n<diff>" }
    ],
    "max_tokens": 4096,
    "temperature": 0.2
  }'

# Swap the model ID for gpt-5.6-terra or gpt-5.6-luna to change the tier.
# The response includes choices[0].message and a usage object
# (prompt_tokens, completion_tokens, and cached tokens when a prefix is
# reused).

Invalid request or unsupported parameter

Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance in the console.

Context length exceeded

Prompt tokens exceed the GPT-5.6 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.

Rate limited (429)

Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.

Content or tool call rejected

Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.

Before you take GPT-5.6 to production

Four checks that decide cost and stability.

01

Pick the exact GPT-5.6 model ID

Use gpt-5.6-sol, gpt-5.6-terra, or gpt-5.6-luna; there is no generic gpt-5.6 alias.

Required
02

Budget for the 272K long-context threshold

Above 272K prompt tokens the whole request uses long-context rates; keep prompts under the line when you can.

Recommended
03

Structure prompts for cache reads

Put stable instructions and reference material first so cache reads replace full-price input on repeated calls.

Recommended
04

Measure with usage, not list price

Reconcile cost from the usage object (uncached input, cache read, cache write, output) before scaling traffic.

Required

Four-part prompt caching

Uncached input, cache reads, cache writes, and output are settled separately from the final usage object. Cache writes are not charged again as ordinary input.

Long-context threshold

At more than 272K prompt tokens, input and both cache roles use 2x rates while output uses 1.5x rates for the entire request.

GPT-5.6 vs other text models

EvoLink

Compare the selected GPT-5.6 tier with comparable frontier text models using EvoLink prices per 1M tokens.

GPT-5.6 Sol
EvoLink price / 1M (input / output)$3.6 / $18
Context window1.05M
Prompt cachingCache read + cache write
Best forHard reasoning, coding agents, and high-value workflows
Claude Opus 4.8
EvoLink price / 1M (input / output)$4.5 / $22.5
Context window1M
Prompt cachingCache read + cache write
Best forAnthropic flagship for deep reasoning, coding, and long-running agents.
Gemini 3.1 Pro
EvoLink price / 1M (input / output)$1.865 / $11.183
Context window1M
Prompt cachingContext caching
Best forGoogle flagship for multimodal reasoning and large-context workloads.

Base-context prices use the current EvoLink user group. Logged-in users see their group rates; signed-out users see Default rates. Long-context multipliers remain in each model's Pricing section.

GPT Model Family

Same API key and balance — switch tiers without changing your integration.

Compare all GPT models
GPT-6.1 Sol

GPT-6.1 Sol

Updated Sol for multi-file coding and Responses-based agents

View model
GPT-6 Astra

GPT-6 Astra

Compare callable GPT-5.6 routes with GPT-6 Astra's published specs and define cost, evaluation, and rollout gates before upgrading

View model
GPT-6 Sol

GPT-6 Sol

GPT-6 model built for complex coding and agentic workflows

View model

Other text models on EvoLink besides GPT-5.6

Claude Opus 4.8

Claude Opus 4.8

Anthropic flagship for deep reasoning, coding, and long-running agents.

View model
Gemini 3.1 Pro

Gemini 3.1 Pro

Google flagship for multimodal reasoning and large-context workloads.

View model
Grok 4.5

Grok 4.5

xAI reasoning and tool-use route with a 500K context window for research and agent workflows.

View model
Kimi K3

Kimi K3

Moonshot’s long-context reasoning route for repository-scale coding and multi-document work.

View model

GPT-5.6 API - FAQ

Which model IDs are available?

Use gpt-5.6-sol, gpt-5.6-terra, or gpt-5.6-luna. A generic gpt-5.6 alias is not available.

How is cached input billed?

Cache reads and cache writes are separate token segments. Neither segment is billed again as ordinary input.

When do long-context rates apply?

When total prompt tokens exceed 272,000, the whole request uses the long-context multipliers shown in the pricing table.

Can I use the OpenAI SDK with GPT-5.6 on EvoLink?

Yes. Point the SDK at the EvoLink base URL with your EvoLink API key and select one of the three GPT-5.6 model IDs.

Which GPT-5.6 tier should I start with?

Start with Terra for most production work, move to Sol when tasks need deeper reasoning, and use Luna for high-volume, cost-sensitive requests.

Does GPT-5.6 support tool calling and structured output?

Yes. All three tiers support function calling and JSON-structured output through the standard Chat Completions fields.