GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

GPT-6.1 Sol API

Use GPT-6.1 Sol for multi-file coding and tool-driven agents with a 1.05M-token context. Compare token rates and retain GPT model choices on one EvoLink API key. Tool calls use Responses.

OpenAIText GenerationAvailable
from $1.800 / 1M input tokens$2.000 official price-10%
API docs
Chat (no tools)Responses toolsReasoning effortPrompt cachingText + image input
Production routeLive
Context
1.05M context · 128K max output
Best For
Complex coding, computer use, professional work
Input
Text + images
Output
Text · JSON (structured output) · tool calls

Choose GPT-6.1 Sol

One exact model ID on the OpenAI-compatible integration. Short- and long-context rates for all four token roles are listed on the card.

GPT-6.1 Sol

One versioned model ID for coding and agent workloads; Responses handles tool calls, while Chat Completions accepts requests without tools.

Selected
Model ID
gpt-6.1-sol
Best for

Multi-file fixes, document questions and multi-step agents with task-level acceptance checks

≤ 272K
Input
$1.800 / 1M-10%
122.4 cr / 1M$2.000official price
Cache read
$0.092 / 1M-9%
6.2 cr / 1M$0.100official price
Cache write
$2.250 / 1M-10%
153 cr / 1M$2.500official price
Output
$9.000 / 1M-10%
612 cr / 1M$10.000official price
> 272K
Input
$3.600 / 1M-10%
244.8 cr / 1M$4.000official price
Cache read
$0.183 / 1M-9%
12.4 cr / 1M$0.200official price
Cache write
$4.500 / 1M-10%
306 cr / 1M$5.000official price
Output
$13.500 / 1M-10%
918 cr / 1M$15.000official price

All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing. At more than 272K prompt tokens, input and both cache roles use 2x rates while output uses 1.5x rates for the entire request.

GPT-6.1 Sol pricing

Estimate what one GPT-6.1 Sol request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.

Request calculator

Enter the token mix for one request.

Estimated request cost

GPT-6.1 Sol
Short-context rate
USD$0.0046
Credits0.3073

Official estimate $0.0051 · save $0.0006 (10%)

Input tokens0.1224 cr
Cache read tokens0.0013 cr
Cache write tokens0 cr
Output tokens0.1836 cr

Minimum charge: 0.01 credits per request. At more than 272K prompt tokens, input and both cache roles use 2x rates while output uses 1.5x rates for the entire request.

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 2212 requests

For quick testing

$50
About 11064 requests

For regular development

$100
About 22128 requests

For production evaluation

GPT-6.1 Sol API

An updated OpenAI Sol model for code changes, document analysis and multi-step agent work. Supply text or images and receive text or structured results; use application tools through Responses to turn a proposed answer into a checked action.

GPT-6.1 Sol is served on EvoLink under the model ID gpt-6.1-sol through Chat Completions · Responses, with the same API key and balance you use for every other model. It offers a 1.05M context window and up to 128K output tokens, plus coding, multi-step agents, tool use.

GPT-6.1 Sol

GPT-6.1 Sol specs and capabilities

Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.

Context window
1.05M tokens
Max output
128K tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Configurable reasoning effort
Tool use
Function calling with multi-step tool sequences
Prompt caching
Cache write + cache read rates
Long-context tier
At more than 272K prompt tokens, input and both cache roles use 2x rates while output uses 1.5x rates for the entire request.
Protocols
Chat Completions · Responses
Model ID
gpt-6.1-sol

GPT-6.1 Sol use cases for coding agents and documents

Start with repository changes, review-heavy document questions and workflows that reuse context. Define the accepted patch, supported answer or final record before testing. Provider benchmark results help select a trial; they do not establish your production quality or latency.

Ship Multi-File Code Changes

Supply a fixed repository revision, failing tests and the issue. Evaluate the resulting patch against required tests and unrelated changes; record retries and reviewer time. This makes the model useful to assess for bug fixes and refactors across files, rather than judging one plausible code snippet.

Run Tool-Driven Agents on Responses

Bring your application’s function tools to a Responses workflow. Keep permissions, tool-result association and final record checks under application control. Replay failures and cancellations before production writes. Chat Completions is for requests without tools; adding this model does not make an old tool loop compatible automatically.

Handle Documents and Business Workflows

Use supplied document text or images to answer questions about specifications, reports and tables. Request traceable evidence and check extracted fields against the source. The 1.05M-token context supports large input/output budgets, but retaining material does not guarantee accurate conclusions or turn text output into image generation.

Tune Reasoning per Request

Start from medium and compare supported effort levels on the same task set. Raise effort for difficult debugging or design reviews only when accepted results justify the output usage and latency. Neither none nor minimal is supported; max uses Responses. Keep a simpler GPT route selectable for routine work.

Two ways to use GPT-6.1 Sol: EvoLink API or Agent

Use the EvoLink API for product backends and batch jobs, or call GPT-6.1 Sol from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated pipelines

Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.

  1. 1Create an EvoLink API key in the console
  2. 2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
  3. 3Send one representative request and read the usage field for input, cached, and output tokens
  4. 4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Option 2

Call it with an Agent

Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini

Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls GPT-6.1 Sol through EvoLink, and returns the answer with token usage.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the task, the inputs to include, and the expected output format
  3. 3Ask the Agent to call GPT-6.1 Sol through EvoLink and show the request before sending
  4. 4Let the Agent report the answer, token usage, and any error body

GPT-6.1 Sol API code example and error handling

This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/chat/completions \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6.1-sol",
    "messages": [
      { "role": "system", "content": "You are a senior engineer fixing a failing test suite." },
      { "role": "user", "content": "These tests fail after the refactor. Find the root cause and propose a minimal patch:\n<test output>" }
    ],
    "reasoning_effort": "medium",
    "max_tokens": 4096
  }'

# Use the exact ID gpt-6.1-sol (with a dot) on both Chat Completions and Responses.
# reasoning_effort accepts low, medium (default), high, and xhigh; max is available
# on Responses. none is not supported — migrate none to low. Reasoning tokens are
# billed as output. Chat Completions requests cannot include tools: use Responses
# for function calling and built-in tools. Remove temperature, top_p, logprobs,
# and top_logprobs from requests.

Invalid request or unsupported parameter

Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance in the console.

Context length exceeded

Prompt tokens exceed the GPT-6.1 Sol context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.

Rate limited (429)

Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.

Content or tool call rejected

Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.

Before you take GPT-6.1 Sol to production

Four checks that decide cost and stability.

01

Pick the exact GPT-6.1 Sol model ID

Use gpt-6.1-sol, with a dot, on both Chat Completions and Responses. It does not replace gpt-6-sol, which stays available under its own ID.

Required
02

Send tool calls to Responses

Function calling and built-in tools work only on the Responses API; a Chat Completions request that includes tools is rejected at any reasoning effort.

Required
03

Clean up parameters carried over from GPT-6 Sol

Change reasoning effort none to low, and remove temperature, top_p, logprobs, and top_logprobs; GPT-6.1 Sol rejects them.

Required
04

Structure prompts for cache reads and budget for 272K

Put stable instructions first so cache reads replace full-price input; above 272K prompt tokens the whole request uses long-context rates.

Recommended

Four-part prompt caching

Read the four token roles separately: uncached input, cache reads, cache writes and output. The usage and pricing sections determine what is billed; a repeated prompt is not proof of a cache hit. OpenAI’s cached-read rate is half the old Sol rate, but output, retries and review still determine whole-task cost.

Long-context threshold

At more than 272K prompt tokens, input and both cache roles use 2x rates while output uses 1.5x rates for the entire request.

GPT-6.1 Sol pricing and GPT model choices

EvoLink

Compare EvoLink input/output rates and workload fit for 6.1 Sol, Astra and the previous Sol. Use the existing Models and Pricing sections for cache roles and long-context rates.

GPT-6.1 Sol
EvoLink price / 1M (input / output)$1.8 / $9
Context window1.05M
Prompt cachingCache read + cache write
Best forMulti-file fixes, document questions and multi-step agents with task-level acceptance checks
GPT-6 Astra
EvoLink price / 1M (input / output)$9 / $45
Context window1.05M
Prompt cachingCache read + cache write
Best forThe hardest end-to-end reasoning, coding, computer-use, and research work
GPT-6 Sol
EvoLink price / 1M (input / output)$1.8 / $9
Context window1.05M
Prompt cachingCache read + cache write
Best forPrevious Sol release; keeps reasoning effort none and function calling on Chat Completions

Base-context prices reflect your account's current pricing; signed-out visitors see Default rates. Long-context multipliers are listed in the Pricing section.

GPT Model Family

Same API key and balance — switch tiers without changing your integration.

Compare all GPT models
GPT-6 Astra

GPT-6 Astra

OpenAI's most capable model for the hardest end-to-end work

View model
GPT-6 Sol

GPT-6 Sol

GPT-6 model built for complex coding and agentic workflows

View model
GPT-6 Luna

GPT-6 Luna

The most efficient GPT-6 model for focused, high-volume tasks

View model

Other text models on EvoLink besides GPT-6.1 Sol

Claude Opus 5.5

Claude Opus 5.5

Anthropic's newest Opus model for the hardest long-horizon coding and agent work, with a 1M context window and thinking always on.

View model
Claude Sonnet 5.5

Claude Sonnet 5.5

Anthropic's newest Sonnet model for everyday coding, tool-using agents, and high-volume production traffic, with a 1M context window.

View model
Grok 4.7

Grok 4.7

xAI reasoning and tool-use model with a 500K context window and metered server-side tools.

View model
Gemini 3.8 Flash

Gemini 3.8 Flash

Google's most capable Flash model, with gains on software engineering and agentic tasks and a 1M context window.

View model

GPT-6.1 Sol guides and related reading

GPT-6.1 Sol release date and rollout

GPT-6.1 Sol release date and rollout

The September 29 announcement, channel availability and confirmed model facts

Read guide
GPT-6.1 Sol vs GPT-6 Sol: upgrade decision

GPT-6.1 Sol vs GPT-6 Sol: upgrade decision

Official evidence, a complete accepted-task cost example, endpoint migration and rollback gates

Read guide
GPT-6 Astra API guide

GPT-6 Astra API guide

Astra-specific integration reference; use this page’s Sol documentation for 6.1 request settings

Read guide

GPT-6.1 Sol API - FAQ

Which model ID should I use?

Use gpt-6.1-sol, with a dot. The same ID works on Chat Completions and Responses. gpt-6-sol is a separate model and stays available under its own ID.

How much does GPT-6.1 Sol cost on EvoLink?

For prompts up to 272K tokens, the displayed EvoLink rates are $1.800 input and $9.000 output per 1M tokens; cache reads are $0.092 and writes $2.250. Signed-out visitors see Default rates, while signed-in accounts use their own rates. OpenAI Standard input/output list rates are $2.000 / $10.000. Longer prompts use the separate multipliers shown on this page.

How is cached input billed?

Cache reads and cache writes are separate token segments, and neither is billed again as ordinary input. OpenAI's cache read price for GPT-6.1 Sol is 95% below uncached input, half the GPT-6 Sol cache read price.

When do long-context rates apply?

When total prompt tokens exceed 272,000, the whole request uses the long-context multipliers: 2x for input and both cache roles, 1.5x for output.

Which reasoning effort levels are supported?

low, medium (the default), high, and xhigh on both endpoints, plus max on the Responses API (reasoning.effort). none and minimal are not supported; requests that still send none should switch to low. Reasoning tokens are billed as output.

Can I use function calling on Chat Completions?

No. GPT-6.1 Sol on Chat Completions accepts requests without tools only. Send function calling and built-in tools to the Responses API with the same gpt-6.1-sol ID.

Are temperature, top_p, and logprobs supported?

No. Remove temperature, top_p, logprobs, and top_logprobs from GPT-6.1 Sol requests, and drop message.output_text.logprobs from the Responses include list; requests that keep them are rejected.

What are the context window and maximum output?

The shared context window is 1,050,000 tokens, with up to 128,000 output tokens. OpenAI lists April 30, 2026 as the knowledge cutoff. Input can include text and images; output is text, not generated images. Do not treat the total context window as a guaranteed maximum-input allowance.

Should I move from GPT-6 Sol to GPT-6.1 Sol?

OpenAI Standard input, output and cache-write rates match the previous Sol; cache-read pricing is halved. Provider evaluations report coding and workflow gains, not your own accepted-task results. Check the removed none setting, Responses-only tools and unsupported sampling/logprobs parameters, then compare accepted-task cost and latency before moving traffic.

GPT-6.1 Sol or GPT-6 Astra: which one should I use?

Use your hardest tasks to compare accepted results and total cost. GPT-6.1 Sol’s OpenAI Standard input/output rates are one fifth of Astra’s, but that does not imply a fivefold whole-task saving or equal quality on every job. Keep Astra for tasks whose measured quality justifies the cost; retain both model IDs and verify each route’s tool contract.