GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5
Editorial illustration of a baseline and candidate separated by a GPT-6.1 Sol vs GPT-6 Sol upgrade evaluation gate
Comparison

GPT-6.1 Sol vs GPT-6 Sol: Same Input/Output Rates, Lower Cache Cost, New Migration Rules

Jacey
Jacey
September 29, 2026
16 min read
Evaluate GPT-6.1 Sol if you want stronger complex-task performance at GPT-6 Sol's ordinary input/output list rates, especially when your agent reuses context. Keep GPT-6 Sol as the baseline until the new tool and reasoning contract passes your tests. OpenAI's September 29 release offers a lower cache-read rate, but it also removes none reasoning and requires Responses for tool calling. An agent that used tools on Chat Completions needs more than a model-ID swap.
This guide is for teams with an existing Sol workload, acceptance tests and a real usage bill. EvoLink's GPT-6 Sol model page gives the current gateway baseline; the GPT collection provides the broader selection path. GPT-6.1 Sol access, feature support and selling rates on EvoLink are not verified here as of September 29, 2026. The comparisons below describe OpenAI's documented contracts and an evaluation method, not an EvoLink head-to-head test.
Compare the configured model, token roles and integration choices on the GPT-6.1 Sol API page. A catalog listing or fallback price is not proof of a successful live request or a settled charge; verify those before routing production traffic.

GPT-6.1 Sol vs GPT-6 Sol at a glance

Decision variableGPT-6 SolGPT-6.1 SolWhat changes for an existing agent
Official model IDgpt-6-solgpt-6.1-solPin the version; do not rely on a generic Sol alias
Input / outputText and image / textText and image / textNo new output modality to design around
Context / max output1,050,000 / 128,000 tokensSame limitsNo larger context window
Knowledge cutoffApril 20, 2026April 30, 2026A newer cutoff is not a quality result on your tasks
Reasoning effortnone, low, medium, high, xhigh, maxlow, medium, high, xhigh, max; no none or minimalRemove unsupported settings from the candidate configuration
Chat Completions function callingOnly at none effortNot supportedMove tool-dependent workflows to a verified Responses route
Responses tool callingSupportedSupportedStill verify your tool loop and gateway behavior
Streaming / structured outputsListed by OpenAIListed by OpenAIProvider support does not prove route-specific support
Sources: the official GPT-6 Sol reference and GPT-6.1 Sol reference, checked September 29. Both default to medium reasoning. The unchanged limits make compatibility and accepted-task cost more useful decision variables than context size.

What the official performance evidence says

OpenAI's launch announcement reports the following improvements over GPT-6 Sol. These are provider evaluations, not EvoLink measurements. A percentage-point change is an absolute score difference, not a relative percentage gain.
EvaluationReported change from GPT-6 SolConditions and decision boundary
DeepSWE v1.1+6.4 percentage points over the old model's best scoreThe candidate used lower effort and cost; this is not an equal-effort comparison
AutomationBench 1.0.6+4.8 percentage pointsSame medium setting; relevant to multi-step tool workflows
OSWorld 2.0+7 percentage points, at less than half the task costMaximum effort; partial reward on the offline v2026.08.08 set

The evidence makes complex code changes and business tool workflows useful trial candidates. OSWorld's partial reward measures progress toward a task; it cannot be substituted for your binary acceptance rate. OpenAI also says its research/API environment can differ from production ChatGPT. Preserve the harness, effort and cost scope when quoting a result.

For routine jobs your current model already passes, the announcement provides less reason for a broad replacement. Use the paired tasks below to establish which gains transfer to your workload. Do not infer your latency, retry rate or EvoLink bill from these results.

Decide the endpoint before judging performance

There are three materially different migrations. Treating them as one trial makes failures hard to interpret.

Existing workflowCandidate pathFirst acceptance check
Chat Completions, no tools, a supported reasoning levelTrial 6.1 Sol's documented no-tools contractResponse parsing, output shape and actual billed usage
Chat Completions tools with noneChange the tool workflow to Responses and select a supported effortTool requests, result association, continuation and retry behavior
Responses with toolsKeep the workflow, pin the new ID and a supported effortFull tool loop, structured results, streaming and cancellation where used
The second row is an integration migration. A failed request there may tell you nothing about the model's ability to solve the task: the old contract is simply unsupported. On the first row, removing none can also alter latency, output usage and behavior even if the endpoint remains unchanged.

For a tool-dependent migration, first inventory the endpoint, effort, tool definitions, result identifiers and continuation handling. Then adapt the tool loop to Responses, including how results are associated with the requesting call. Replay a successful tool action, a failed action and a cancelled run in a sandbox; inspect resulting records as well as the final text. Check structured output parsing and streaming if your client uses them. Only then start the quality comparison.

Keep compatibility failures separate from rejected tasks. Verify the exact gateway route's contract and billing before a pilot, and retain the old model, endpoint and parser together as a rollback configuration.

The cache-read discount is smaller than a whole-job discount

The following table uses OpenAI Standard USD per one million tokens, not EvoLink selling rates. It summarizes the comparison; current gateway rates belong on the corresponding model page.
Token categoryGPT-6 Sol, up to 272K inputGPT-6.1 Sol, up to 272K inputGPT-6.1 Sol, above 272K input
Uncached input$2.00$2.00$4.00
Cached input read$0.20$0.10$0.20
Cache write$2.50$2.50$5.00
Output$10.00$10.00$15.00
Source: OpenAI pricing, checked September 29, 2026. For prompts exceeding 272K input tokens, the higher rates apply to the corresponding token categories across the full request, not just the tokens above the threshold. Batch, Flex, Fast and regional processing have separate conditions; this example uses Standard only.

The direct change from 6 Sol is a 50% lower cache-read rate. Relative to 6.1 Sol's own ordinary input rate, a cached read costs 95% less. Neither statement means that an entire job is 50% or 95% cheaper. Output, uncached input and cache writes have not received that same reduction.

Consider a batch of requests with one million total input tokens and 100,000 output tokens. Each request stays within the short-input band. Assume the usage report confirms the stated cached-input share; exclude cache-write charges, tools, failed attempts and other processing premiums from this simplified illustration.
Cached share of the one million input tokensGPT-6 Sol input + outputGPT-6.1 Sol input + outputDirect saving
40%$1.20 uncached + $0.08 cached + $1.00 output = $2.28$1.20 + $0.04 + $1.00 = $2.24$0.04
90%$0.20 uncached + $0.18 cached + $1.00 output = $1.38$0.20 + $0.09 + $1.00 = $1.29$0.09

Those shares are assumptions, not measured hit rates or a promise that a reused prompt will qualify for caching. If the candidate produces more reasoning output or needs an extra retry, that usage can absorb the illustrated saving. Conversely, a higher accepted-task rate could create value much larger than the cache difference. Measure both effects instead of declaring success from the token-price table.

Build a paired evaluation around six real jobs

Existing tasks flowing through controls, measurement and a guarded rollout while a candidate remains separate, illustrating GPT-6.1 Sol upgrade evaluation
Existing tasks flowing through controls, measurement and a guarded rollout while a candidate remains separate, illustrating GPT-6.1 Sol upgrade evaluation
Editorial illustration of baseline, measurement and rollout stages; no measured model result is encoded in the image.

Freeze recent tasks, repository versions and acceptance rules before running either model. Include examples where the old agent failed or needed human intervention, alongside routine jobs. Keep tool permissions and retry limits consistent; report necessary endpoint or harness changes instead of pretending the experiment is perfectly controlled.

JobInput and expected resultAcceptance signalCost or failure signalTrial priority
Multi-file bug fixFixed repository and issue → patchRequired tests pass without unrelated changesRetries, rework and reviewer interventionStart with failures and review-heavy changes
PR reviewFixed diff and conventions → findingsConfirmed defects with tolerable false positivesTime spent checking false alarmsKeep the baseline if extra findings add review noise
Repeated-context agentSaved context and allowed tools → completed taskConstraints retained; no duplicate side effectsCache reads/writes and reasoning outputTrial when usage confirms substantial cache reads
Document questionsFixed documents and questions → supported answersCorrect fields and traceable evidenceUnsupported conclusions and review timeTest tables and conflicting evidence, not only easy lookup
Business tool workflowGoal and sandbox tools → expected final recordsCorrect sequence and record stateDuplicate writes or mismatched tool resultsVerify the tool loop before judging model quality
Difficult-task escalationFixed task queue → accepted resultsQuality threshold met across the queueCandidate, escalation and fallback costs combinedTrial a separate difficult-task route first

Choose rejection rules before the run. For example, a business workflow may reject any duplicate write even if its final textual answer looks correct. A patch may need hidden tests as well as the tests the agent can see. JSON validity alone cannot establish that extracted fields are accurate.

Track task ID, model identity, endpoint, effort, tool calls, usage, retries, acceptance and reviewer time for every attempt. Summarize sample size and paired disagreements; a small task set can identify a blocker but cannot establish a reliable population-wide improvement. Inspect individual failures as well as averages, particularly when one long-running job dominates the bill.

Compare accepted-task cost, then stage the switch

Use this accounting definition:

Accepted-task cost = total trial cost, including failed attempts, retries, tools and review ÷ accepted tasks.

Keep human review cost explicit: either price it with a documented rate or report review minutes alongside API cost. Do not quietly omit it from one model's result. If no task is accepted, the ratio is undefined; record a failed trial rather than reporting a cheap result.

A complete 100-task accounting example

All four columns below are illustrative assumptions, not observed model results. Each represents the same 100-task queue. Token totals include initial attempts, retries and failed work; every request stays within the Standard short-input band. Treat the four token categories as separately reported billable buckets. Tool charges are assumed totals, and review uses an illustrative internal rate of $30/hour. Exclude regional or other processing premiums.
Metric6 Sol baseline6.1: Cache only6.1: Less rework6.1: Output/retries
Uncached input / cached reads, million tokens4 / 64 / 63.6 / 5.44.8 / 7.2
Cache writes / output, million tokens0.4 / 10.4 / 10.36 / 0.90.48 / 1.8
Additional retry attempts20201030
Token charges$20.20$19.60$17.64$29.52
Assumed tool charges$3.00$3.00$2.70$3.60
Review minutes / cost240 / $120240 / $120180 / $90300 / $150
Total trial cost$143.20$142.60$110.34$183.12
Accepted tasks, out of 10080809075
Cost per accepted task$1.79$1.78$1.23$2.44
For the baseline, token charges are 4 × $2 + 6 × $0.20 + 0.4 × $2.50 + 1 × $10 = $20.20. Adding tools and review gives $143.20 ÷ 80 = $1.79 per accepted task. The less-rework scenario is $110.34 ÷ 90 ≈ $1.23.

With behavior unchanged, the cache reduction saves only $0.60 across this queue. Less rework can create a larger benefit; more output, retries and review can erase it. The example does not predict which behavior 6.1 Sol will produce. Replace every assumption with paired usage, acceptance and review records, and use verified gateway rates for an EvoLink decision.

Set the pilot gate before running the candidate

Copy this scorecard and replace its example policy with your team's service requirements. These are editorial starting points, not OpenAI recommendations or a statistical guarantee.

MetricRecord for both modelsExample gate
Accepted tasksAccepted/assigned tasks and paired disagreementsCandidate not below baseline; review critical-task regressions separately
Effective costAll-attempt cost / accepted tasksNo higher than the baseline, unless a quality premium was agreed beforehand
LatencyEnd-to-end p95 including tools and retriesWithin a team-defined 75-second budget in this example
Tool correctnessWrong actions, duplicate writes and final record stateZero duplicate or unauthorized writes in the trial; any event blocks a pilot
Contract and billingReturned identity, supported features, usage and actual chargeVerified for the candidate route before production traffic

The following decisions continue the fictional 100-task example; latency and action outcomes are additional assumptions.

OutcomeExample evidenceNext action
Begin a limited pilot90 accepted vs baseline 80; $1.23 vs $1.79; p95 70s; no wrong writes; route/billing verifiedSend a small selected cohort, such as 5%, then recheck the same gates before expanding
Continue offline evaluationNo hard gate breached, but reviewer disagreements leave the quality result unresolvedRe-score disputed tasks and extend the paired set; retain the baseline for production
Reject or roll back75 accepted and $2.44, or any duplicate/unauthorized writeRestore the baseline configuration and diagnose the failing layer
A 100-task trial can expose blockers; it does not establish population-wide superiority or a rare-action failure rate. Inspect failures individually and monitor the pilot. Keep the old GPT-6 Sol configuration available and use the GPT collection for difficult-task alternatives. Save the actual model used so fallbacks do not hide candidate failures.

Rollback restores model, endpoint, effort, tool schema and parser together. Check action state before retrying on another model to avoid duplicate side effects. A unified gateway helps retain model choices; route contracts and retry semantics still need their own checks.

When keeping GPT-6 Sol is the better decision

Stay with the baseline while a tool-dependent client cannot use a verified Responses route, while the new gateway features or bill remain unverified, or while the trial fails your quality and latency thresholds. A low-cache, output-heavy workload may gain little from the direct pricing change. If the current model already passes routine jobs with little rework, prioritize a focused difficult-task trial rather than a broad migration.

Do not call 6 Sol deprecated merely because 6.1 Sol is newer. The references distinguish the models; this article has no confirmed sunset instruction for the old route. Keep the exact old identifier selectable until there is a documented reason to change it.

For the release date and channel details, see the GPT-6.1 Sol release article. For a production switch, the next step is verified access followed by your own paired trial—not assuming that the official announcement proves gateway readiness.

FAQ

Is GPT-6.1 Sol cheaper than GPT-6 Sol?

At OpenAI Standard short-input list rates, ordinary input and output are unchanged; cache reads cost half as much. Whole-job cost depends on cache usage, output, retries and accepted results. EvoLink rates need separate verification.

Can I upgrade by changing only the model ID?

Possibly for a compatible no-tools workflow, after checking supported settings and response handling. A Chat Completions agent using tools at none needs an endpoint and reasoning migration.

Does GPT-6.1 Sol support none reasoning?

No. OpenAI lists low, medium, high, xhigh and max, with medium as the default. Neither none nor minimal is supported.

Does the new Sol have a larger context window?

No. Both references list a 1,050,000-token context window and 128,000 maximum output. Do not mistake the combined context budget for maximum input.

This article has not verified that route, its supported features or selling rates as of September 29. Check the current catalog and documentation before using a gateway-specific configuration.

What is the most useful upgrade metric?

Accepted-task cost together with quality, latency and unsafe-action rejection rules. Include failures and fallback cost, and keep reviewer time visible. A cheaper successful response is not enough if it fails the task.

Sources

Official sources checked September 29, 2026. Workload examples are illustrative calculations and test methods, not measured usage or a guarantee of savings.

Ready to Reduce Your AI Costs by 89%?

Start using EvoLink today and experience the power of intelligent API routing.