
GPT-6.1 Sol vs GPT-6 Sol: Same Input/Output Rates, Lower Cache Cost, New Migration Rules
none reasoning and requires Responses for tool calling. An agent that used tools on Chat Completions needs more than a model-ID swap.GPT-6.1 Sol vs GPT-6 Sol at a glance
| Decision variable | GPT-6 Sol | GPT-6.1 Sol | What changes for an existing agent |
|---|---|---|---|
| Official model ID | gpt-6-sol | gpt-6.1-sol | Pin the version; do not rely on a generic Sol alias |
| Input / output | Text and image / text | Text and image / text | No new output modality to design around |
| Context / max output | 1,050,000 / 128,000 tokens | Same limits | No larger context window |
| Knowledge cutoff | April 20, 2026 | April 30, 2026 | A newer cutoff is not a quality result on your tasks |
| Reasoning effort | none, low, medium, high, xhigh, max | low, medium, high, xhigh, max; no none or minimal | Remove unsupported settings from the candidate configuration |
| Chat Completions function calling | Only at none effort | Not supported | Move tool-dependent workflows to a verified Responses route |
| Responses tool calling | Supported | Supported | Still verify your tool loop and gateway behavior |
| Streaming / structured outputs | Listed by OpenAI | Listed by OpenAI | Provider support does not prove route-specific support |
medium reasoning. The unchanged limits make compatibility and accepted-task cost more useful decision variables than context size.What the official performance evidence says
| Evaluation | Reported change from GPT-6 Sol | Conditions and decision boundary |
|---|---|---|
| DeepSWE v1.1 | +6.4 percentage points over the old model's best score | The candidate used lower effort and cost; this is not an equal-effort comparison |
| AutomationBench 1.0.6 | +4.8 percentage points | Same medium setting; relevant to multi-step tool workflows |
| OSWorld 2.0 | +7 percentage points, at less than half the task cost | Maximum effort; partial reward on the offline v2026.08.08 set |
The evidence makes complex code changes and business tool workflows useful trial candidates. OSWorld's partial reward measures progress toward a task; it cannot be substituted for your binary acceptance rate. OpenAI also says its research/API environment can differ from production ChatGPT. Preserve the harness, effort and cost scope when quoting a result.
For routine jobs your current model already passes, the announcement provides less reason for a broad replacement. Use the paired tasks below to establish which gains transfer to your workload. Do not infer your latency, retry rate or EvoLink bill from these results.
Decide the endpoint before judging performance
There are three materially different migrations. Treating them as one trial makes failures hard to interpret.
| Existing workflow | Candidate path | First acceptance check |
|---|---|---|
| Chat Completions, no tools, a supported reasoning level | Trial 6.1 Sol's documented no-tools contract | Response parsing, output shape and actual billed usage |
Chat Completions tools with none | Change the tool workflow to Responses and select a supported effort | Tool requests, result association, continuation and retry behavior |
| Responses with tools | Keep the workflow, pin the new ID and a supported effort | Full tool loop, structured results, streaming and cancellation where used |
none can also alter latency, output usage and behavior even if the endpoint remains unchanged.For a tool-dependent migration, first inventory the endpoint, effort, tool definitions, result identifiers and continuation handling. Then adapt the tool loop to Responses, including how results are associated with the requesting call. Replay a successful tool action, a failed action and a cancelled run in a sandbox; inspect resulting records as well as the final text. Check structured output parsing and streaming if your client uses them. Only then start the quality comparison.
Keep compatibility failures separate from rejected tasks. Verify the exact gateway route's contract and billing before a pilot, and retain the old model, endpoint and parser together as a rollback configuration.
The cache-read discount is smaller than a whole-job discount
| Token category | GPT-6 Sol, up to 272K input | GPT-6.1 Sol, up to 272K input | GPT-6.1 Sol, above 272K input |
|---|---|---|---|
| Uncached input | $2.00 | $2.00 | $4.00 |
| Cached input read | $0.20 | $0.10 | $0.20 |
| Cache write | $2.50 | $2.50 | $5.00 |
| Output | $10.00 | $10.00 | $15.00 |
The direct change from 6 Sol is a 50% lower cache-read rate. Relative to 6.1 Sol's own ordinary input rate, a cached read costs 95% less. Neither statement means that an entire job is 50% or 95% cheaper. Output, uncached input and cache writes have not received that same reduction.
| Cached share of the one million input tokens | GPT-6 Sol input + output | GPT-6.1 Sol input + output | Direct saving |
|---|---|---|---|
| 40% | $1.20 uncached + $0.08 cached + $1.00 output = $2.28 | $1.20 + $0.04 + $1.00 = $2.24 | $0.04 |
| 90% | $0.20 uncached + $0.18 cached + $1.00 output = $1.38 | $0.20 + $0.09 + $1.00 = $1.29 | $0.09 |
Those shares are assumptions, not measured hit rates or a promise that a reused prompt will qualify for caching. If the candidate produces more reasoning output or needs an extra retry, that usage can absorb the illustrated saving. Conversely, a higher accepted-task rate could create value much larger than the cache difference. Measure both effects instead of declaring success from the token-price table.
Build a paired evaluation around six real jobs

Freeze recent tasks, repository versions and acceptance rules before running either model. Include examples where the old agent failed or needed human intervention, alongside routine jobs. Keep tool permissions and retry limits consistent; report necessary endpoint or harness changes instead of pretending the experiment is perfectly controlled.
| Job | Input and expected result | Acceptance signal | Cost or failure signal | Trial priority |
|---|---|---|---|---|
| Multi-file bug fix | Fixed repository and issue → patch | Required tests pass without unrelated changes | Retries, rework and reviewer intervention | Start with failures and review-heavy changes |
| PR review | Fixed diff and conventions → findings | Confirmed defects with tolerable false positives | Time spent checking false alarms | Keep the baseline if extra findings add review noise |
| Repeated-context agent | Saved context and allowed tools → completed task | Constraints retained; no duplicate side effects | Cache reads/writes and reasoning output | Trial when usage confirms substantial cache reads |
| Document questions | Fixed documents and questions → supported answers | Correct fields and traceable evidence | Unsupported conclusions and review time | Test tables and conflicting evidence, not only easy lookup |
| Business tool workflow | Goal and sandbox tools → expected final records | Correct sequence and record state | Duplicate writes or mismatched tool results | Verify the tool loop before judging model quality |
| Difficult-task escalation | Fixed task queue → accepted results | Quality threshold met across the queue | Candidate, escalation and fallback costs combined | Trial a separate difficult-task route first |
Choose rejection rules before the run. For example, a business workflow may reject any duplicate write even if its final textual answer looks correct. A patch may need hidden tests as well as the tests the agent can see. JSON validity alone cannot establish that extracted fields are accurate.
Track task ID, model identity, endpoint, effort, tool calls, usage, retries, acceptance and reviewer time for every attempt. Summarize sample size and paired disagreements; a small task set can identify a blocker but cannot establish a reliable population-wide improvement. Inspect individual failures as well as averages, particularly when one long-running job dominates the bill.
Compare accepted-task cost, then stage the switch
Use this accounting definition:
Keep human review cost explicit: either price it with a documented rate or report review minutes alongside API cost. Do not quietly omit it from one model's result. If no task is accepted, the ratio is undefined; record a failed trial rather than reporting a cheap result.
A complete 100-task accounting example
| Metric | 6 Sol baseline | 6.1: Cache only | 6.1: Less rework | 6.1: Output/retries |
|---|---|---|---|---|
| Uncached input / cached reads, million tokens | 4 / 6 | 4 / 6 | 3.6 / 5.4 | 4.8 / 7.2 |
| Cache writes / output, million tokens | 0.4 / 1 | 0.4 / 1 | 0.36 / 0.9 | 0.48 / 1.8 |
| Additional retry attempts | 20 | 20 | 10 | 30 |
| Token charges | $20.20 | $19.60 | $17.64 | $29.52 |
| Assumed tool charges | $3.00 | $3.00 | $2.70 | $3.60 |
| Review minutes / cost | 240 / $120 | 240 / $120 | 180 / $90 | 300 / $150 |
| Total trial cost | $143.20 | $142.60 | $110.34 | $183.12 |
| Accepted tasks, out of 100 | 80 | 80 | 90 | 75 |
| Cost per accepted task | $1.79 | $1.78 | $1.23 | $2.44 |
4 × $2 + 6 × $0.20 + 0.4 × $2.50 + 1 × $10 = $20.20. Adding tools and review gives $143.20 ÷ 80 = $1.79 per accepted task. The less-rework scenario is $110.34 ÷ 90 ≈ $1.23.With behavior unchanged, the cache reduction saves only $0.60 across this queue. Less rework can create a larger benefit; more output, retries and review can erase it. The example does not predict which behavior 6.1 Sol will produce. Replace every assumption with paired usage, acceptance and review records, and use verified gateway rates for an EvoLink decision.
Set the pilot gate before running the candidate
Copy this scorecard and replace its example policy with your team's service requirements. These are editorial starting points, not OpenAI recommendations or a statistical guarantee.
| Metric | Record for both models | Example gate |
|---|---|---|
| Accepted tasks | Accepted/assigned tasks and paired disagreements | Candidate not below baseline; review critical-task regressions separately |
| Effective cost | All-attempt cost / accepted tasks | No higher than the baseline, unless a quality premium was agreed beforehand |
| Latency | End-to-end p95 including tools and retries | Within a team-defined 75-second budget in this example |
| Tool correctness | Wrong actions, duplicate writes and final record state | Zero duplicate or unauthorized writes in the trial; any event blocks a pilot |
| Contract and billing | Returned identity, supported features, usage and actual charge | Verified for the candidate route before production traffic |
The following decisions continue the fictional 100-task example; latency and action outcomes are additional assumptions.
| Outcome | Example evidence | Next action |
|---|---|---|
| Begin a limited pilot | 90 accepted vs baseline 80; $1.23 vs $1.79; p95 70s; no wrong writes; route/billing verified | Send a small selected cohort, such as 5%, then recheck the same gates before expanding |
| Continue offline evaluation | No hard gate breached, but reviewer disagreements leave the quality result unresolved | Re-score disputed tasks and extend the paired set; retain the baseline for production |
| Reject or roll back | 75 accepted and $2.44, or any duplicate/unauthorized write | Restore the baseline configuration and diagnose the failing layer |
Rollback restores model, endpoint, effort, tool schema and parser together. Check action state before retrying on another model to avoid duplicate side effects. A unified gateway helps retain model choices; route contracts and retry semantics still need their own checks.
When keeping GPT-6 Sol is the better decision
Stay with the baseline while a tool-dependent client cannot use a verified Responses route, while the new gateway features or bill remain unverified, or while the trial fails your quality and latency thresholds. A low-cache, output-heavy workload may gain little from the direct pricing change. If the current model already passes routine jobs with little rework, prioritize a focused difficult-task trial rather than a broad migration.
Do not call 6 Sol deprecated merely because 6.1 Sol is newer. The references distinguish the models; this article has no confirmed sunset instruction for the old route. Keep the exact old identifier selectable until there is a documented reason to change it.
FAQ
Is GPT-6.1 Sol cheaper than GPT-6 Sol?
At OpenAI Standard short-input list rates, ordinary input and output are unchanged; cache reads cost half as much. Whole-job cost depends on cache usage, output, retries and accepted results. EvoLink rates need separate verification.
Can I upgrade by changing only the model ID?
none needs an endpoint and reasoning migration.Does GPT-6.1 Sol support none reasoning?
low, medium, high, xhigh and max, with medium as the default. Neither none nor minimal is supported.Does the new Sol have a larger context window?
No. Both references list a 1,050,000-token context window and 128,000 maximum output. Do not mistake the combined context budget for maximum input.
Is GPT-6.1 Sol available through EvoLink?
This article has not verified that route, its supported features or selling rates as of September 29. Check the current catalog and documentation before using a gateway-specific configuration.
What is the most useful upgrade metric?
Accepted-task cost together with quality, latency and unsafe-action rejection rules. Include failures and fallback cost, and keep reviewer time visible. A cheaper successful response is not enough if it fails the task.
Sources
- GPT-6.1 Sol model reference — new contract and limits.
- GPT-6 Sol model reference — baseline contract and limits.
- OpenAI API pricing — Standard rates, cache charges and long-input conditions.
- GPT-6.1 Sol launch announcement — launch framing and vendor-reported evaluations; not an EvoLink reproduction.
Official sources checked September 29, 2026. Workload examples are illustrative calculations and test methods, not measured usage or a guarantee of savings.
