GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5
Editorial illustration of a managed agent core coordinating tasks and a finished report
Product Launch

OpenAI Agents API Release: Codex Harness, Tools & Costs

Jessie
Jessie
COO
October 2, 2026
14 min read
OpenAI Agents API brings the Codex harness to application developers as a managed service. The September 10, 2026 announcement describes a public beta: OpenAI runs the agent loop, while your product supplies the task, business tools and an execution environment when needed. This changes the build-versus-buy decision for long-running agents; it does not introduce a new foundation model. Official announcement.

The practical questions behind the launch are more specific than “what is a managed agent?” Can a task survive a disconnected client? Does context compaction preserve the constraints that matter? Will parallel subagents improve completion time enough to justify their cost? And what remains in your application if OpenAI runs the harness?

This release analysis connects those questions to documented mechanisms and a concrete evaluation workload. Checked October 2, 2026: EvoLink integration is in progress and is not yet available. The product page owns platform access and notification updates. The examples below are proposed evaluations, not EvoLink test results.

What has actually been released?

QuestionCurrent answerWhat it means for adoption
Is this an announcement or a usable provider product?OpenAI announced a public beta on September 10Evaluate against beta documentation, not a promised GA contract
Is this Agents SDK under a new name?No: Agents API is a hosted runtime; the SDK runs in your applicationThe migration changes operating responsibility
Is the Codex harness the same as a model?No: it coordinates model calls, tools and ongoing workAssess the complete workflow, not just model quality
Can it be used through EvoLink?Integration is underway; calls are not open yetKeep the existing verified path while preparing a candidate workload
Does a published API imply open model weights?No weight or model-license claim follows from this runtime releaseSeparate runtime source availability, hosted service terms and model licensing

The name collision matters. A tutorial that installs the Agents SDK may be useful, but it does not demonstrate the new managed API. Equally, an OpenAI-compatible model endpoint does not establish support for durable agent sessions. Before reusing an example, identify which of those products it calls.

Why exposing the Codex harness matters

For a coding or analysis product, the model call is only one part of the work. The application also needs to decide the next tool action, carry results forward, keep enough context, and continue after an interruption. The managed service packages a runtime around those interactions. Its value is greatest where maintaining that runtime is a substantial part of the team's work.

That does not mean every application should migrate. If a product has a short, fixed sequence—classify an input, validate it, return a structured answer—the orchestrator may already be simple and dependable. The release becomes more relevant when the next action depends on intermediate findings and tasks span many tool calls. The API/SDK/Responses comparison evaluates that choice in detail.

Use the release as an opportunity to inventory engineering work. Separate effort spent improving the business task from effort spent maintaining execution machinery. A managed runtime can only deliver a useful operational saving if the latter is material and the service boundary fits the product.

OpenAI Agents API workflow concept: input files, iterative execution, result inspection and delivery
OpenAI Agents API workflow concept: input files, iterative execution, result inspection and delivery
Workflow concept: input → agent execution → application acceptance → delivery. The application checks the actual file and business rules before treating the task as complete.

Long-running sessions and context compaction: continuity is the feature

OpenAI describes automatic compaction as part of the managed harness, allowing work to continue across context windows. The meaningful promise to evaluate is continuity through a long task, not unlimited perfect memory. Launch explanation.

Consider a proposed repository-maintenance task: fix pagination, preserve the public response shape, do not modify authentication, and attach the relevant test output. After several investigation and editing steps, steer the agent toward an additional edge case. The useful question is whether the final patch still respects the original constraints. A long transcript alone cannot answer that.

Keep the acceptance rules in the application's task record. At review time, compare the changed files, response fixture and test evidence against those rules. This lets you distinguish a context-handling failure from a weak initial task description. It also gives the team an independent record if it later evaluates another runtime.

The same principle applies to data analysis. A report can remain conversationally coherent while changing its accounting window or dropping a requested exclusion. Evaluate those invariants explicitly. Do not turn “persistent session” into an untested claim that every business constraint will survive every continuation.

Tool search and programmatic calling solve different sources of overhead

Tool search loads relevant tool definitions when needed. Programmatic Tool Calling lets code combine or filter tool results before returning a smaller result to the model. The first concerns which capabilities enter context; the second concerns how work using those capabilities is organized.

For a proposed account-analysis assistant, there may be many available CRM operations but only a few relevant to a specific question. Separately, the assistant might need to retrieve several account records and calculate a summary. Those are different optimization opportunities: discovering the right operation does not remove the cost of fetching or processing its data.

An evaluation should therefore keep two records: whether the correct tool was selected, and whether the computed result matches an independently calculated answer. A small final response is not evidence of a correct aggregation. Include missing records, empty results and inconsistent units in the fixture.

Our recommendation is to keep final customer-facing writes outside an opaque aggregation step. Let the analysis produce a proposed change, then have application code validate the target and authorization before applying it. That is a workflow design choice, not a claim that EvoLink currently exposes either tool mechanism.

Parallel subagents: useful for separable work, not an automatic speed win

The multi-agent documentation describes subagents with their own context, coordinated by a main agent. That makes document-by-document review or independent investigation plausible candidates. It does not establish a fixed speedup for a particular workload.

For a proposed incident investigation, separate deployment changes, error samples and dependency health into independent read-only tasks. Require each to return evidence, a hypothesis and a confidence limitation. The coordinating agent then checks whether the explanations agree before proposing a mitigation.

Contrast that with three agents modifying the same configuration file. The apparent parallelism can create conflicting edits and extra reconciliation. Start with separate evidence gathering and one owner for the final modification. Measure end-to-end accepted completion time, including synthesis and conflict resolution—not only the duration of the fastest subtask.

A useful comparison is one agent versus a bounded subagent configuration on the same incident fixtures. Count total spend and manual corrections alongside time. If more parallel work makes the final answer faster but less verifiable, the optimization has missed the product requirement.

Hosted sandbox, your own environment, or no sandbox?

OpenAI documents three environment modes: no execution environment, an OpenAI-hosted sandbox, or a self-hosted environment. A task using external tools need not automatically receive a filesystem and shell. Architecture.
Proposed workloadStarting point to evaluateWhyApplication responsibility
Account lookup through business functionsNo sandboxThe task needs service results, not local computationAuthenticate the user and enforce record access
CSV analysis that produces a workbookHosted sandboxFiles, packages and output artifacts are centralValidate totals and save the accepted deliverable
Repository work with specialized dependenciesSelf-hosted environmentExisting compute or tooling may be importantProvision, isolate, recover and retire compute
The ecosystem discussion around “bring your own sandbox” is relevant, but it addresses execution placement. The self-hosted guide describes an executor connecting outward to the managed service. Running the executor on your infrastructure does not relocate the hosted harness.
For file-producing applications, also distinguish workspace files from published artifacts. OpenAI's files guide describes different retrieval paths for hosted and self-hosted environments. A report workflow should finish by retrieving and validating the actual deliverable. A message saying “the file is ready” is not the deliverable.
The retrieval path also changes the application code. In an OpenAI-hosted environment, the guide says files written under /workspace/outputs are published as immutable artifacts when the turn completes. Match both the completed turn and file path when retrieving a report, so a revised analysis does not accidentally serve the previous version. Published artifacts survive environment expiration; save retained copies before deleting the session.
In a self-hosted environment, retrieve files through your infrastructure or sandbox provider and copy them into application storage before that environment expires. Writing to /workspace/outputs does not publish those files through OpenAI's Artifacts API. For the same CSV workflow, this means two different download adapters behind one customer-facing “Download report” action. File retrieval and lifetime.
OpenAI Agents API file delivery concept: two environment paths bring reports into application storage
OpenAI Agents API file delivery concept: two environment paths bring reports into application storage
Top: the hosted artifact retrieval path. Bottom: file retrieval through your infrastructure. Both require application-side retention and delivery.

From an agent run to a downloadable report

Here is a suggested product flow for that CSV task. These are application states, not API event names:

Customer seesApplication actionWhat allows the next step?
ProcessingAssociate the business job with the managed session and follow progressA result is ready to inspect
Checking resultsRetrieve the correct report version and check totals and exceptionsThe workbook passes the saved acceptance rules
Ready to downloadSave the accepted result and enforce the user's download permissionsThe file is actually retrievable
ReconnectingRead the existing session and saved work after a stream interruptionEstablish whether the original job is still running or has a result
Needs attentionPreserve useful output and explain the unresolved issueA corrected input, review decision or controlled retry
The quickstart instructs developers to inspect execution results after turn completion and retrieve saved session state before retrying a disconnected run. That is why a UI should show “checking results” before announcing a completed business report. This flow also keeps a transient browser disconnect from turning into a second billable job by default.

What “no additional API fee” does—and does not—mean

The announcement states that using Agents API adds no separate API fee. Model, tool and hosted-container charges still matter. This is not a free-agent offer and does not establish EvoLink's future price. Announcement, official cost categories.

For budget planning, build a task ledger with model charges, tool charges, environment charges, failed attempts and accepted outputs. Keep your own operating and review time in a separate column. This avoids comparing a token-only estimate with a fully loaded production cost.

Illustrative arithmetic, not a quote or benchmark: suppose a batch costs $24 in model usage, $6 in tools and $10 in environment usage. If 80 outputs pass acceptance, the direct cost is $40 / 80 = $0.50 per accepted result. Dividing by 100 submitted tasks would produce $0.40, but that would conceal the 20 failures. If ten of those failures are rescued by another $8 of work, the new batch cost is $48 / 90, approximately $0.53 per accepted result. More completed tasks can be worthwhile even when the unit cost rises.
For idle environments, measure the actual lifecycle and applicable bill rather than projecting from one active run. Session duration, compute lifetime and customer waiting time are separate observations. The sandbox lifecycle guide explicitly distinguishes the session from its environment. That distinction should be present in the budget worksheet.

Beta constraints that can change the decision

As checked on October 2, the overview states US-only data residency and no Zero Data Retention (ZDR) support, including with a self-hosted sandbox. If the application requires otherwise, treat that as an adoption blocker for the current service rather than a setting to fix later.
Tool execution also needs a precise boundary. The functions guide says the application handles function calls and returns their results; attaching a sandbox does not automatically move those handlers into it. Budget for those workers and their failure handling.
Finally, a completed turn is not proof that every intended action succeeded. The quickstart distinguishes turn completion, failure and idle state. Product acceptance should be based on the artifact or destination-system result, not a single lifecycle label.

A first evaluation that produces a decision

Use a deliberately bounded report task: reconcile two CSV exports, explain unmatched rows, and produce a workbook plus a summary. The following is an editorial test design; no results are claimed here.

Test caseFixture or interruptionAcceptance evidence
Ordinary inputTwo exports with a known matching totalWorkbook totals equal an independent calculation
Ambiguous dataDuplicate identifiers and a missing currencyAmbiguities are surfaced; no silent fabricated match
Client disconnectClose the stream after work beginsExisting work is located; the job is not blindly resubmitted
Changed instructionAdd an exclusion during the taskFinal totals and explanation reflect the revised scope
Output retrievalEnd compute after preserving the resultApplication can retrieve and open the accepted deliverable
Budget reviewInclude all attempts in the batchSpend and accepted-output count can be reconciled
Record task version, model, runtime configuration, route, timestamps and acceptance result together. Keep failed cases instead of presenting only the attractive demo. Promote the workload only if it meets your current quality requirement and offers a measurable benefit in completion time, operational work or cost. For model requests that must ship now, retain the verified integration or inspect EvoLink's model catalog; managed-session support is a separate decision.

FAQ

When was OpenAI Agents API released?

OpenAI announced the public beta on September 10, 2026. That date does not establish GA or an EvoLink launch date.

Is this the Codex model or the Agents SDK?

It is a managed service built around the Codex harness. Model choice and application-run SDK orchestration are separate concepts.

Does automatic compaction mean unlimited memory?

No. Evaluate whether the task's constraints and evidence remain usable across long work. Keep the business acceptance record outside the conversation.

Will subagents always reduce cost or latency?

No workload-independent result follows from parallel execution. Include coordination, duplicated work and final verification in the comparison.

Is Agents API free?

No separate API fee does not eliminate model, tool or environment charges. This article's cost arithmetic is illustrative; EvoLink pricing is not announced.

Can I keep everything inside my VPC?

Do not infer that from a self-hosted sandbox. The execution environment and OpenAI's managed service remain different boundaries; the current residency and retention limits still matter.

Not yet. Subscribe to access updates for the integration under preparation. A subscription is a notification request, not an API credential.

What would cause this article to change?

A new beta/GA contract, changed environment or data controls, verified gateway support, or reproducible workload results. Those events should update the relevant facts and recommendations, rather than merely refreshing the date.

Technical sources are linked beside the mechanisms they support. The earlier market research also identified the HN API-versus-SDK question and the framework-replacement debate. Those discussions explain the questions addressed here; they are not evidence for API capabilities or measured savings.
Continue with the runtime comparison for framework choice, paired evaluation and migration rollback, or the EvoLink status page for gateway availability.