
OpenAI Agents API Release: Codex Harness, Tools & Costs
The practical questions behind the launch are more specific than “what is a managed agent?” Can a task survive a disconnected client? Does context compaction preserve the constraints that matter? Will parallel subagents improve completion time enough to justify their cost? And what remains in your application if OpenAI runs the harness?
What has actually been released?
| Question | Current answer | What it means for adoption |
|---|---|---|
| Is this an announcement or a usable provider product? | OpenAI announced a public beta on September 10 | Evaluate against beta documentation, not a promised GA contract |
| Is this Agents SDK under a new name? | No: Agents API is a hosted runtime; the SDK runs in your application | The migration changes operating responsibility |
| Is the Codex harness the same as a model? | No: it coordinates model calls, tools and ongoing work | Assess the complete workflow, not just model quality |
| Can it be used through EvoLink? | Integration is underway; calls are not open yet | Keep the existing verified path while preparing a candidate workload |
| Does a published API imply open model weights? | No weight or model-license claim follows from this runtime release | Separate runtime source availability, hosted service terms and model licensing |
The name collision matters. A tutorial that installs the Agents SDK may be useful, but it does not demonstrate the new managed API. Equally, an OpenAI-compatible model endpoint does not establish support for durable agent sessions. Before reusing an example, identify which of those products it calls.
Why exposing the Codex harness matters
For a coding or analysis product, the model call is only one part of the work. The application also needs to decide the next tool action, carry results forward, keep enough context, and continue after an interruption. The managed service packages a runtime around those interactions. Its value is greatest where maintaining that runtime is a substantial part of the team's work.
Use the release as an opportunity to inventory engineering work. Separate effort spent improving the business task from effort spent maintaining execution machinery. A managed runtime can only deliver a useful operational saving if the latter is material and the service boundary fits the product.

Long-running sessions and context compaction: continuity is the feature
Consider a proposed repository-maintenance task: fix pagination, preserve the public response shape, do not modify authentication, and attach the relevant test output. After several investigation and editing steps, steer the agent toward an additional edge case. The useful question is whether the final patch still respects the original constraints. A long transcript alone cannot answer that.
Keep the acceptance rules in the application's task record. At review time, compare the changed files, response fixture and test evidence against those rules. This lets you distinguish a context-handling failure from a weak initial task description. It also gives the team an independent record if it later evaluates another runtime.
The same principle applies to data analysis. A report can remain conversationally coherent while changing its accounting window or dropping a requested exclusion. Evaluate those invariants explicitly. Do not turn “persistent session” into an untested claim that every business constraint will survive every continuation.
Tool search and programmatic calling solve different sources of overhead
For a proposed account-analysis assistant, there may be many available CRM operations but only a few relevant to a specific question. Separately, the assistant might need to retrieve several account records and calculate a summary. Those are different optimization opportunities: discovering the right operation does not remove the cost of fetching or processing its data.
An evaluation should therefore keep two records: whether the correct tool was selected, and whether the computed result matches an independently calculated answer. A small final response is not evidence of a correct aggregation. Include missing records, empty results and inconsistent units in the fixture.
Our recommendation is to keep final customer-facing writes outside an opaque aggregation step. Let the analysis produce a proposed change, then have application code validate the target and authorization before applying it. That is a workflow design choice, not a claim that EvoLink currently exposes either tool mechanism.
Parallel subagents: useful for separable work, not an automatic speed win
For a proposed incident investigation, separate deployment changes, error samples and dependency health into independent read-only tasks. Require each to return evidence, a hypothesis and a confidence limitation. The coordinating agent then checks whether the explanations agree before proposing a mitigation.
Contrast that with three agents modifying the same configuration file. The apparent parallelism can create conflicting edits and extra reconciliation. Start with separate evidence gathering and one owner for the final modification. Measure end-to-end accepted completion time, including synthesis and conflict resolution—not only the duration of the fastest subtask.
A useful comparison is one agent versus a bounded subagent configuration on the same incident fixtures. Count total spend and manual corrections alongside time. If more parallel work makes the final answer faster but less verifiable, the optimization has missed the product requirement.
Hosted sandbox, your own environment, or no sandbox?
| Proposed workload | Starting point to evaluate | Why | Application responsibility |
|---|---|---|---|
| Account lookup through business functions | No sandbox | The task needs service results, not local computation | Authenticate the user and enforce record access |
| CSV analysis that produces a workbook | Hosted sandbox | Files, packages and output artifacts are central | Validate totals and save the accepted deliverable |
| Repository work with specialized dependencies | Self-hosted environment | Existing compute or tooling may be important | Provision, isolate, recover and retire compute |
/workspace/outputs are published as immutable artifacts when the turn completes. Match both the completed turn and file path when retrieving a report, so a revised analysis does not accidentally serve the previous version. Published artifacts survive environment expiration; save retained copies before deleting the session./workspace/outputs does not publish those files through OpenAI's Artifacts API. For the same CSV workflow, this means two different download adapters behind one customer-facing “Download report” action. File retrieval and lifetime.
From an agent run to a downloadable report
Here is a suggested product flow for that CSV task. These are application states, not API event names:
| Customer sees | Application action | What allows the next step? |
|---|---|---|
| Processing | Associate the business job with the managed session and follow progress | A result is ready to inspect |
| Checking results | Retrieve the correct report version and check totals and exceptions | The workbook passes the saved acceptance rules |
| Ready to download | Save the accepted result and enforce the user's download permissions | The file is actually retrievable |
| Reconnecting | Read the existing session and saved work after a stream interruption | Establish whether the original job is still running or has a result |
| Needs attention | Preserve useful output and explain the unresolved issue | A corrected input, review decision or controlled retry |
What “no additional API fee” does—and does not—mean
For budget planning, build a task ledger with model charges, tool charges, environment charges, failed attempts and accepted outputs. Keep your own operating and review time in a separate column. This avoids comparing a token-only estimate with a fully loaded production cost.
Beta constraints that can change the decision
A first evaluation that produces a decision
Use a deliberately bounded report task: reconcile two CSV exports, explain unmatched rows, and produce a workbook plus a summary. The following is an editorial test design; no results are claimed here.
| Test case | Fixture or interruption | Acceptance evidence |
|---|---|---|
| Ordinary input | Two exports with a known matching total | Workbook totals equal an independent calculation |
| Ambiguous data | Duplicate identifiers and a missing currency | Ambiguities are surfaced; no silent fabricated match |
| Client disconnect | Close the stream after work begins | Existing work is located; the job is not blindly resubmitted |
| Changed instruction | Add an exclusion during the task | Final totals and explanation reflect the revised scope |
| Output retrieval | End compute after preserving the result | Application can retrieve and open the accepted deliverable |
| Budget review | Include all attempts in the batch | Spend and accepted-output count can be reconciled |
FAQ
When was OpenAI Agents API released?
OpenAI announced the public beta on September 10, 2026. That date does not establish GA or an EvoLink launch date.
Is this the Codex model or the Agents SDK?
It is a managed service built around the Codex harness. Model choice and application-run SDK orchestration are separate concepts.
Does automatic compaction mean unlimited memory?
No. Evaluate whether the task's constraints and evidence remain usable across long work. Keep the business acceptance record outside the conversation.
Will subagents always reduce cost or latency?
No workload-independent result follows from parallel execution. Include coordination, duplicated work and final verification in the comparison.
Is Agents API free?
No separate API fee does not eliminate model, tool or environment charges. This article's cost arithmetic is illustrative; EvoLink pricing is not announced.
Can I keep everything inside my VPC?
Do not infer that from a self-hosted sandbox. The execution environment and OpenAI's managed service remain different boundaries; the current residency and retention limits still matter.
Is it available through EvoLink?
What would cause this article to change?
A new beta/GA contract, changed environment or data controls, verified gateway support, or reproducible workload results. Those events should update the relevant facts and recommendations, rather than merely refreshing the date.


