GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5
Editorial illustration comparing managed runtime, application orchestration and a direct model interface
Comparison

OpenAI Agents API vs Agents SDK: Who Runs Your Agent?

Jerry
Jerry
CGO
October 2, 2026
16 min read
Evaluate Agents API when maintaining the agent runtime is the work you want to outsource. Keep Agents SDK when orchestration behavior is part of your application's design. Use Responses API directly when your existing workflow engine already owns the sequence, state and recovery. A team can use more than one approach for different workloads.
The API-versus-SDK question appeared in the launch discussion on HN, while a separate developer discussion asked whether managed agents displace frameworks. These are useful adoption questions, not a benchmark. This guide answers them through responsibility boundaries, concrete workload choices and a migration test.
Checked October 2, 2026. This is a documented-capability comparison with editorial evaluation methods, not a head-to-head runtime test. EvoLink is still preparing Agents API integration; platform availability is separate from the architectural choice.

API, SDK and Responses: what changes in the application?

Agents API provides managed orchestration. Agents SDK supplies orchestration components that execute in your application. Responses API is the lower-level model/tool interface around which you can compose your own workflow. The SDK uses Responses by default for OpenAI models, so SDK versus Responses is not necessarily a choice between unrelated backends. Agents SDK overview.
DecisionAgents APIAgents SDKDirect Responses API
Who runs orchestration?OpenAI's managed serviceYour application runs the SDKYour application or existing workflow engine
Where is continuing work represented?Managed sessions linked to your business tasksYour chosen session/state integrationYour workflow record plus API state features used
How do business tools run?Application handlers still execute function toolsSDK tools integrate with your application's codeYour dispatcher handles client-owned tool work
What is the main engineering trade?Less runtime operation, an external service boundaryRuntime control plus deployment responsibilityDirect composition plus explicit workflow ownership
What survives a runtime change?Only what you make portable outside the serviceBusiness records and adapters you keep independentBusiness records and workflow contracts you retain

The last row is an architectural recommendation. No product label guarantees portability. A tool implementation may be reusable while its pending-call record, approval state and result envelope still need adaptation.

Runtime ownership concept for OpenAI Agents API, Agents SDK and direct Responses calls
Runtime ownership concept for OpenAI Agents API, Agents SDK and direct Responses calls
Left to right: a managed runtime, orchestration inside the application, and model calls organized by the application. Business responsibilities remain with the application in all three.

Will Agents API replace LangGraph or your existing agent framework?

Our assessment is that replacement depends on what the framework does in your product. If it primarily maintains a generic model/tool loop, a managed runtime may replace a substantial portion of that work. If it encodes business routing, approval transitions, deadlines and durable domain state, those responsibilities still need a home.

For example, an insurance-document workflow might extract information, pause for an authorized reviewer, and send a signed-off result to a downstream system. The reasoning step can change runtime without changing who is allowed to approve. Replacing the whole workflow because one execution step became managed would mix business policy with infrastructure selection.

Inventory each existing component as one of three things: business rule, execution mechanism, or integration adapter. Then identify which execution mechanisms the managed service can actually replace. This is a more useful migration estimate than comparing the number of lines in two quickstarts.

A hybrid is also legitimate. Keep a deterministic outer workflow and delegate a bounded investigation to Agents API. Define a single input, expected evidence and a return condition for that step. Avoid letting both the outer workflow and the inner runtime independently decide when to repeat the same external write.

This is not a claim that every named framework lacks managed features or that one framework is obsolete. The decision concerns the responsibilities in your implementation, not a universal ranking of framework brands.

Sessions, memory and approvals are different kinds of state

It would be inaccurate to describe Agents SDK as “stateless.” Its session documentation includes persistent session implementations. Its human-in-the-loop flow supports paused runs and serialized RunState. The distinction is who deploys and operates the mechanism, not whether the SDK has memory or approvals.

For a refund-review assistant, keep three records conceptually separate:

RecordExample contentsWhy a conversation alone is insufficient
Working contextEvidence gathered and candidate explanationIt helps reasoning but is not the authorization record
Execution stateCurrent run, pending tool call, continuation referenceIt identifies where work can resume
Business stateCustomer, proposed amount, reviewer, final transaction identifierIt establishes what was permitted and what happened

These are suggested application records, not three required tables or an OpenAI schema. Their purpose is to make a resumed run answer “is this action still authorized?” before performing it. If a customer cancels while approval is pending, resuming execution should not resurrect the old permission.

With managed sessions, map the service reference to the business job. With the SDK, decide where session and paused-run data are stored and how a worker retrieves them. With direct Responses, define the equivalent continuation in the existing workflow. In every case, test a process restart during approval rather than assuming a successful in-memory demo proves recovery.

Is a private sandbox equivalent to self-hosting the agent?

No. OpenAI's self-hosted environment guide separates your executor from the managed harness. The executor connects outward and performs work in your environment. That gives control over execution infrastructure, not ownership of the entire hosted service.
The current Agents API overview states US-only data residency and no ZDR eligibility, including with self-hosted sandboxes. For a workload requiring an incompatible data boundary, stop that candidate at the architecture review. Running SDK code locally does not automatically solve the issue either: its model, tracing and tool destinations still need their own review.

Map the actual data path. For a database investigation, distinguish the database query, the returned rows, the tool result sent to the model, and the trace retained for debugging. Keeping the database in a VPC does not mean the query result never leaves it.

Tool placement matters too. The sandbox security guide distinguishes remote MCP connections from connections originating in the executor. A private tool endpoint that is reachable from your worker may not be reachable from the managed service. Resolve that placement before treating “MCP supported” as an integration result.

What if multi-model choice is a requirement?

Keep the model interface and runtime interface separate. For an application using a gateway, access to several models can simplify model selection, but it does not make every provider implement the same session lifecycle or tool protocol.

A practical evaluation has two stages. First compare runtime options with the same model, tools, data and acceptance rules where that is possible. Then evaluate different models within the chosen architecture. If one candidate requires a different model or tool configuration, label the comparison as a whole-system comparison; do not attribute every improvement to the runtime alone.

For an SDK-based route, verify the adapter's actual capabilities: input messages, tool arguments and results, structured output, streaming, usage and error handling. A basic text response is insufficient evidence for a tool-heavy agent. For the managed API, do not infer support for arbitrary gateway models from SDK provider flexibility.

The useful fallback boundary is often a new business task. Route that task to a verified alternative with a fresh execution record. Moving an in-progress managed session to another model or runtime requires a deliberate state conversion and replay policy; changing a base URL is not that policy.

Choose by workload, including when to keep what you have

Workload and existing systemFirst candidateWhy it is plausibleWhat would reverse the choice?
Small team, variable-length research tasks, little existing orchestrationAgents APIRuntime operation is a substantial new burdenData boundary mismatch or no quality/operating benefit
Product with custom approval paths and application workersAgents SDKExecution can remain close to existing controlsWorker and state operations outweigh the control benefit
Reliable workflow engine with a few fixed model stepsDirect ResponsesReuse the existing state machineTasks require an adaptive loop the team cannot maintain economically
File-heavy work using specialized internal computeSDK or managed API with a self-hosted environmentBoth deserve evaluation against the actual infrastructureRequired connection, isolation or retrieval path cannot be satisfied
Multiple independent evidence-gathering tasksManaged or SDK orchestrationParallel work may shorten the critical pathSynthesis cost, duplicate work or verification effort erases the benefit
These are starting hypotheses. The release's context compaction, tool search and subagents are mechanisms to test against a bottleneck, not reasons to introduce all of them. The release analysis explains those mechanisms. Here the decision is whether they replace work your team currently needs to do.

Test recovery at the point where a duplicate action becomes possible

Use a non-production ticketing fixture. The task should investigate an issue and create exactly one ticket after approval. Interrupt the client after the destination has accepted the ticket but before the application has recorded the tool result. This is a proposed failure-injection scenario, not a reported provider defect.

On recovery, inspect the application's action record and the destination's ticket identifier before retrying. “The stream stopped” describes the observer, not necessarily the job. OpenAI's errors guide directs developers to saved state for execution failures. The application must then reconcile that state with its own business result.
For function tools, the documented flow identifies pending calls through required actions; a historical call item alone does not mean it is still waiting for execution. That detail is important when replaying or reconnecting.

Pass this scenario only if the application can explain whether the ticket exists, avoid creating a duplicate, and resume or terminate the task in a known state. If the result is ambiguous, route it to review. Blind retry may make an apparently resilient demo less reliable in production.

Compare accepted-result cost, not just tokens

Use direct charges per accepted result and keep engineering effort separate. Include failed attempts, tools, environments and any rescue work. A runtime can improve operating effort while increasing direct API spend, or vice versa; the trade should be visible.

Illustrative paired batch, not measured performance: both configurations receive the same 100 tasks. Configuration A costs $60 and yields 90 accepted results; B costs $48 and yields 72. Both cost about $0.67 per accepted result, despite B's smaller invoice. If rescuing 18 of B's failures costs another $18, B reaches 90 accepted results at $66 / 90, about $0.73. The initial invoice did not identify the cheaper way to deliver 90 usable outputs.

Migration also has a break-even point. Suppose integration and validation cost an internally estimated $1,200, and a later verified steady workload saves $0.04 per accepted task. Recovering that investment requires 30,000 accepted tasks, before recurring operating differences. These are hypothetical inputs; replace them with your team's data. If the workload will not reach that volume, a small unit saving may not justify migration.

Measure latency in comparable terms: time to first progress, time to accepted artifact, and manual review time. Do not compare a streaming first token with a fully verified report. For environments, record setup and cleanup as well as active processing; that reveals whether cold starts or long waits dominate the result.

Trace export helps diagnosis; it is not a business audit by itself

Current observability documentation describes OTLP JSON session trace export. A comparison based on an older “no export” claim is no longer a sound basis for choosing the SDK.

Nevertheless, an exported trace and an accepted business outcome answer different questions. In the ticket fixture, connect the application job to the runtime trace, tool call and destination ticket. Ask a reviewer who did not run the test to explain why the ticket was created and whether it was authorized. If the evidence cannot support that explanation, exporting more spans alone will not solve the audit gap.

Keep model/tool observations and billed charges as separate inputs until reconciled. Do not treat a dashboard screenshot as proof of final unit economics, or a recorded tool invocation as proof that the downstream change committed.

A migration plan with an actual rollback boundary

  1. Freeze a baseline. Save task fixtures, tool versions, acceptance checks and current results. Include a long task, an ambiguous input, an approval pause and a failed external action.
  2. Run paired trials. Use the same model and budgets where possible; record differences where not. Keep multiple attempts for variable tasks and report the sample size, not just the best output.
  3. Shadow read-only work. Compare candidates without allowing the shadow agent to send messages or duplicate writes. Judge outputs against the same acceptance rules.
  4. Canary new tasks. Assign the runtime when a new business job starts. Keep that ownership stable through its lifecycle; do not split one ongoing task between two unsynchronized controllers.
  5. Rollback deliberately. Route new jobs back to the baseline. For in-flight jobs, either drain the existing runtime or reconcile tool effects and artifacts before starting a replacement. Preserve the evidence needed to explain the transition.

Set thresholds before viewing results. Quality must meet the product's existing acceptance requirement; unauthorized writes and duplicate side effects should block rollout; cost and latency budgets should come from the business workload. There is no universal “95% is production-ready” score that replaces those requirements.

For EvoLink users, this separates two projects: choosing the runtime and validating the gateway surface. EvoLink's unified API positioning is relevant to model access, credentials and cost management, but Agents API sessions require their own integration evidence. Follow the status page; keep verified routes for existing applications while that work proceeds. Use the model catalog to investigate alternatives against the operations you need.

A migration example: keep the support workflow, replace its investigation step

Suppose an existing SDK application receives a support issue, gathers account evidence, pauses for approval and creates a ticket. A useful first migration is to replace only evidence gathering. The managed task returns a draft and references; the existing application retains approval and ticket creation. This is a proposed design, not a tested migration result.

Incremental Agents API migration concept: replace the investigation module while preserving intake, approval and ticket delivery
Incremental Agents API migration concept: replace the investigation module while preserving intake, approval and ticket delivery
Migration concept: replace the investigation step first, preserve approval and delivery, and retain the original route for new tasks.
Existing componentKeep or adapt?Concrete migration work
User identity, account permissions and ticket schemaKeep the business contractGive the candidate the same permitted records and required output fields
SDK investigation runnerReplace for the pilot stepStart a managed session and map its reference to the existing business job
Tool implementationsReuse where their contract fits; adapt dispatchTranslate arguments and results, preserve permission checks and record tool failures
Pending approval and stored RunStateKeep with the current owner for in-flight jobsFinish or reconcile the old run; do not treat its serialized state as a managed-session import
UI progress and final resultAdapt the application-facing mappingDistinguish investigation progress, a draft awaiting review and a successfully created ticket
Trace and billing recordsAdd the new referencesAssociate each candidate run with the same task, accepted outcome and cost ledger

The first pilot can therefore end at “draft ready for review.” It need not transfer every part of the workflow at once. If the investigation improves but approval recovery regresses, keep that approval path in the application and narrow the migration instead of accepting an all-or-nothing result.

The SDK's approval guide also matters for what stays: serialized RunState includes pending work and approval decisions, but deserialization does not authenticate who supplied it. Store it under application control, validate the reviewer against the pending action and coordinate resumption so the same approval is not consumed twice. This is concrete application work that remains even when the investigation runtime changes.

FAQ

Does Agents API make existing frameworks obsolete?

It may replace generic runtime work. Business policy, approval transitions and domain state still need an owner. Evaluate components rather than replacing a framework by name.

Is Agents SDK stateless?

No. The SDK documents sessions and persistent implementations. Your application still operates their deployment and storage choices.

Can I preserve human approval with the SDK?

Yes, the documented flow includes interrupted runs and resumable state. Test restart and authorization changes in your own deployment rather than relying only on an in-memory demo.

Does self-hosted compute make Agents API ZDR-compatible?

No, according to the current Agents API documentation. Review the full data path for SDK-based alternatives too; local orchestration alone is not a retention guarantee.

Can I switch runtimes by changing the base URL?

Do not assume that. Tool implementations may be reusable, but sessions, pending calls, state and result handling need an explicit adapter and tests.

Which choice is cheapest?

Measure the same accepted business outcome, including failures and rescue attempts, then include migration and operating effort. The arithmetic above is hypothetical, not a winner declaration.

Can I export traces from Agents API?

The current official docs describe OTLP JSON export. Plan how the exported trace will connect to your monitoring system; consult the EvoLink launch documentation for the available integration path.

Not yet as of October 2, 2026. Request integration updates. That action subscribes to notifications; it does not grant API access.

Sources and scope

Primary documentation is linked beside each technical claim. HN and Reddit are used only to identify the API/SDK and framework questions. Workloads, arithmetic and rollout recommendations are editorial proposals; this article does not claim a controlled runtime benchmark, universal cost saving or production gateway compatibility.