OpenAI Launched the Agents API: What Builders Need to Know
The public beta packages the Codex agent harness, durable sessions, subagents, tools, and selectable execution environments—but production ownership still sits with the builder.

Bottom line
OpenAI's new Agents API manages long-running sessions and multi-agent orchestration while letting builders choose where code executes. Here is what the beta changes.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 1
- Last checked
- 2026-09-11
Important limits
- • Features, availability, and pricing can change after publication; confirm consequential details with the provider.
In this guide
*This research-based analysis covers OpenAI's September 10, 2026 launch. Feature descriptions and customer results are OpenAI's reported information unless otherwise stated. The product is in public beta, so interfaces, limits, eligibility, and prices can change.*
The short answer
OpenAI's Agents API is a public-beta service for running long-lived cloud agents with the same managed harness family used by Codex. A developer supplies the task, model, tools, and environment; OpenAI manages session state, context, orchestration, recovery, and optional subagents. Execution can occur in an OpenAI-hosted sandbox, the customer's infrastructure, or a supported partner sandbox.
This is more than another tool-calling endpoint. It moves difficult agent infrastructure—durable sessions, context compression, parallel work, checkpoints, and tool coordination—into a managed layer. It does not transfer responsibility for permissions, evaluation, observability, human approval, or the consequences of actions.
What changed?
The API can keep a task running across long sessions, preserve intermediate files and state, coordinate multiple subagents, and expose streamed events and session history. OpenAI says the harness evolves as its Codex infrastructure improves. Builders can use MCP tools and capability directories while deciding whether the working environment is hosted or customer controlled.
That separation matters. The harness decides how an agent organizes work; the environment determines what files, credentials, networks, and commands it can reach. A convenient hosted sandbox is not permission to expose production secrets or grant broad write access.
How much does the Agents API cost?
The total is not one flat subscription. Builders should model at least four meters: model tokens, hosted sandbox or partner compute, tool charges such as web or file search, and their own storage, logging, review, and recovery costs. The launch page does not establish a universal all-in price for every agent run.
Measure cost per accepted outcome, not cost per API call. A long session that finishes a reviewed task may be cheaper than many failed short calls; an autonomous run that creates rework is expensive even when token spend looks low.
Should existing agent teams migrate?
Run a matched evaluation first. Use representative tasks and compare completion rate, reviewer corrections, latency, total metered cost, recovery from tool failures, and the clarity of the audit trail. Migration is attractive when a team is maintaining its own session, context, retry, and subagent machinery without creating differentiated value.
Keep a custom orchestrator when execution policy, model portability, deterministic workflows, or specialized scheduling is core to the product. Public beta is also a reason to place an adapter around the API rather than coupling every application component to a changing interface.
Production safety checklist
Give each agent the minimum tools and credentials required. Use short-lived secrets, destination allowlists, spending and runtime limits, approval gates for irreversible actions, immutable logs, idempotent tools, and tested cancellation. Separate development, evaluation, and production environments.
Test failures deliberately: unavailable tools, malicious retrieved text, poisoned files, partial writes, expired credentials, conflicting subagents, context loss, and an interrupted session. The agent should fail safely and leave enough evidence for a person to understand what happened.
The verdict
The Agents API is important because it turns a sophisticated agent harness into managed infrastructure and preserves a choice of execution environment. It is most compelling for teams whose orchestration plumbing is slowing product work.
Treat the beta as infrastructure to evaluate, not autonomy to trust by default. The winning implementation will be the one with the best bounded-task completion rate, evidence trail, and recovery behavior—not the largest number of subagents.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What is the OpenAI Agents API?
It is a public-beta API for building and running long-lived agents with managed sessions, context, tool use, recovery, and optional subagent orchestration.
Does the Agents API require an OpenAI-hosted sandbox?
No. OpenAI says builders can choose an OpenAI-hosted environment, their own infrastructure, or a supported sandbox partner.
Is the Agents API the same as the Responses API?
No. The Agents API adds a durable agent harness and orchestration layer for longer-running work; builders should verify how it fits with the underlying models and tools they already use.
Should production teams migrate immediately?
Not without a matched evaluation. Compare completion quality, review effort, failure recovery, latency, total cost, security boundaries, and interface stability on representative tasks.
Found this useful?
Get the next one in your inbox.
One five-minute briefing a week: a meaningful change, a practical workflow, and a clearer tool decision—already filtered for lean teams.
Free · one email a week · unsubscribe any time
Tools mentioned in this article
ChatGPT
The general-purpose AI assistant that started it all
OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.
Read next
Recommended for you

GPT-6 Astra vs Claude for Business: Which Should Your Team Choose?
Astra is the higher-control agent bet; Claude remains the more flexible model family for many cost-sensitive knowledge and coding workflows.
GPT-6 Astra is best tested for difficult end-to-end agent work, while Claude offers more model and price choices for everyday business workloads.
Read guide
Aomni Review 2026: AI Sales Research, Account Plans, and Pricing
DocsBot AI Review 2026: Pricing, Accuracy, Privacy, and Fit
Scira AI Review 2026: Pricing, Sources, Privacy, and Fit