Ten minutes into testing OpenAI’s new Agents API, I typed the same request I always use to sanity-check a new agent runtime: “Create tree.py, a script that prints a directory tree, run it, and tell me what it actually printed.” Boring on purpose. If a harness can’t handle a five-line script honestly, nothing else about it matters.
The Agents API, public beta since September 10, handled it fine. Session created, sandbox spun up, script written, executed, output reported back through an event stream. Then I ran the identical prompt through Claude Code with a subagent doing the same job. Also fine. Same result, different shape of trust required to get there.
That’s the actual story here — not which one is “better,” but where each one puts the seams you have to reason about when something breaks at 2am.
What OpenAI shipped
Four objects: an Agent (model, instructions, tools, MCP servers), an optional Environment sandbox, a durable Session, and the event stream a session pushes back at you. You pick a hosted sandbox — OpenAI’s own, or a partner one from E2B, Modal, Daytona, Cloudflare, Vercel, and a handful of others — or bring nothing and let the model reason without executing code.
client.beta.agents.sessions.create(
agent={
"model": "gpt-6-astra",
"instructions": "Write clean code, run it, and report the actual output."
},
environment={"type": "openai_hosted"},
input="Create tree.py, a Python script that prints a readable tree of the current directory.",
stream=True,
)
Everything after that call is OpenAI’s problem: keeping the session alive, compacting context once it fills, recovering from a crash mid-turn, coordinating sub-agents if the task spawns them. You consume events — agent.session.turn.completed, turn.failed, session.failed — and steer from the outside. It’s the Codex harness, the same one running behind ChatGPT’s coding features, now callable directly, billed only in model tokens, tool calls, and sandbox compute.
What Claude Code does differently
Claude Code doesn’t hand you an event stream to a black box — it runs the orchestration loop in your own process, and the subagents it spawns are visible, interruptible, and inspectable while they run. When I ran the tree.py task through a Claude Code subagent, I could watch every tool call as it happened, not just the terminal event. If the subagent had gone sideways — say, it decided to rm -rf something it shouldn’t — I’d have seen the tool-use block before it executed, in the same context I was already working in.
That’s the trade being made. OpenAI’s model treats the harness as infrastructure: you rent reliability, you don’t own the failure modes. Claude Code’s model treats the harness as something you operate inside of: more visibility, but you’re also the one who has to build the guardrails, since nothing is compacting your context or recovering your session for you unless you wire that up yourself.
Neither is wrong. They’re built for different failure budgets.
Where this actually matters for a team
We run three categories of automated agent work at any given time: customer-facing triage (low tolerance for silent failure, needs a human loop fast), internal data cleanup (medium tolerance, batchable, can retry), and one-off exploratory scripts an engineer kicks off and walks away from (high tolerance, nobody’s watching).
The exploratory category is where OpenAI’s Agents API earns its keep immediately. You don’t want to babysit context compaction for a script somebody fired off before lunch. Session recovery being someone else’s job is a feature, not a gap.
The triage category is where I’d keep Claude Code, specifically because the failure needs to be visible to a human in the same breath it happens, not reconstructed afterward from an event log. When our support-ticket agent misclassified a churn-risk ticket as low priority last month, the fix took four minutes because I was watching the reasoning trace live in the same session. An event stream I had to query after the fact would have cost us the same afternoon.
The part nobody’s benchmark will tell you
Cost isn’t the differentiator people expect it to be — both land in a similar band once you account for sandbox compute versus local execution. The differentiator is organizational: does your team want an agent runtime that behaves like a managed cloud service (opaque, reliable, billed cleanly), or one that behaves like a library you import into work you’re already directly responsible for?
If you’re a five-person startup shipping features fast, the managed version buys you back engineering hours you’d otherwise spend building your own compaction and retry logic. If you’re running agents against production data with a compliance team asking “show me exactly what the model saw and did,” you want the harness in your own process where you control the logging, not behind someone else’s event API.
We ended up running both, on purpose, split exactly along that line. Not because it’s the elegant answer — because it’s the honest one.
Try it yourself in fifteen minutes
export OPENAI_API_KEY=sk-...
python3 -c "
from openai import OpenAI
client = OpenAI()
stream = client.beta.agents.sessions.create(
agent={'model': 'gpt-6-astra', 'instructions': 'Be concise. Run code, report real output.'},
environment={'type': 'openai_hosted'},
input='Write and run a script that lists the 5 largest files in /tmp.',
stream=True,
)
for event in stream:
print(event.type)
"
Watch the event types scroll by. Then go run the same task through a Claude Code subagent and watch the tool calls scroll by in your own terminal instead. The gap between those two experiences is the actual decision you’re making — not a benchmark score.