Last week a teammate pinged me with “should we just move our agent runtime to OpenAI’s new Agents API and stop maintaining our own sandbox fleet?” Fair question. OpenAI opened the public beta on September 10, and the pitch is clean: you submit a task, OpenAI runs the reasoning loop on its own infrastructure, and your shell commands execute in a hosted sandbox unless you bring your own. No more babysitting container images.

I spent two evenings reading the pricing docs and a teardown from tokencost.app before answering him. My answer ended up being “not yet, and here’s exactly why.”

What you’re actually buying

Strip away the marketing and the Agents API is a Codex harness as a service. You send it a task, it plans, calls tools, runs code, and comes back with a result — the same shape as what we already run through Claude Code and MCP servers today. The difference is where the loop executes: on OpenAI’s infrastructure instead of ours.

The sandbox pricing has four memory tiers, billed per 20-minute session:

TierPrice / 20 minPrice / hour
1 GB$0.03$0.09
4 GB$0.12$0.36
16 GB$0.48$1.44
64 GB$1.92$5.76

On top of that, tool calls are billed separately — web search runs $10 per 1,000 calls plus token cost, file search is $2.50 per 1,000 calls plus $0.10/GB/day of storage after the first gigabyte. None of that is unreasonable on its own.

The part that made me close the laptop and not sign anything

Here’s where it gets murky, and where I’d tell anyone doing capacity planning to slow down.

You can’t pick your tier. The Create Session API has no memory parameter. Nothing in the docs says which of the four tiers you land on — the working assumption floating around is that you default to the 1 GB tier, but that’s inference, not documentation. If your agent needs to unpack a large dependency tree or hold a big dataset in memory, you have no lever to pull, and no way to know in advance whether you’ll get billed at $0.03 or $1.92 per session.

“Eligible” sessions are undefined. The pricing page says “eligible container sessions will be billed by the minute, with a 5-minute minimum” — and never explains what makes a session eligible, or when the billing clock starts and stops.

Idle time is a black box. Sandboxes stay alive for up to an hour after a task finishes, and they receive keep-alives between agent turns. Whether that idle window counts toward your bill isn’t stated anywhere I could find. For a long-running agent that pauses between tool calls waiting on a human approval — which describes most of the agent workflows I’ve shipped this year — that ambiguity is not a rounding error.

If I can’t answer “what would 10,000 agent sessions cost us next month” with a spreadsheet and public documentation, I can’t put it in front of finance.

Doing the comparison math

I priced out three self-hosted alternatives for the same 20-minute session:

  • E2B, Daytona, and Blaxel cluster around $0.055 per 20-minute session — roughly what you’d pay OpenAI for the 4 GB tier, assuming you land there.
  • Cloudflare and Vercel bill by active CPU time rather than wall-clock session time, which means an agent that spends most of its 20 minutes waiting on a tool response or a human approval costs close to nothing on those platforms and full price on OpenAI’s.

So OpenAI’s sandbox is running at roughly double the going rate of the smaller self-hosted providers — and that’s before you factor in that those providers let you pick your tier and bill only for active compute. What you get in exchange is one less system to operate: no image builds, no patching, no capacity planning for your own container fleet.

The framework I actually used

I don’t think “OpenAI’s sandboxes are pricier” settles the build-vs-buy question by itself. The real question is which line item dominates your agent’s cost structure.

If your agents are short, bursty, and infrequent — a support bot that spins up a sandbox for 90 seconds a handful of times an hour — the 2x sandbox premium is a rounding error next to the engineering time you’d spend running your own fleet. Buy.

If your agents are long-running with lots of idle time between tool calls — anything with a human-in-the-loop approval step, or anything doing iterative research with pauses for rate-limited APIs — the idle-billing ambiguity could double or triple your real cost versus a CPU-metered provider. That’s the profile of most of the internal tooling agents I’ve built this year, which is why I told my teammate to hold off.

If you’re already running MCP servers and a sandbox fleet for Claude Code, the marginal cost of adding OpenAI as a second model provider inside your existing runtime is close to zero, and you keep the pricing predictability you already have. That’s usually the better move than swapping runtimes wholesale.

What I’m actually doing about it

We’re not moving off our own sandbox fleet. We are wiring the Agents API in as an alternate model backend behind the same MCP tool layer we already run, gated to workloads under two minutes where the tier ambiguity can’t hurt us much. If OpenAI publishes the missing pieces — tier selection, an eligibility definition, and idle-time billing rules — I’ll redo this math. Until then, I’m not going to sign a contract I can’t forecast.

One thing I’d flag to anyone else evaluating this: ask your OpenAI rep directly, in writing, which tier a bare session defaults to. Three different engineers I talked to this week got three different answers from sales, which tells you everything about how new this pricing model actually is.

Export for reading

Comments