Three months ago I had a fight with a security reviewer over Cursor’s Cloud Agents. Not a real fight — Slack messages, mostly polite — but the sticking point was real: our client’s compliance policy required all code execution, including AI agent sandboxes, to stay inside their VPC. Cursor’s cloud agents ran on Cursor’s machines. Full stop. We ended up disabling the feature for that client and telling the team to use local agents only, which is a worse experience and everyone knew it.

On September 3, 2026, Cursor shipped the fix, six months later than I wanted it, in the form of a Vercel changelog entry most people probably skimmed past: “Cursor Cloud Agents can now run in Vercel Sandbox.” Read that as a one-liner and it sounds like a minor integration. Read the architecture underneath it and it’s Cursor telling you, out loud, which part of the product it actually owns.

What shipped

The core move is called Self-Hosted Machines. Cursor keeps the agent harness — the loop that decides what to read, what to edit, when to run tests, when to stop — and lets you supply the execution environment where that loop actually does its work: cloning the repo, writing files, running the test suite. That execution environment used to be Cursor’s own fleet. Now it can be one of eight backends: Vercel, Cloudflare, Daytona, E2B, Modal, Coder, Namespace, or Lambda, all routed through one worker pool.

Vercel’s implementation is the one I looked at closely because we already run infra there. Each agent request gets its own isolated Firecracker microVM, provisioned on demand. Vercel Functions and Vercel Workflows act as the control plane: they claim a queued agent job, spin up a worker, watch the session, and tear it down when it’s done. No idle fleet of long-lived VMs sitting around waiting for someone to kick off a task. Scale-to-zero, dedicated isolation per request, short-lived credentials scoped to that one job.

This is not a new pattern if you’ve built serverless CI runners before. It’s the same shape as GitHub Actions’ ephemeral runners, or how CodeBuild spins up throwaway containers. What’s new is applying it to an agent loop that can run for twenty minutes, touch forty files, and needs durable retries if the underlying worker dies mid-task — because unlike a CI job, you don’t want to just fail and restart from zero, you want to resume the agent’s actual progress.

Why this matters more than it looks

Here’s the part that took me a second read to appreciate: Self-Hosted Machines requires a Cursor Enterprise plan. That’s the tell. Cursor isn’t giving away infrastructure flexibility out of generosity — it’s drawing a line between “the agent brain, which we sell you” and “the compute it runs on, which we no longer insist on selling you.” That’s a company being honest about its actual defensibility. The value isn’t the sandbox. Anyone can spin up a Firecracker VM. The value is the harness: the part that decides what “done” means for a coding task, retries intelligently, and doesn’t burn your token budget re-reading files it already saw.

Compare this to how OpenAI structured the Agents API a week earlier (September 10) with nine sandbox partners of its own — Vercel, Cloudflare, Daytona, E2B, Modal, DigitalOcean, Oracle, Runloop, Blaxel. Different company, near-identical list, same underlying admission. Both are converging on: harness is the product, sandbox is a commodity you plug in. If two competitors independently land on the same architecture within a week of each other, that’s not coincidence, that’s the shape of the problem.

What I’d actually check before adopting this

I ran a quick trial deploying the Vercel Sandbox reference implementation to a personal account, mostly to see what the failure modes look like before I’d trust it with client work:

# Vercel's reference deployment for the Cursor sandbox worker
git clone https://github.com/vercel/cursor-sandbox-worker
cd cursor-sandbox-worker
vercel link
vercel env pull .env.local
vercel deploy --prod

The deploy itself took about ninety seconds. Wiring it into Cursor’s Self-Hosted Machines settings and pointing a test agent at it worked on the first try, which honestly surprised me — most “bring your own infra” integrations don’t survive first contact.

What I couldn’t verify from the docs, and what I’d push hard on before putting this in front of a compliance-sensitive client again: exactly what telemetry crosses back from the sandbox to Cursor’s harness. The execution happens in your VPC-adjacent environment, sure, but the harness making decisions still lives with Cursor, and decisions require context — meaning file contents, diffs, command output. “Your infrastructure” doesn’t automatically mean “your data never leaves.” That’s a question for Cursor’s enterprise sales team, not something you can infer from a changelog post, and I’d want it in writing before I re-enabled Cloud Agents for that client.

The actual lesson for anyone building agent tooling

If you’re designing agent infrastructure internally — and a few teams I talk to are, because not everyone wants to hand a coding agent to a third party — the lesson isn’t “use Firecracker” or “copy Cursor’s architecture.” It’s that the two things people conflate, the reasoning loop and the execution sandbox, are separable, and separating them is what lets you sell to the customer who has a hard requirement you can’t talk them out of. We spent two months building a workaround for a limitation that turned out to be a product decision, not a technical one. Cursor changed the decision. The technical work to support it — Firecracker isolation, ephemeral credentials, a durable control plane — was already sitting there waiting to be plugged in, mostly because Vercel had already built Sandbox for other reasons entirely.

Next time a vendor tells you “that’s just how the architecture works,” ask whether it’s actually load-bearing, or whether it’s a business decision wearing a technical excuse.

Export for reading

Comments