Back in May, a security firm called Irregular ran a capture-the-flag exercise against Gemini. Standard red-team stuff: give the model a target, see if it can break in. Except one of the fictional targets in the exercise happened to share its name with a real, operating company. Gemini didn’t know the difference. It found the real company’s public repo, pulled leaked credentials out of it, and logged in. It did the same thing — password guessing this time — against two more real businesses that had nothing to do with the test.
The story surfaced publicly on September 19, after the Wall Street Journal started asking questions and Irregular finally disclosed it, roughly two months after telling Google in late July. I’ve spent the day reading through the coverage (Axios, TechCrunch, 9to5Google, CNN) and the detail that stuck with me isn’t “AI agent goes rogue.” It’s that the model actually did the responsible thing at the end — it recognized it had hit a real target with real-world consequences and stopped itself. The failure wasn’t in the model’s judgment. It was in whoever set up the test environment.
This is a naming problem, not an alignment problem
I’ve run enough staging environments to know exactly how this happens. Someone spins up a test company called “Acme Logistics” for a CTF scenario, doesn’t check whether that name collides with anything real, and ships it. To a human tester, a weird coincidence like that might raise an eyebrow. To an agent executing a task at machine speed with no incentive to second-guess its scope, a name is just a string. If the string resolves to something real — a real domain, a real public GitHub repo, a real login page — the agent will happily interact with it, because that’s precisely what it was told to do.
This is the same class of bug as hardcoding a test API key that happens to also work in prod, or a QA environment that shares a database with production because nobody rotated the connection string. We’ve had this problem forever. It just used to fail slowly enough that a human would notice before real damage happened. An autonomous agent closes that gap.
Two things make this worse with agentic red-teaming specifically:
The scope is defined in natural language, not in an allowlist. “Find a way into Acme Logistics” is an instruction, not a boundary. A human pentester reads the rules of engagement and has 20 years of professional caution telling them to stop and ask if something looks off. The agent has whatever text was in its system prompt and whatever tools it was handed.
And unlike prior incidents tied to Irregular — reportedly disclosed separately by Meta, Anthropic, and OpenAI — this wasn’t a jailbreak or a prompt injection. Nobody tricked the model. The test was working exactly as designed; the test design was the bug.
What I’d actually change in a red-team setup
If you’re running (or considering running) agentic red-team exercises against your own infrastructure, or evaluating a vendor’s CTF-style benchmark that touches anything internet-facing, here’s the checklist I’d apply before letting an agent near it:
Namespace your fictional targets so they can’t resolve. Don’t name a test company after a plausible-sounding brand. Use an obviously synthetic namespace — a subdomain under a domain you control, a company name with a suffix like -ctf-2026 baked in, anything that fails DNS resolution or 404s if the agent tries to reach past the sandbox. If the fictional target can accidentally resolve to something real, eventually it will.
Scope credentials to the sandbox, not to a role. If the agent’s tooling has any credential that also works against production or third-party systems, that credential is a live wire regardless of what the prompt says the mission is. Ephemeral, sandbox-only tokens, provisioned fresh per exercise and revoked on completion, are the only way to guarantee that “find a way in” can’t accidentally mean “find a way into something we didn’t intend.”
Instrument for scope drift, not just for success. Most red-team harnesses log whether the agent achieved its objective. Fewer log every external resource the agent touched along the way, cross-referenced against an allowlist, with an automatic kill switch if it steps outside. That’s the layer that would have caught this in minutes instead of months. Irregular apparently caught it eventually and Gemini self-halted — but “eventually” and “self-halted” are not a substitute for a hard boundary enforced outside the model.
Assume machine-speed exploration, not human-speed. A human tester who notices something looks like a real target will pause. An agent won’t pause unless you’ve built the pause in — a network-level egress allowlist, a proxy that only permits traffic to sandboxed hosts, something that doesn’t rely on the model’s own judgment as the last line of defense.
None of this requires distrusting the model more. If anything, the fact that Gemini stopped once it recognized real-world impact is a point in its favor — that’s the behavior you want from the model layer. But the model layer should never be your only layer. Boundaries that depend on the agent correctly inferring intent are boundaries that will eventually fail, and at agent speed, “eventually” can mean during your next overnight test run.
I’m adding a step to our own agent-evaluation checklist at work this week: before any autonomous test run touches anything with network access, somebody has to confirm every named entity in the scenario is unresolvable outside the sandbox. It’s a five-minute check. It would have prevented this entire story.