DoorDash’s experimentation platform runs more than 60,000 feature flags across roughly 623 repositories, adding about 2,300 new ones every month. Nobody deletes flags on purpose. They just accumulate until someone finally audits them, usually right before a migration or an incident review turns up a flag that’s been at 100% rollout for a year and a half but is still sitting in every call site as an if-statement. I’ve cleaned up exactly this kind of debt by hand. It’s tedious, low-risk-but-not-zero-risk work, and it’s exactly the kind of task everyone agrees should be automated and nobody prioritizes building the automation for.
DoorDash built it anyway, and — this is the part I actually care about — they published the cost and success-rate numbers instead of just a launch post. 50 stale flags, evaluated end to end: 45 produced a usable pull request, averaging 13.8 minutes and $4.79 each. That’s the first time I’ve seen a real dollar figure attached to “AI agent does a whole engineering task,” not a demo, and it’s worth taking apart.
The architecture is two models doing two different jobs
The system is an orchestrator plus worker pattern. A Claude Sonnet orchestrator reads Jira tickets and pulls flag metadata — cheap, high-volume, not much reasoning required. Once a flag is scoped, a Claude Opus agent does the actual removal: it works inside an isolated git worktree (DoorDash runs up to four concurrently per repo, Gradle with --no-daemon to avoid state bleeding across worktrees), traces the call chain, deletes the dead branches, and opens a PR. Everything gets a one-hour hard timeout.
That’s a cost-tiering decision, not a capability one — Sonnet could probably do the metadata pass fine on its own, but there’s no reason to pay Opus rates for reading a Jira ticket. I’ve made the same call in our own pipelines: route by task complexity, not by defaulting everything to the most capable model available. The numbers back it up. Simple flags: 100% success, 7.5 minutes, $2.69. Complex flags with deep call chains: 85% success, 17.7 minutes, $6.20. The cost curve tracks reasoning depth almost linearly, which is a nice property if you’re trying to forecast a budget for this kind of work at scale.
The part that actually makes this safe to run unattended
The number that matters more than $4.79 is 95 — the minimum JaCoCo patch coverage required before a PR gets opened, on top of passing build, passing tests, and Detekt static analysis. Four deterministic gates between “agent thinks it’s done” and “a human sees a PR.” Of the 14 revisions needed across the 50-flag sample, 6 were coverage failures and 8 were incomplete removals — both categories of bug the gates were built to catch, and did.
The failure mode worth noting: when the system gets it wrong, it’s conservative-wrong, not dangerous-wrong. DoorDash’s own writeup says the agent is “more likely to leave incomplete dead-code removal in complex call chains than to make an unsafe semantic change” — leftover unused variables deep in a call chain, not a flag flipped to the wrong default in production. That’s the failure mode you want from an automated system touching live conditional logic. An agent that occasionally leaves dead code behind is an annoyance. An agent that occasionally inverts a rollout percentage is an incident.
The other detail that explains a lot of the 90% success rate: the removal agent doesn’t just read the repo, it queries live experimentation data over MCP — actual rollout percentages and target values, not whatever the code’s default branch implies. That’s the difference between “the flag looks like it’s at 100% based on the code I can see” and “the flag is actually at 100% right now.” Inferring flag state purely from source is exactly how you’d silently change production behavior while believing you’re just deleting dead code. Giving the agent a live data source instead of a static snapshot is a small architectural choice that removes an entire class of risk.
Where I’d actually use this pattern
The Sonnet-orchestrator-plus-Opus-worker split, gated by deterministic checks before anything reaches a human, generalizes well past feature flags — dependency upgrades, deprecated-API migrations, dead-code removal after a product sunset, anything where the task is “mechanical but requires call-chain awareness, and a wrong answer is worse than a slow one.” The template isn’t the LLM. It’s cheap-model-scopes-the-work, expensive-model-executes-the-work, and neither model’s output reaches main without passing tests, coverage, and static analysis first.
What I wouldn’t do is skip the coverage gate to save the 6-8% of PRs it catches. $4.79 already assumes the gate exists — take it away and you’re not saving cost, you’re moving the cost to whoever finds the coverage regression three sprints later.