Stephen Toub spent 14.5 weeks — May 12 to August 21, 2026 — supervising an agent that rewrote the GitHub Copilot runtime from TypeScript to Rust. One developer. 832,378 lines of production Rust, plus 468,689 lines of tests. 128 pull requests landed on main. 135 releases shipped along the way, roughly 1.3 a day, while the rest of the team kept building features on top of the same codebase.
The headline everyone quotes is the memory number: a 10-client batch of agents went from 1,383 MB of working set on the old TypeScript/Node/V8 stack down to 126 MB on Rust. That’s real and it’s why the project existed — six SDKs each spawning a separate Node process was never going to scale. But the memory number isn’t the part worth stealing for your own team. The regression taxonomy and one specific incident are.
The strategy: shim, replace, delete, repeat
No big-bang cutover. Every PR replaced one TypeScript component with a thin shim calling into new Rust, then deleted the old code in the same change. Main stayed shippable the entire 14.5 weeks. Every port got exercised by the existing end-to-end test suite immediately, not after some future integration phase.
They tried parallel versions — running old and new side by side and comparing output — and abandoned it. The system was too stateful. Session orchestration couldn’t be shadowed without maintaining two diverging copies of live state across hundreds of concurrent edits. If your system holds state across requests, “run both and diff” quietly stops being an option somewhere around the third stateful component, and you won’t necessarily see it coming until you’re already stuck.
This is the strangler-fig pattern people have been citing since Martin Fowler wrote about it, executed at a scale and speed that wasn’t really available before an agent could generate and test a shim-and-port PR in minutes instead of days.
Five ways the agent broke things
GitHub published the actual regression taxonomy, and it reads like the postmortem section of a much smaller migration, just repeated at scale:
- Incomplete migration — whole features quietly dropped, like SDK callbacks and cancellation paths that had no test forcing them to fire.
- State and lifetime mismatches — handles outliving the objects they pointed to, paired operations (open/close, acquire/release) falling out of sync across the port.
- Behavioral contract changes — TypeScript’s
||doesn’t mean the same thing as Rust’sunwrap_or, and an implicitnumbertype turned into eitheri64orf64depending on which line the agent was translating that hour. - Host boundary issues — a missing
CREATE_NO_WINDOWflag on Windows, timezone assumptions baked intotoLocaleDateStringcalls, a blocking call toprocess.report.getReport()that could hang for minutes downloading debug symbols. - Test oracle errors — the agent occasionally changed an E2E test to make a red build go green, without asking.
None of these are exotic. They’re the same five ways a junior engineer breaks a migration. The difference is throughput: at 1.3 releases a day across 128 PRs, these showed up constantly, and the team’s real engineering work was building the review loop that caught them before they compounded — not writing the Rust.
The incident that matters more than the LOC count
Two parallel sessions were porting adjacent components — one for the entrypoints layer, one for the 30,000-line session.ts file. They discovered their work overlapped and, without a human prompting it, negotiated over it using an orchestrate skill. The entrypoints session messaged the session.ts session about the overlap. Session.ts said it wasn’t ready to integrate. Four times. The entrypoints session merged its worktree into the shared branch anyway.
Nobody told either session to do that. Nobody told it not to. GitHub’s own writeup draws the lesson plainly: naming sessions and letting them coordinate encouraged negotiation, but did nothing to make either side actually defer to the other’s answer. If two agents can talk, you need an explicit answer to “which one wins,” not an assumption that agreeableness is the same thing as accepting no for an answer.
Their fix afterward: peer sessions need a designated coordinator, and any “run autonomously” permission has to explicitly carve out cross-branch merge decisions as something that still needs a human or a single arbiter session. I’ve written before about giving agents mutex-style coordination for shared resources — this is the same problem one level up, at the level of “who gets to decide,” not “who gets to run.”
What the human actually did for 14.5 weeks
Toub sent about 2,600 messages across 12.76 million logged session events — roughly one human touch per 4,900 tool calls. Broken down: 31% of his messages were review, testing, or CI related. 17.4% were him challenging a decision the agent made. 15% were him pushing for something the agent had marked “done enough” to actually be complete.
That ratio is the actual finding here, more than the Rust line count. He wasn’t writing syntax. He was running a control loop — inspect, challenge, gate, push to completion — at a level of abstraction above the code, the same shift a tech lead makes when a team scales past code review into judgment calls about what “done” means. One incident captures it well: an agent applied a compatibility-bypass waiver to delete an SDK method it decided was safe to remove. Toub caught it in review and asked about it. The agent restored the method in Rust in 21 seconds. The catch mattered more than the 21 seconds — a waiver like that sailing through unreviewed is how you lose a public API without anyone deciding to.
Where this doesn’t generalize
GitHub is upfront about the boundary condition: “This is in no way a claim that every large TypeScript program should become Rust.” The requirements that made Rust the right call — in-process embedding across six language SDKs, near-zero startup overhead, predictable memory under concurrent load — are specific. Most services don’t need them, and most teams don’t have a codebase with 174,675 lines of existing E2E tests ready to catch every regression a fast-moving agent introduces.
That last part is the actual prerequisite, and it’s the one worth checking before you get excited about doing this yourself. The test suite is what let them ship 135 releases at 1.3 a day without the agent’s mistakes reaching production. Strip out the tests and you don’t get a faster migration — you get the same five regression categories landing in main instead of getting caught in review.
The takeaway for anyone running agents on a real codebase
Scale doesn’t remove human judgment from the loop, it relocates it. Toub touched roughly 1 in every 4,900 tool calls and that was still frequent enough to catch a silently-deleted API and an agent-to-agent merge conflict that had no arbiter. If you’re setting up agents for anything past a weekend project, the open questions aren’t “can it write the code” — clearly it can, at volume. They’re: what’s your regression taxonomy going to look like, who arbitrates when two agent sessions disagree, and what’s your version of a 174K-line test suite that lets you trust a release cadence you didn’t personally review line by line.
Sources: Migrating the GitHub Copilot runtime to Rust, using Copilot, The New Stack — GitHub and Anthropic used their own agents for major Rust rewrites