I’ve been paged for a retry storm exactly once, at 2 a.m., for a service that had nothing wrong with it except being twenty hops downstream of one that did. Every layer in between saw a timeout, assumed it owned the failure, and retried. The actual broken service got hit with 20x its normal load from a single degraded dependency. Uber just published the numbers on the version of this problem they had, and the fix is smaller than I expected.

Retries assume the wrong thing by default

Standard retry logic treats every error the same way: something failed, try again, maybe with backoff. What it doesn’t know is whether the error originated at this hop or arrived from three services upstream, already retried twice by someone else. In a deep call chain, that ignorance compounds. Service A calls B calls C calls D. D degrades. C sees a failure and retries. B sees C’s failure (or C’s retry-induced latency) and retries. A does the same. One degraded leaf node now generates traffic multiplied by the depth of every layer that doesn’t know it isn’t the origin.

Uber’s engineering team — Deepanshu Mehndiratta, Alok Srivastava, Vibhor Dhingra, Ankit Srivastava — call this “error ownership,” and the framing is the actual insight, not the implementation. An error belongs to whichever service generated it. Every other hop that merely observed it while forwarding a request is not the owner, and shouldn’t be spending its own retry budget on it.

The mechanism is one header, not a new subsystem

They added x-uber-error-claim to their retry middleware. When a service generates an error, it claims it. When a downstream error passes through a service unchanged, that service doesn’t claim it — it just forwards the claim that already exists. Only the service that owns the error, per its own retry budget and backoff policy, gets to decide whether a retry happens. Everyone upstream of it inherits that decision instead of independently reinventing it.

That’s the entire mechanism. No new consensus protocol, no additional round trip, no service mesh redesign — a header carrying ownership metadata through the existing call chain, checked before a retry decision is made instead of after.

The numbers made me actually believe it

During a real Core Entity service degradation in November 2025, this blocked 9.5 million spurious retry requests. Uber’s own estimate is that without error ownership, the traffic spike on the already-degraded service would have been 46 to 135 percent higher — on top of a service that was already failing.

The blast-radius numbers are the part I keep rereading. Retry-storm depth went from 25 layers to 3. Average retry depth across user-facing APIs dropped from 20 to 2. That’s not a marginal tuning win. That’s the difference between “one bad dependency degrades the whole call graph” and “one bad dependency degrades the three services closest to it.”

I’d take the 9.5 million number with a grain of salt as a generalizable constant — it’s specific to Uber’s call-graph depth and their retry defaults before the fix, and a shallower architecture won’t see the same multiplier. But the mechanism generalizes cleanly regardless of scale, and the depth-25-to-3 number is the one that actually tells you why.

Where I’d wire this in

If you’re running a service mesh or even a handful of services with retry logic layered at every hop, the pattern is portable without needing Uber’s infrastructure. The header doesn’t have to be Uber’s header — it has to encode two things: who generated this error, and has a retry already been attempted for it. A gRPC interceptor or an HTTP client middleware can carry that metadata just as well as a custom header, as long as every hop in your chain respects it instead of making an independent retry decision from its own local view.

The part I’d get wrong if I implemented this carelessly: trusting the claim without validating it. A malicious or buggy service could claim ownership of an error it didn’t generate, suppressing legitimate retries elsewhere in the chain. Uber’s writeup doesn’t go deep on how they guard against that internally, and it’s the first thing I’d want answered before shipping this pattern into a system with less trust between services than Uber’s internal mesh has.

What actually changed my mind reading this isn’t the specific header format. It’s that retry logic almost everywhere I’ve worked treats “did this call fail” as the only signal that matters, when “who actually owns this failure” is the one that determines whether retrying helps or just adds load to something already on fire.

Export for reading

Comments