On November 18, 2025, a Core Entity service more than five hops deep inside Uber’s call graph started failing from an infrastructure issue. Nothing dramatic at first — just elevated errors on one node, buried under layers of callers who had no idea it was the actual source. With the retry policy Uber had running until recently, that would have meant a 46%-135% traffic spike hitting the already-degraded service, from every layer above it retrying in parallel, each unaware the others were retrying too. Instead, Uber’s engineering team reports they stopped 9.5 million spurious requests across the mesh, because a pattern they call error ownership had shipped months earlier. The immediate callers of the failing service were blocked from sending up to 200,000 additional requests each.
I’ve debugged this exact failure mode without a name for it. A service four layers up from the actual problem, retrying a request that was doomed the moment it left the building, making the outage worse instead of riding it out. Uber’s writeup is the first time I’ve seen someone put real math on why it happens and a real fix that isn’t “tune your backoff harder.”
The bug: every hop treats every error as its own
Standard retry middleware doesn’t know the difference between “I failed” and “someone I called failed.” It just sees an error and retries, at every layer, uniformly. In a chain A → B → C → D, if D fails, C retries, B retries, A retries — three retry attempts stacked on top of one real failure, and that’s a shallow chain. Uber’s own retry-amplification formula makes the shape of the problem obvious:
Requests at depth d, no retry budget: R^d × Ƞ
Requests at depth d, retry budget B: (1+B)^d × Ƞ
Where R is retries per hop and Ƞ is the baseline request rate. Run their own six-hop example: one retry per hop, node D fails at depth 3. Without any budget, nodes D through G each see 8Ƞ requests from a single upstream failure. With a 10% retry budget instead of a blind retry, that drops to 1.33Ƞ. With error ownership restricting retries to just the edge that actually failed — C calling D — nodes A, B, and C see no extra traffic at all, and D through G see only 1.1Ƞ.
That’s the entire pitch in one table: a retry budget alone helps, but knowing exactly which edge owns the failure and retrying only there is an order of magnitude better, and it’s the difference between a self-healing system and a pile-on.
The fix: a service claims its own errors
Uber’s rule, straight from the post: if a service calls out to N downstreams while handling a request, and one of those downstream calls fails, and that’s why the service returns an error — it’s not the owner. It’s a symptom, and the error propagates upstream unclaimed. But if none of its outbound calls failed and it still errors out, that service is the cause, and it claims the error before returning it.
The claim gets carried on the error itself — Uber’s diagrams show it as a header their retry middleware reads before deciding whether to retry. Claimed errors coming back from a direct downstream are eligible for a retry at that hop, because you know retrying might actually help. Unclaimed errors mean the real failure is further down the chain, retrying here just adds load to an already-struggling service, so the middleware doesn’t.
Here’s the shape of it as middleware, stripped down to the logic that matters — this is not Uber’s actual code, just the pattern applied to a typical Node/Express service mesh:
async function callDownstream(req, res, next) {
const outboundResults = await Promise.allSettled(
req.dependencies.map(dep => dep.call(req))
);
const anyDownstreamFailed = outboundResults.some(r => r.status === 'rejected');
try {
const result = await handleRequest(req, outboundResults);
return res.json(result);
} catch (err) {
// I failed, but not because a downstream failed — I own this.
if (!anyDownstreamFailed) {
res.set('x-error-claim', 'claimed');
} else {
// A downstream failed and that's why I'm failing — pass it on unclaimed.
res.set('x-error-claim', 'unclaimed');
}
return res.status(err.status || 500).json({ error: err.message });
}
}
function shouldRetry(response) {
return response.headers['x-error-claim'] === 'claimed';
}
The production result, aggregated across all of Uber’s user-facing APIs, not just the one incident: max retry-storm radius — the depth a retry storm can travel before error ownership stops it — dropped from 25 hops to 3. The average dropped from 20 to 2.
Where this actually matters, and where it doesn’t
If your service graph is two or three hops deep, skip this. The math only bites once you have enough layers for one failure to get retried, and re-retried, and re-retried again by callers who have no visibility into each other. If you’re running anything past four or five hops of internal service calls — which describes most microservice estates I’ve worked in past a certain company size — this is worth an afternoon.
One honest caveat Uber’s own post raises: ownership attribution isn’t perfect. Their worst-case estimate is that around 2% of the time, a caller that also genuinely failed on its own gets its error incorrectly marked unclaimed, because a downstream call happened to fail around the same time by coincidence. They handle this by feeding historical failure patterns back into the decision, not a single point-in-time check. If you build this yourself, budget for that same false-negative case rather than assuming a clean binary signal — a coincidental downstream blip shouldn’t silently mask a real bug in your own service.
The part I’d steal first, even before the full ownership machinery: the retry-budget math above. (1+B)^d versus R^d is a one-line config change in most retry libraries, and it alone buys you most of the safety margin before you’ve written a line of ownership logic.