Retry storms historically impact business operations and brand trust. While retry configuration tuning and retry budgets provide meaningful mitigation at the service level, they’re manually configured and lack visibility into cross-service amplification caused by deep dependency chains and fan-out patterns. As a result, it can be difficult to shield infrastructure against the domino effect triggered by a single service outage deeper in the stack.
A key reason is that retry behavior today isn’t context-aware. While we can control how many retries occur, we can’t precisely control when they occur. This stems from the challenge of reliably distinguishing between errors generated by a service and those merely propagated through it.
As a result, retries are applied uniformly rather than conditionally.
This approach works for transient or low-rate failures. However, during moderate or severe degradation, it becomes counterproductive. Aggressively retrying against an already struggling service increases load, accelerates failure, and amplifies retry traffic across upstream dependencies. What begins as a localized outage can quickly escalate into a stack-wide incident—ultimately degrading, or in the worst case, completely breaking, the end user experience.
One might argue that error codes from downstream services could be translated upstream to provide context for retries. While theoretically possible, this approach doesn't scale at Uber due to large fan-in and fan-out, evolving call flows, and the need for frequent adaptive changes. Therefore, we developed a context-aware mechanism in shared infrastructure to handle errors more efficiently. This blog explains the mechanism.
Consider a simple call chain as shown in Figure 1, where the total number of requests arriving at NodeA is Ƞ. By deduction, all nodes B, C, D, E, F, and G serve Ƞ requests in the steady state (when no node errors out).
Figure 1: Call-chain with 1:1 fan-out, where a node calls its downstream exactly once for any incoming request.
DM
Deepanshu Mehndiratta
Senior Staff Engineer
Deepanshu Mehndiratta is a Senior Staff Engineer in Uber's Business Platform org, where he leads Reliability and AI Engineering. His AI work spans the MCP Gateway and Uber's frontier deep-agent ecosystem, connecting all Uber services to AI agents and leveraged by tens of thousands of employees.
AS
Alok Srivastava
Principal Engineer
Alok Srivastava is a Principal Engineer on Uber's Business Platform team. He leads Uber's Edge Platform, the ingress and egress tier for Uber's business traffic, spanning APIs, content, and push messaging across all mobile and web surfaces.
VD
Vibhor Dhingra
Sr Software Engineer
Vibhor Dhingra is a Senior Software Engineer at Uber. Previously, as part of Uber's Business Platform team, he owned Uber’s core caching proxy microservice, redesigning it to serve over 4M QPS from 1.2M QPS in 2 years.
AS
Ankit Srivastava
Distinguished Engineer
Ankit Srivastava is a Distinguished Engineer at Uber, where he works on the development of core business platforms that scale to millions of people who use Uber across the world.