Back Original

How Uber Protects Against Retry Storms

Retry storms historically impact business operations and brand trust. While retry configuration tuning and retry budgets provide meaningful mitigation at the service level, they’re manually configured and lack visibility into cross-service amplification caused by deep dependency chains and fan-out patterns. As a result, it can be difficult to shield infrastructure against the domino effect triggered by a single service outage deeper in the stack. 

A key reason is that retry behavior today isn’t context-aware. While we can control how many retries occur, we can’t precisely control when they occur. This stems from the challenge of reliably distinguishing between errors generated by a service and those merely propagated through it.

As a result, retries are applied uniformly rather than conditionally.

This approach works for transient or low-rate failures. However, during moderate or severe degradation, it becomes counterproductive. Aggressively retrying against an already struggling service increases load, accelerates failure, and amplifies retry traffic across upstream dependencies. What begins as a localized outage can quickly escalate into a stack-wide incident—ultimately degrading, or in the worst case, completely breaking, the end user experience.

One might argue that error codes from downstream services could be translated upstream to provide context for retries. While theoretically possible, this approach doesn't scale at Uber due to large fan-in and fan-out, evolving call flows, and the need for frequent adaptive changes. Therefore, we developed a context-aware mechanism in shared infrastructure to handle errors more efficiently. This blog explains the mechanism.

Background

Consider a simple call chain as shown in Figure 1, where the total number of requests arriving at NodeA is Ƞ. By deduction, all nodes B, C, D, E, F, and G serve Ƞ requests in the steady state (when no node errors out).

Seven blue circles labeled A to G connected by rightward arrows in a straight horizontal line.

Figure 1: Call-chain with 1:1 fan-out, where a node calls its downstream exactly once for any incoming request.

Deepanshu Mehndiratta

DM

Deepanshu Mehndiratta

Senior Staff Engineer

Deepanshu Mehndiratta is a Senior Staff Engineer in Uber's Business Platform org, where he leads Reliability and AI Engineering. His AI work spans the MCP Gateway and Uber's frontier deep-agent ecosystem, connecting all Uber services to AI agents and leveraged by tens of thousands of employees.

Alok Srivastava

AS

Alok Srivastava

Principal Engineer

Alok Srivastava is a Principal Engineer on Uber's Business Platform team. He leads Uber's Edge Platform, the ingress and egress tier for Uber's business traffic, spanning APIs, content, and push messaging across all mobile and web surfaces.

Vibhor Dhingra

VD

Vibhor Dhingra

Sr Software Engineer

Vibhor Dhingra is a Senior Software Engineer at Uber. Previously, as part of Uber's Business Platform team, he owned Uber’s core caching proxy microservice, redesigning it to serve over 4M QPS from 1.2M QPS in 2 years.

Ankit Srivastava

AS

Ankit Srivastava

Distinguished Engineer

Ankit Srivastava is a Distinguished Engineer at Uber, where he works on the development of core business platforms that scale to millions of people who use Uber across the world.