NSDI 26 - Ubers Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale... @UsenixOrg
NSDI 26 - Ubers Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale...  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
NSDI '26 - Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure

Mayank Bansal, Milind Chabbi, Kenneth Bøgh, Srikanth Prodduturi, Kevin Xu, Amit Kumar, David Bell, Ranjib Dey, Yufei Ren, Sachin Sharma, Juan Marcano, Shriniket Kale, and Subhav Pradhan, Uber Technologies; Ivan Beschastnikh, University of British Columbia; Miguel Covarrubias, Chien-Chih Liao, Sandeep Koushik Sheshadri, Wen Luo, Kai Song, Ashish Samant, Sahil Rihan, Nimish Sheth, Albert Greenberg, and Uday Kiran Medisetty, Uber Technologies

Operating a global, real-time platform at Uber’s scale requires infrastructure that is both resilient and cost-efficient. Historically, reliability was ensured through a costly 2× capacity model—each service provisioned to handle global traffic independently across two regions—leaving half the fleet idle. We present Uber’s Failover Architecture (UFA), which replaces the uniform 2× model with a differentiated architecture aligned to business criticality. Critical services retain failover guarantees, while non-critical services opportunistically use failover buffer capacity reserved for critical services during steady state. During rare “full-peak” failovers, non-critical services are selectively preempted and rapidly restored, with differentiated Service-Level Agreements (SLAs) using on-demand capacity. Automated safeguards, including dependency analysis and regression gates, ensure critical services continue to function even while non-critical services are unavailable. The quantitative impact is significant: UFA reduces steady-state provisioning from 2× to 1.3×, raising utilization from 20% toward 30% while sustaining 99.97% availability. To date, UFA has hardened over 4,000 unsafe dependencies, eliminated 575K CPU cores, and projected to reduce over one million cores from a baseline of about 4 million cores.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Ubers Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale...SREcon26 Americas - Taming the Unpredictable: Reliability in ChaosNSDI 26 - RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail BatchingNSDI 26 - KeepON: Supporting Deterministic Traffic on Standard NICsNSDI 26 - FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable SwitchesComputer Security and Voting, Invited Talk by David Dill at USENIX Security 07PEPR 26 - Toward Provably Private Insights into AI UsePEPR 26 - DPSynth: From Research to Production—Engineering Differentially Private Synthetic...NSDI 26 - FRCC: Towards Provably Fair and Robust Congestion ControlUSENIX Security 25 - ALERT: Machine Learning-Enhanced Risk Estimation for Databases Supporting...NSDI 26 - cc-pipe: Breaking Systemic Bottlenecks in RPKI Data Supply Chain with Concurrent and...NSDI 26 - Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application Management
USENIX |

NSDI '26 - Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale...

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER