Implementing Decentralized Per-Partition Automatic Failover in Azure Cosmos DB

๐Ÿ“… 2025-05-20
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Azure Cosmos DB struggles to simultaneously achieve fine-grained recovery, low recovery time objective (RTO) and recovery point objective (RPO), and strong consistency under node- to region-level failures. Method: This paper proposes the first partition-level, decentralized cross-region automatic failover architecture. It leverages distributed consensus protocols and partition-granular failure detection and traffic rerouting to fully decentralize metadata coordination and state-machine fault tolerance. Clients can flexibly configure per-partition consistency levels and RPO/RTO targets. Contribution/Results: Experiments demonstrate millisecond-scale RTO for critical partitions, optional RPO = 0 (zero data loss), and robust self-healing across full operational scenarios at scaleโ€”supporting over 20 million vCores and 100+ PB of data. This work breaks the conventional region-level disaster recovery paradigm, establishing a new high-availability framework for hyperscale distributed databases.

Technology Category

Search and Optimization: Distributed SearchConstraint Satisfaction and Optimization: Distributed CSP/OptimizationData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Decentralized Web and Fediverse systemsResponsible Web: Data and user privacy-enhancing technologies for the WebSecurity and Privacy: Blockchains and distributed ledgers
๐Ÿ“ Abstract
Azure Cosmos DB is a cloud-native distributed database, operating at a massive scale, powering Microsoft Cloud. Think 10s of millions of database partitions (replica-sets), 100+ PBs of data under management, 20M+ vCores. Failovers are an integral part of distributed databases to provide data availability during outages (partial or full regional outages). While failovers within a replica-set within a single region are well understood and commonly exercised, geo failovers in databases across regions are not as common and usually left as a disaster recovery scenario. An upcoming release of Azure Cosmos DB introduces a fine grained (partition-level) automatic failover solution for geo failovers that minimizes the Recovery Time Objective (RTO) and honors customer-chosen consistency level and Recovery Point Objective (RPO) at any scale. This is achieved thanks to a decentralized architecture which offers seamless horizontal scaling to allow us to handle outages ranging from node-level faults to full-scale regional outages. Our solution is designed to handle a broad spectrum of hardware and software faults, including node failures, crashes, power events and most network partitions, that span beyond the scope of a single fault domain or an availability zone.
Problem

Research questions and friction points this paper is trying to address.

Implementing partition-level automatic geo failover in Azure Cosmos DB
Minimizing Recovery Time Objective while maintaining consistency
Handling diverse faults from node failures to regional outages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decentralized per-partition automatic failover solution
Minimizes Recovery Time Objective (RTO)
Handles node to regional outages seamlessly
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
J
Josh Rowe
Microsoft
M
Mikael Horal
Microsoft
H
Hari Sudan Sundar
Microsoft
M
Muthukumaran Arumugam
Microsoft
B
Burak Kose
Microsoft
S
Sravani Mitra Palivela
Microsoft
G
Geni Marsh
Microsoft
V
Varun Jain
Microsoft
A
Abhishek Kumar
Microsoft
D
Dhaval Patel
Microsoft