Uploaded November 2025 | Updated September 2026, 3 weeks ago
Detect, Decide, Fence: Resolving Netsplits in 3-AZ Stretch Cluster - Kamoltat (Junior) Sirivadhna, IBM
Last Cephalocon, I introduced the architecture behind supporting 3 Availability Zone (AZ) stretch clusters in Ceph. One of the biggest challenges we faced in this design is handling a network partition (netsplit) between datacenters. Such events can lead to a split brain state between monitors and OSDs, preventing placement groups from peering and effectively halting all I/O operations. This talk presents a practical solution for detecting and resolving split brain in a 3-AZ stretch cluster. We'll walk through the design and implementation of a location-aware netsplit detection mechanism using monitor connection graphs, followed by maximal clique detection via the Bron–Kerbosch algorithm. From there, we’ll discuss a heuristic scoring system to select the surviving site and safely fence OSDs on the isolated side to restore availability without risking data inconsistency. Finally, we’ll examine how fencing is lifted once connectivity is re-established, ensuring minimal disruption and a smooth return to a healthy cluster state. Attendees will leave with a deeper understanding and greater confidence in managing a 3-AZ stretch cluster setup.
Detect, Decide, Fence: Resolving Netsplits in 3-AZ Stretch Cluster - Kamoltat (Junior) Sirivadhna, IBM
Last Cephalocon, I introduced the architecture behind supporting 3 Availability Zone (AZ) stretch clusters in Ceph. One of the biggest challenges we faced in this design is handling a network partition (netsplit) between datacenters. Such events can lead to a split brain state between monitors and OSDs, preventing placement groups from peering and effectively halting all I/O operations. This talk presents a practical solution for detecting and resolving split brain in a 3-AZ stretch cluster. We'll walk through the design and implementation of a location-aware netsplit detection mechanism using monitor connection graphs, followed by maximal clique detection via the Bron–Kerbosch algorithm. From there, we’ll discuss a heuristic scoring system to select the surviving site and safely fence OSDs on the isolated side to restore availability without risking data inconsistency. Finally, we’ll examine how fencing is lifted once connectivity is re-established, ensuring minimal disruption and a smooth return to a healthy cluster state. Attendees will leave with a deeper understanding and greater confidence in managing a 3-AZ stretch cluster setup.










