Level 1 · Multiplying Machines
Session 13: Every Server Is Up, the System Is Still Down: SPOF, High Availability, & Split Brain
What if we had 10 high-speed application servers, but all of them were connected through a single device that has no twin?
1. The Core Problem: The Illusion of Partial Redundancy (SPOF)
Plenty of teams love to brag about their architecture: "We've got 10 web servers, 5 workers, and auto-scaling!"
But trace their network diagram and you'll often find: - All 10 of those web servers sit behind 1 single Load Balancer unit. - Or every server is powered by 1 rack Power Distribution Unit (PDU). - Or all traffic passes through 1 network Core Switch.
INTERNET
│
▼
┌───────────────────┐
│ SINGLE LB / SWITCH│ ◄── SINGLE POINT OF FAILURE (SPOF)!
└─────────┬─────────┘ (If this dies, EVERYTHING DIES)
┌─────────────┼─────────────┐
▼ ▼ ▼
┌───────────┐ ┌───────────┐ ┌───────────┐
│ Server 1 │ │ Server 2 │ │ Server 10 │
│ (100% UP) │ │ (100% UP) │ │ (100% UP) │
└───────────┘ └───────────┘ └───────────┘
The result: All ten application servers are 100% healthy and running, yet the system is completely down for every user in the world. This is what we call a Single Point of Failure (SPOF).
2. High Availability (HA): Eliminating the Single Point of Failure
The core principle of High Availability (HA) is simple: No component on the critical transaction path may stand alone without a redundant partner.
Two HA Redundancy Patterns:
A. Active-Passive (Master-Standby)
- How it works: The Primary node (
Active) serves 100% of the traffic. The Secondary node (Passive / Standby) waits on standby, monitoring the heartbeat. - Failover: If the Primary node dies, the Secondary node takes over the Virtual IP (VIP) address via the VRRP / Keepalived protocol.
- Pros: Very predictable, with no two-way data synchronization headaches.
- Cons: 50% of your hardware capacity sits idle (idle capacity waste).
B. Active-Active
- How it works: Both nodes (
Node 1andNode 2) are active, receiving and processing traffic at the same time. - Failover: If one dies, the remaining node carries 100% of the load.
- Pros: Maximum hardware efficiency, no machine goes to waste.
- Cons: Cluster capacity must be designed with 50% headroom, so that when 1 node goes down, the surviving node doesn't blow up too from being overloaded.
3. A Deadly Danger: Split Brain & Quorum
What happens if the network connection (the heartbeat cable link) between two redundant nodes is cut, while both nodes themselves are still alive?
┌──────────────┐ ┌──────────────┐
│ NODE A │ X [HEARTBEAT] X │ NODE B │
│ "B is down!" │ CUT │ "A is down!" │
│ I AM MASTER! │ │ I AM MASTER! │
└──────┬───────┘ └──────┬───────┘
│ │
▼ ▼
[Write to Disk] [Write to Disk]
└───► TOTAL DATA CORRUPTION! ◄──┘
The Split Brain Phenomenon:
- Node A thinks Node B is dead ──► Node A declares itself
MASTER. - Node B thinks Node A is dead ──► Node B also declares itself
MASTER. - Both accept traffic and write conflicting data to storage/database ──► Permanent data corruption (a destroyed database).
The Fix for Split Brain: Quorum & Odd Number Nodes (The Odd Rule)
Modern distributed systems (like etcd, Zookeeper, Raft, Paxos) forbid even-sized clusters (2 nodes). They require at least 3 nodes (or an odd number: 3, 5, 7).
- In a 3-node cluster: \text{Quorum} = \lfloor 3/2 \rfloor + 1 = 2.
- If a network partition happens (1 node gets isolated from the other 2):
- The majority side (2 nodes) has Quorum (2/2) ──► Keeps Serving Traffic.
- The minority side (1 node) has no Quorum (1/2) ──► Automatically Steps Down (Self-Isolate / Read-Only).
- Split brain becomes impossible!
4. Disaster Readiness Metrics: RTO and RPO
When designing HA and Disaster Recovery systems, these two numbers must be agreed on with the business right from the start:
DISASTER STRIKES
│
▼
◄────────────┼────────────────────────► TIME
RPO │ RTO
(Data Lost) │ (System Downtime)
| Metric | Full Name | Business Question | Ideal Target |
|---|---|---|---|
| RTO | Recovery Time Objective | "How long can the system be down before it comes back up?" | < 5 seconds (Auto failover) |
| RPO | Recovery Point Objective | "How much of the most recent transaction data can we afford to lose?" | 0 bytes (Zero Data Loss) |
5. The Biggest Trap: Redundancy Without Failover Testing Is Just an Assumption
Lots of companies feel safe because they have backup servers, but when a real disaster hits: 1. The automatic failover script turns out to error out because of a file permission. 2. The SSL certificate on the backup server turns out to have expired 6 months ago. 3. The backup server's capacity was never tested to handle 100% of peak load.
The Hard Law of SRE: "A failover system that has never been deliberately tested (GameDay / Chaos Engineering) is guaranteed to fail when the real disaster arrives."
6. HA Design Summary
| Component | SPOF Problem | HA Solution |
|---|---|---|
| Entry Point (VIP) | 1 Load Balancer dies | Dual LB with VRRP / Keepalived (Active-Passive / Active-Active) |
| Application Servers | Server crash / memory leak | Stateless Pool + Health Check Drain |
| Shared State (Redis) | 1 Redis dies | Redis Sentinel / Redis Cluster (3-node Quorum) |
| Database | 1 DB dies | Primary-Replica + Auto Failover (Patroni / Orchestrator) |
| Cluster Coordination | Split Brain | Raft Consensus + Odd Quorum (3 or 5 nodes) |
7. Key Terms
SPOF (Single Point of Failure): A single point that, if it breaks, takes down the entire system.HA (High Availability): A system design with no SPOF that guarantees high uptime.Active-Passive: One serves, one stands by.Active-Active: All serve at the same time, within a headroom limit.Failover: Automatically shifting the workload to a backup node.Split Brain: A condition where two nodes both claim to be Master because their communication link was cut.Quorum: The majority vote count (\lfloor N/2 \rfloor + 1) required to make a valid decision.RTO / RPO: The tolerance limit for recovery time and the tolerance limit for data loss.