UNDER PRESSURE

Level 1 · Multiplying Machines

Session 13: Every Server Is Up, the System Is Still Down: SPOF, High Availability, & Split Brain

What if we had 10 high-speed application servers, but all of them were connected through a single device that has no twin?

Session 13 / 344 min read

1. The Core Problem: The Illusion of Partial Redundancy (SPOF)

Plenty of teams love to brag about their architecture: "We've got 10 web servers, 5 workers, and auto-scaling!"

But trace their network diagram and you'll often find: - All 10 of those web servers sit behind 1 single Load Balancer unit. - Or every server is powered by 1 rack Power Distribution Unit (PDU). - Or all traffic passes through 1 network Core Switch.

                 INTERNET
                    │
                    ▼
          ┌───────────────────┐
          │ SINGLE LB / SWITCH│  ◄── SINGLE POINT OF FAILURE (SPOF)!
          └─────────┬─────────┘      (If this dies, EVERYTHING DIES)
      ┌─────────────┼─────────────┐
      ▼             ▼             ▼
┌───────────┐ ┌───────────┐ ┌───────────┐
│ Server 1  │ │ Server 2  │ │ Server 10 │
│ (100% UP) │ │ (100% UP) │ │ (100% UP) │
└───────────┘ └───────────┘ └───────────┘

The result: All ten application servers are 100% healthy and running, yet the system is completely down for every user in the world. This is what we call a Single Point of Failure (SPOF).


2. High Availability (HA): Eliminating the Single Point of Failure

The core principle of High Availability (HA) is simple: No component on the critical transaction path may stand alone without a redundant partner.

Two HA Redundancy Patterns:

A. Active-Passive (Master-Standby)

  • How it works: The Primary node (Active) serves 100% of the traffic. The Secondary node (Passive / Standby) waits on standby, monitoring the heartbeat.
  • Failover: If the Primary node dies, the Secondary node takes over the Virtual IP (VIP) address via the VRRP / Keepalived protocol.
  • Pros: Very predictable, with no two-way data synchronization headaches.
  • Cons: 50% of your hardware capacity sits idle (idle capacity waste).

B. Active-Active

  • How it works: Both nodes (Node 1 and Node 2) are active, receiving and processing traffic at the same time.
  • Failover: If one dies, the remaining node carries 100% of the load.
  • Pros: Maximum hardware efficiency, no machine goes to waste.
  • Cons: Cluster capacity must be designed with 50% headroom, so that when 1 node goes down, the surviving node doesn't blow up too from being overloaded.

3. A Deadly Danger: Split Brain & Quorum

What happens if the network connection (the heartbeat cable link) between two redundant nodes is cut, while both nodes themselves are still alive?

┌──────────────┐                 ┌──────────────┐
│    NODE A    │  X [HEARTBEAT] X │    NODE B    │
│ "B is down!" │      CUT        │ "A is down!" │
│ I AM MASTER! │                 │ I AM MASTER! │
└──────┬───────┘                 └──────┬───────┘
       │                                │
       ▼                                ▼
[Write to Disk]                  [Write to Disk]
       └───► TOTAL DATA CORRUPTION! ◄──┘

The Split Brain Phenomenon:

  • Node A thinks Node B is dead ──► Node A declares itself MASTER.
  • Node B thinks Node A is dead ──► Node B also declares itself MASTER.
  • Both accept traffic and write conflicting data to storage/database ──► Permanent data corruption (a destroyed database).

The Fix for Split Brain: Quorum & Odd Number Nodes (The Odd Rule)

Modern distributed systems (like etcd, Zookeeper, Raft, Paxos) forbid even-sized clusters (2 nodes). They require at least 3 nodes (or an odd number: 3, 5, 7).

\text{Quorum} = \left\lfloor \frac{N}{2} \right\rfloor + 1
  • In a 3-node cluster: \text{Quorum} = \lfloor 3/2 \rfloor + 1 = 2.
  • If a network partition happens (1 node gets isolated from the other 2):
  • The majority side (2 nodes) has Quorum (2/2) ──► Keeps Serving Traffic.
  • The minority side (1 node) has no Quorum (1/2) ──► Automatically Steps Down (Self-Isolate / Read-Only).
  • Split brain becomes impossible!

4. Disaster Readiness Metrics: RTO and RPO

When designing HA and Disaster Recovery systems, these two numbers must be agreed on with the business right from the start:

        DISASTER STRIKES
             │
             ▼
◄────────────┼────────────────────────► TIME
      RPO    │           RTO
(Data Lost)  │ (System Downtime)
Metric Full Name Business Question Ideal Target
RTO Recovery Time Objective "How long can the system be down before it comes back up?" < 5 seconds (Auto failover)
RPO Recovery Point Objective "How much of the most recent transaction data can we afford to lose?" 0 bytes (Zero Data Loss)

5. The Biggest Trap: Redundancy Without Failover Testing Is Just an Assumption

Lots of companies feel safe because they have backup servers, but when a real disaster hits: 1. The automatic failover script turns out to error out because of a file permission. 2. The SSL certificate on the backup server turns out to have expired 6 months ago. 3. The backup server's capacity was never tested to handle 100% of peak load.

The Hard Law of SRE: "A failover system that has never been deliberately tested (GameDay / Chaos Engineering) is guaranteed to fail when the real disaster arrives."


6. HA Design Summary

Component SPOF Problem HA Solution
Entry Point (VIP) 1 Load Balancer dies Dual LB with VRRP / Keepalived (Active-Passive / Active-Active)
Application Servers Server crash / memory leak Stateless Pool + Health Check Drain
Shared State (Redis) 1 Redis dies Redis Sentinel / Redis Cluster (3-node Quorum)
Database 1 DB dies Primary-Replica + Auto Failover (Patroni / Orchestrator)
Cluster Coordination Split Brain Raft Consensus + Odd Quorum (3 or 5 nodes)

7. Key Terms

  • SPOF (Single Point of Failure): A single point that, if it breaks, takes down the entire system.
  • HA (High Availability): A system design with no SPOF that guarantees high uptime.
  • Active-Passive: One serves, one stands by.
  • Active-Active: All serve at the same time, within a headroom limit.
  • Failover: Automatically shifting the workload to a backup node.
  • Split Brain: A condition where two nodes both claim to be Master because their communication link was cut.
  • Quorum: The majority vote count (\lfloor N/2 \rfloor + 1) required to make a valid decision.
  • RTO / RPO: The tolerance limit for recovery time and the tolerance limit for data loss.