UNDER PRESSURE

Level 1 · Multiplying Machines

Session 9: If There Are Two Servers, Who Decides? (Load Balancer & HAProxy)

Imagine if there were two identical service counters, and one attendant stood at the entrance directing visitors. How would that attendant choose where each person should go?

Session 9 / 344 min read

1. The Core Problem: The Illusion of Two Servers

When you grow from 1 server to 2 (Server A and Server B), a new problem shows up at the front door: - Users on the internet only know 1 domain (for example api.tokomu.com). - That domain points to only 1 public IP address. - So who decides which request goes to Server A and which goes to Server B?

That's the job of a Load Balancer (Reverse Dispatcher / Traffic Director) such as HAProxy, NGINX, or AWS ALB.


2. Four Selection Strategies (Load Balancing Algorithms)

The attendant (the Load Balancer) has a few logical strategies to choose from:

A. Round Robin (Plain Alternation)

  • Logic: Request 1 to A, Request 2 to B, Request 3 to A, Request 4 to B.
  • When It Fits: Every request carries exactly the same amount of work and every server has identical capacity.
  • When It Goes Wrong (The Trap): Say Server A gets a heavy request (generating a 10-second PDF report) while Server B gets a light one (a 10 ms status check). Round robin will keep sending new requests to Server A anyway, leaving Server A overloaded.

B. Weighted Round Robin (Weighted Load)

  • Logic: Server A (8 cores) gets weight 2, Server B (4 cores) gets weight 1. Server A receives 2 requests for every 1 request that goes to Server B.
  • When It Fits: The servers in the cluster have different hardware specs (a heterogeneous fleet).

C. Least Connections (Fewest Connections)

  • Logic: Send the next request to whichever server is currently handling the fewest active connections.
  • When It Fits: Request durations vary a lot (some take 10 ms, some take 5 seconds, some are streaming/WebSocket).
  • Result: It keeps load from piling up on a single server that's stuck on a heavy query.

D. IP Hash / Source Hashing

  • Logic: hash(IP_Klien) % Jumlah_Server. A client with the same IP always lands on the same server.
  • When It Fits: A stopgap when the application isn't stateless yet (a simple sticky session).
  • The Trap: If 10,000 users come in through the same office proxy / NAT, all of them get thrown at the same single server (hotspot imbalance).

3. Health Checks: The Load Balancer's Eyes and Ears

A load balancer must never send traffic to a server that's already dead or dying.

                  ┌──────────────┐
                  │ Load Balancer│
                  └──────┬───────┘
          Health Check   │   Health Check
         GET /healthz    │   GET /healthz
               ┌─────────┴─────────┐
               ▼                   ▼
        ┌─────────────┐     ┌─────────────┐
        │  Server A   │     │  Server B   │
        │ HTTP 200 OK │     │ HTTP 500 /  │
        │  (HEALTHY)  │     │ TIMEOUT     │
        └─────────────┘     └─────────────┘
                                   │
                             [DRAINED / OFF]

The Shallow vs Deep Health Check Trap:

  1. Shallow Health Check (/ping returning a static 200 OK): - Only proves the web server (Nginx) is alive. - Doesn't prove the database pool or the application's worker threads can actually process queries. - A server whose database is failing still looks healthy and keeps getting flooded with traffic.
  2. Deep Health Check (/healthz checking DB, Redis, Disk): - Makes sure the server is truly ready to serve end-to-end. - Watch out: Don't let the health check query against the DB get too heavy every 2 seconds, because hundreds of LB health checks can end up loading the database themselves (self-inflicted DoS).

4. Graceful Drain (Maintenance Without Downtime)

When you want to update the application on Server B: 1. Set Server B's status in the Load Balancer to DRAIN. 2. The Load Balancer stops sending new connections to Server B. 3. Requests already running on Server B are allowed to finish (given a drain timeout, say 30 seconds). 4. Once active connections hit 0, Server B is restarted / deployed with the new code. 5. Run the health check, then switch the status back to READY / UP.


5. Architecture Design Summary

Component Key Role
Virtual IP (VIP) The single public entry point that DNS sees
HAProxy / ALB Picks the target server based on the algorithm & health metrics
Backend Pool The set of worker nodes ready to take load
Health Check An automatic detection sensor that isolates broken servers
Graceful Drain A way to update the application without cutting off user transactions

6. Key Terms

  • Load Balancer (LB): Directs the flow of network traffic.
  • HAProxy: The industry-standard high-performance TCP/HTTP load balancer.
  • Backend Pool / Target Group: The list of target servers.
  • Round Robin: Sequential, take-turns distribution.
  • Least Connections: Distribution to the server with the fewest active connections.
  • Health Check: A periodic check of whether a server is alive.
  • Drain: Gradually stopping new traffic so a server can be maintained.