UNDER PRESSURE

Level 1 · Multiplying Machines

Session 14: When One Building Isn't Enough: GTM, Multi-DC, & Anycast DNS

What if the entire building housing our system lost power all at once — and it turned out the backup server that saved us back in Session 13 was sitting in that very same building?

Session 14 / 344 min read

1. Imagine If

A bank installs two state-of-the-art vaults that duplicate each other's data in real time. If vault A breaks, vault B takes over within seconds. Everyone feels safe.

But one night, the main electrical substation on the road in front of the bank explodes, and the main water pipe bursts, flooding the building's entire basement. Both vaults die at the same moment. Local backups (Local High Availability) are helpless when the whole physical environment around them collapses.


2. What Actually Happens

Every load balancing technique we covered from Session 9 through Session 13 lives at the LTM (Local Traffic Manager) level. The LTM's job is to spread load within the walls of a single data center.

When we want to survive physical disasters (fires, massive blackouts, severed undersea fiber optic cables), the system has to jump to the global level:

  1. Availability Zone (AZ) vs Region: - AZ: One or more separate data centers a few kilometers apart (with independent power, generators, cooling, and internet connectivity), yet with very low latency between AZs (< 1–2 ms). - Region: Geographic areas separated by hundreds to thousands of kilometers (for example: Jakarta vs Singapore). Latency between regions is very real (15–30 ms).

  2. GTM (Global Traffic Manager) / GSLB (Global Server Load Balancing): - It doesn't distribute TCP/HTTP data packets directly; instead it works at the DNS level. - When a user's browser asks for the IP of app.perusahaan.com, the GTM checks the health status of each data center and the user's location:

    • Users from Jakarta are sent to the Jakarta DC IP (103.x.x.x).
    • Users from Tokyo are sent to the Tokyo DC IP (133.x.x.x).
    • If the Jakarta DC goes down, the GTM automatically answers DNS queries with the backup DC's IP.
  3. Anycast Routing (BGP): - Instead of relying on DNS TTLs that ISP caches are slow to update, several data centers around the world announce the exact same public IP address via BGP (Border Gateway Protocol). - Internet routers automatically send data packets to the data center that's closest in network topology terms. If one DC dies, its BGP route is pulled (withdrawn), and traffic shifts instantly without waiting for a DNS refresh.


3. What If We Try...

"Let's just set up DNS Failover with a 60-second TTL!"

Why does this so often fail completely during an incident? - Stubborn DNS Caching: Users' local ISP resolvers often ignore low TTLs and cache DNS records for hours (TTL overriding). - The result: the primary DC is completely dark, yet 40% of users keep sending traffic to the ruins of the old DC because their DNS cache hasn't expired yet. - Multi-DC Replication Lag: Sending data synchronously between continents runs into the speed of light in fiber optics. If replication is asynchronous, then when disaster strikes, the last 5 seconds of data in the primary DC haven't made it across to the backup DC yet (resulting in data loss, or RPO > 0).


4. The Official Name

  • GTM / GSLB (Global Traffic Manager): A smart, multi-region, DNS-based load balancer.
  • LTM (Local Traffic Manager): A load balancer inside the data center (HAProxy, F5 BIG-IP LTM, NGINX).
  • Anycast BGP: A routing method that maps one IP address to many data center locations.
  • Split Horizon / GeoDNS: Returns different IPs based on the geographic location of the client IP.
  • Dual-Region Active-Passive / Active-Active: Multi-region architecture patterns.

5. In Our World

On platforms like AWS/GCP/Cloudflare: - Cloudflare / AWS Route 53 / Akamai GTM: Handle DNS routing and endpoint health checking in each region. - Anycast CDN: The edge network accepts TLS connections in the city closest to the user, then forwards them to the origin data center over a private backbone network. - Chaos Engineering (Region Evacuation Drill): Engineering teams deliberately cut DNS to a primary region during working hours to prove the secondary region can absorb 100% of the load without knocking over the database.


6. The Performance Tester's Lens

Critical metrics and tests: 1. DNS Failover Convergence Time: How many real minutes does it take until 99% of global traffic has truly moved when one data center is shut down? 2. Cross-Region Latency Penalty: How much does RTT (Round Trip Time) jump when Indonesian users are forced over to a data center in Singapore or Japan? 3. Replication Lag Under Load: How many milliseconds of database replication lag are there between DCs at peak write load? 4. Capacity Headroom: Does the backup DC have enough server capacity and DB connections to absorb a sudden double load?


7. Question for the Next Round

Now we know how to steer traffic to buildings all over the world. But inside those buildings, who's actually responsible for spawning, killing, and healing hundreds of server containers in a matter of seconds, with no human help?

The answer is in Session 15: The Robot That Keeps Pods Alive (Kubernetes & Container Orchestration).