UNDER PRESSURE

Level 2 · The Database Is the Bottleneck

Session 19: Holding Questions Back from the Database: Redis Cache, Hit Ratio, & Thundering Herd

What if 10,000 people ask the exact same thing within one second — and we ask the database 10,000 times?

Session 19 / 344 min read

1. Imagine If

A library has a single reference clerk. Every time a visitor asks "Where's the World History book?", the clerk walks 5 minutes down to the basement storeroom to check the shelves, then comes back to answer. If 1,000 people ask the same question in an hour, the clerk collapses from exhaustion.

So the clerk sticks a single memo (sticky note) on the reception desk: "The World History book is on Shelf 4B". Now the next 999 people who ask just read the memo in 1 second, without the clerk taking a single step.


2. What Actually Happens

In-memory caching (Redis / Memcached) is the most powerful shock absorber for a database:

  1. Cache-Aside Pattern (Lazy Loading): - The application checks Redis first (GET key). - Cache Hit: The data is found in Redis RAM -> Return it straight to the client in sub-millisecond time (0.5 ms). - Cache Miss: The data isn't there -> The application queries the Database (15–50 ms) -> Stores the query result in Redis with an expiration time (TTL) -> Returns the data to the client.

  2. The Arithmetic of Hit Ratio Effectiveness (H): - Effective response time:

    T_{eff} = (H \times T_{cache}) + ((1 - H) \times T_{db})
    - If T_{cache} = 1\text{ ms}, T_{db} = 50\text{ ms}:

    • Hit Ratio 99% \rightarrow T_{eff} = (0.99 \times 1) + (0.01 \times 50) = \mathbf{1.49\text{ ms}} (The database only receives 1% of the traffic!).
    • Hit Ratio 80% \rightarrow T_{eff} = (0.80 \times 1) + (0.20 \times 50) = \mathbf{10.8\text{ ms}} (The database gets hit with 20% of the traffic, 20 times heavier!).

3. Three Ways a Cache Can Kill a System

  1. Cache Stampede / Thundering Herd: - One very popular data key (e.g., a flash sale item) has a TTL of 60 seconds. - At second 60, the TTL runs out. At that exact millisecond, there are 5,000 concurrent requests that all hit a Cache Miss at once. - All 5,000 of those requests simultaneously send the exact same heavy query to the Database. The database collapses instantly! - Solution: Mutual Exclusion Lock (Mutex via Redis SETNX) or Probabilistic Early Expiration (XFetch algorithm).

  2. Hot Key & Big Key: - Hot Key: A single key accessed millions of times per second, pushing one Redis core thread on a particular node to 100% CPU. (Solution: Local In-Memory Cache in the Pod, like Caffeine / Go-Cache). - Big Key: A key with a 50 MB JSON payload. Its serialization and network transfer block the Redis event loop.

  3. Cold Start (Empty Cache After a Restart): - A Redis cluster has just been restarted or freshly deployed. The entire cache is empty (H = 0\%). - The full load of millions of requests goes straight through and slams into the database with no filter. - Solution: Cache Warming (pre-populating popular data before opening the traffic gateway).


4. The Official Name

  • Cache-Aside / Read-Through: Patterns for fetching data through a cache layer.
  • Cache Hit Ratio (H): The percentage of requests successfully served from cache RAM.
  • TTL (Time-To-Live): How long data is kept before it's automatically deleted.
  • Cache Stampede / Thundering Herd: A storm of queries to the database caused by popular keys expiring at the same time.
  • Cache Warming / Pre-heating: Filling the cache with critical data before the system takes public traffic.

5. In Our World

  • Redis Cluster vs Sentinel: Sentinel is for High Availability master-replica failover; Cluster is for sharding RAM data horizontally across 16,384 hash slots.
  • Two-Tier Caching (L1 + L2): L1 in the application pod's memory (sub-microsecond) + L2 in a centralized Redis (sub-millisecond) + L3 in the Database.

6. The Performance Tester's Lens

Crucial metrics and tests: 1. Hit Ratio Sensitivity Test: Simulate the Hit Ratio degrading from 99% down to 85%. What's the database's capacity limit before it falls over? 2. Key Expiration Stampede Simulation: Load test the moment a viral key is invalidated or expires. Does the mutex lock actually prevent a thundering herd? 3. Cold Cache Load Test: Start the load test with Redis empty to measure how well the database holds up in a mass-restart scenario.


7. Question for the Next Round

Cache and read replicas have freed the database from 98% of the read load. But what if the volume of our transaction write data has reached tens of terabytes, so that even the biggest database disk in the world can no longer hold it?

The answer is in Session 20: When One Database Isn't Enough (Database Sharding & Partitioning).