UNDER PRESSURE

Level 3 · Systems That Depend on Each Other

Session 26: When Everyone Tries Again: Retry Storms, Jitter, & Backoff

What if the good intentions of an app on 10,000 smartphones to "try again" all at once end up killing a server that has only just finished restarting, turning a 10-second outage into a 3-hour disaster?

Session 26 / 343 min read

1. Imagine If

A gate at a concert stadium suddenly shuts because its hinge jams. In front of the gate, 10,000 people are waiting in line. The staff manage to fix the hinge in 10 seconds.

The moment the gate opens a crack, all 10,000 people slam their bodies into it in sync, in the very same second (simultaneous retry). The gate shatters, the staff get trampled, and the concert is cancelled outright.

If those 10,000 people instead stepped back a few paces and waited for random, varying intervals (Exponential Backoff with Full Jitter), the crowd would flow in calmly without wrecking the gate.


2. What Actually Happens

When a server or downstream connection hits a temporary glitch (transient error):

  1. Blind Immediate Retry (Retrying Blindly With No Pause): - 5,000 clients fail a request at second 0. - With no pause, all 5,000 clients send a retry at the same time at second 0.1. - Server load that started at 5,000 QPS instantly jumps to 10,000 QPS (2x) right when the server is gasping for air. - This is a Retry Storm.

  2. The Fixed Delay Trap (Every Retry Lands in the Same Second): - Clients are given a fixed 1-second delay. - The result: the load spike just moves from second 0 to second 1 in a harmonic pattern (waves of traffic spikes). The server still gets hammered by periodic shock waves.


3. Three Weapons to Calm the Storm

  1. Exponential Backoff: - Increase the wait time exponentially with each consecutive failure:

    \text{Wait Time} = \text{Base} \times 2^{\text{attempt}}
    - Attempt 1: 100 ms \rightarrow Attempt 2: 200 ms \rightarrow Attempt 3: 400 ms \rightarrow Attempt 4: 800 ms.

  2. Full Jitter (Adding Random Noise): - Exponential backoff alone can still trigger harmonic waves if thousands of clients fail in the same second. - The Amazon AWS solution: Add a purely random component (Full Jitter):

    \text{Sleep} = \text{random}(0, \text{Base} \times 2^{\text{attempt}})
    - The dense retry traffic instantly untangles and spreads out evenly along the timeline (flat smooth distribution).

  3. Idempotency Key & Retry Budget: - Idempotency Key (X-Idempotency-Key: uuid): Guarantees that if the server receives the same payment request 3 times because of a network timeout on the client side, the user's credit card is charged only once. - Retry Budget (Envoy / Finagle): Caps retry requests at a maximum of 10% of total traffic. Once the retry quota runs out, requests fail fast immediately.


4. The Official Name

  • Retry Storm: A destructive traffic surge caused by synchronized attempts to repeat failed requests.
  • Exponential Backoff: A delay algorithm in which the wait duration doubles with every failure.
  • Jitter: Random variation injected into the backoff duration to break up the resonance of synchronized traffic.
  • Idempotency: The property of an operation where executing a request many times produces exactly the same end effect as executing it once.
  • Retry Budget: A mechanism that limits the retry ratio at the network proxy layer.

5. In Our World

  • AWS SDK / Google Cloud Client: Include Exponential Backoff with Full Jitter on all API calls by default.
  • Stripe / Midtrans Payment API: Require an Idempotency-Key header on the /charges endpoint to prevent double charges when automatic retries happen.

6. The Performance Tester's Lens

Crucial metrics and tests: 1. Traffic Spike vs Smooth Distribution: Simulate 5,000 clients failing at once with Fixed Delay vs Full Jitter. Compare the peak QPS charts (peak surge). 2. Double Charge Prevention: Send 10 parallel payment requests with an identical Idempotency-Key. Make sure there is only 1 bank mutation record and that the other 9 responses replay the result of the first transaction. 3. Retry Budget Depletion: Test how Envoy behaves when the downstream is completely dead: verify that the retry ratio stops at the 10% limit and doesn't multiply incoming traffic.


7. Question for the Next Round

When transaction data keeps pouring in and services have to communicate without waiting on each other directly, what architecture is used to safely hold millions of events in the background?

The answer is in Session 27: Absorbing Load Without Waiting: Message Queue, Backpressure, & DLQ.