Level 3 · Systems That Depend on Each Other
Session 28: Throttling the Flow: Rate Limiting, Leaky Bucket, & Token Bucket
What if the most important ability of a high-capacity system isn't serving every request, but having the courage to say "NO" to the excess 10% of requests in order to save the other 90% of users?
1. Imagine If
A fancy nightclub can only hold 200 people inside. At the entrance stands a burly bouncer with a velvet rope: - If 500 drunk people try to force their way in all at once, the bouncer doesn't let everyone in (because if they all got in, the dance floor would collapse and everyone would get hurt). - The bouncer only lets people in based on how many ticket wristbands are available (Token Bucket), or lets 1 person in every 5 seconds at a steady pace (Leaky Bucket). Everyone else outside is politely told: "Please wait your turn, come back in 10 minutes" (HTTP 429 Too Many Requests).
2. What Actually Happens
Without a flow limiter (Rate Limiter), a system is highly vulnerable to: - Layer 7 DDoS (HTTP Flood) attacks. - Rogue Scrapers / Bots that suck up the catalog database 500 times per second. - Users who accidentally run a script loop with no pause.
When a server receives load beyond its peak capacity, latency spikes sharply and the server collapses. Rate Limiting acts as a shield at the outermost gate (Reverse Proxy / API Gateway).
3. Four Flow-Limiting Algorithms
-
Fixed Window Counter: - Limit to 100 requests per minute (00:00 - 01:00). - Weakness (Boundary Burst): If 100 requests arrive at 00:59 and another 100 at 01:01, the server gets hit with 200 requests in 2 seconds!
-
Sliding Window Log / Counter: - Counts the number of requests within the last 60-second sliding window precisely using timestamps. Eliminates the gap at the window edges (boundary burst).
-
Token Bucket: - The bucket has a maximum capacity of B tokens. Tokens are refilled at a constant rate of R tokens/second. - Every incoming request must take 1 token. If the bucket is empty \rightarrow
429 Too Many Requests. - Advantage: Allows short traffic bursts (as long as there are tokens left in the bucket), then caps the average rate. -
Leaky Bucket: - Requests go into the bucket and drain out at a constant rate, like a leaky faucet. - Advantage: Produces an extremely smooth outgoing traffic flow with no spikes at all (smooth constant outflow).
4. Distributed Rate Limiter with Redis
In a multi-server system, the rate limiter has to be global and centralized:
- Redis Atomic Counter & Lua Scripting:
- Don't use GET followed by SET in your application code (prone to race conditions).
- Use a Redis Lua Script or the redis-cell module to execute the Token Bucket atomically in <1 millisecond.
- Standard HTTP Response Headers:
HTTP/1.1 429 Too Many RequestsRetry-After: 30(Try again in 30 seconds)X-RateLimit-Limit: 100(Quota limit)X-RateLimit-Remaining: 0(Remaining quota)
5. The Official Name
- Rate Limiting: Controlling the rate of network traffic coming into or going out of a service interface.
- HTTP 429: A standard HTTP status code indicating the client has exceeded the allowed request quota within a given time window.
- Token Bucket Algorithm: A packet control algorithm that allows controlled bursts based on a deposit of tokens.
- Leaky Bucket Algorithm: A rate control algorithm that normalizes the output rate to a constant.
6. In Our World
- GitHub / OpenAI API: Enforce strict limits (e.g. OpenAI Tier 1 = 500 RPM / 30,000 TPM) with
X-RateLimit-*headers. - Cloudflare / AWS WAF: Rate limiting based on IP and TLS fingerprint at the network edge.
7. The Performance Tester's Lens
Crucial metrics and tests:
1. Burst Capacity vs Throttling: Fire 500 requests within 100 ms at an endpoint with a 100 RPM quota. Make sure exactly 100 requests get through (200 OK) and the other 400 are rejected with status 429.
2. Distributed Redis Latency: Measure the latency overhead of the token check in Redis: it must stay under 2 milliseconds at a load of 10,000 QPS.
3. Retry-After Header Accuracy: Make sure a client that waits the number of seconds in the Retry-After header gets a 200 OK status on its next request.
8. Question for the Next Round
Now that our system is protected from request floods, what if we want to test the limits of its strength before a real disaster strikes on product launch day?
The answer is in Session 29: Shooting at Your Own Server: Load Testing, Spike Testing, & Chaos Engineering.