UNDER PRESSURE

Level 1 · Multiplying Machines

Session 11: One Door for Everyone: Reverse Proxy, Buffering, & Rate Limiting

What if one visitor filled out their form extremely slowly, and the door attendant had to wait for them to finish before serving anyone else? What happens if 1,000 of those slow visitors show up at once?

Session 11 / 345 min read

1. Reverse Proxy vs Load Balancer: Two Roles, Often One Product

A Load Balancer focuses on distributing load across many backend servers. A Reverse Proxy focuses on terminating client connections, buffering, caching, compression, and protection.

In modern practice (NGINX, HAProxy, Envoy, AWS ALB, Cloudflare), both roles often live in the same binary/service. The product accepts requests from the internet, acts as a reverse proxy (TLS termination, reading HTTP), then forwards them to the backend pool as a load balancer.

INTERNET
    │
    ▼
┌─────────────────────────────┐
│  REVERSE PROXY / LB (NGINX) │  ◄── ONE ENTRY DOOR
└──────────────┬──────────────┘
    │           │
    │ TLS Term  │ Buffering & Compression
    ▼           ▼
┌─────────────────────────────┐
│   BACKEND APPLICATION       │
└─────────────────────────────┘

2. The Slow Client Problem (Slow Client / Slowloris)

One of the biggest hidden dangers for an application server is the Slow Client: a client that sends its request very slowly (say, 1 byte per second) or reads the response very slowly.

Without a Reverse Proxy (Direct to App):

Slow Client ──(TCP Connection Held Open)──► App Worker Thread (BLOCKED)
  • Every slow connection eats up 1 thread/worker on the application server (Gunicorn, Node.js event loop, PHP-FPM).
  • Just a few hundred slow clients are enough to cripple the server's entire capacity (a Slowloris attack, or an unintentional DoS from bad mobile networks).

With a Reverse Proxy (Buffering):

Slow Client ──► Reverse Proxy (FULL BUFFER) ──► Backend (FAST)
                     │
              [Holds the entire request
               in a memory/disk buffer]
  • The reverse proxy captures the whole slow request in its buffer (memory or disk).
  • Only once the request is complete and finished does the proxy forward it to the backend in one fast burst.
  • The backend worker only spends business logic processing time, not time waiting on the client's network.

3. Request/Response Buffering

Type Description Why It Matters
Request Buffering The proxy reads the entire request body from the slow client, then sends it to the backend Protects the backend from Slow Client / Slowloris
Response Buffering The proxy receives the full response from the fast backend, then sends it slowly to the client The backend frees its thread as fast as possible; the slow client becomes the proxy's problem

Critical NGINX Configuration:

# Buffer request body (default on)
client_body_buffer_size 16k;       # Memory buffer before spilling to disk
client_max_body_size 10m;          # Limit upload size

# Buffer responses from the backend
proxy_buffering on;
proxy_buffers 8 16k;               # 8 buffers x 16KB = 128KB memory
proxy_busy_buffers_size 32k;       # Buffers being sent to the client

4. Static Offload & Compression: The Proxy's Job, Not the App's

Don't let the backend application (Python, Node, Go, Java) serve static files (CSS, JS, images, fonts) or do Gzip/Brotli compression. That wastes expensive CPU and worker threads.

At the Reverse Proxy (Do It Here):

  1. Static File Serving: nginx location /static/ { alias /var/www/assets/; expires 1y; add_header Cache-Control "public, immutable"; }
  2. Compression (Gzip / Brotli): nginx gzip on; gzip_types text/css application/javascript application/json; brotli on; brotli_types text/css application/javascript;
  3. Cache-Control Header: - Cache-Control: public, max-age=31536000, immutable for assets with hashed filenames. - Cache-Control: no-store for dynamic HTML / APIs.

The Result:

  • > 90% of static traffic is served directly by the proxy without ever touching the backend application.
  • The backend focuses only on business logic (APIs, DB queries, auth).

5. Rate Limiting: The First Line of Defense

Before a request reaches the business logic (which is expensive: DB queries, Redis hits, calls to other services), limit it at the front door.

Token Bucket / Leaky Bucket Algorithm:

  • Each IP/user gets "tokens" every second (for example 10 req/second).
  • A request without a token → HTTP 429 Too Many Requests (stopped at the proxy, the backend is never touched).
limit_req_zone $binary_remote_addr zone=api_limit:10m rate=10r/s;

location /api/ {
    limit_req zone=api_limit burst=20 nodelay;
    # burst=20: allow a spike of 20 req at once
    # nodelay: don't delay, reject immediately if the burst is exceeded
}

Layers of Protection:

Level Target Example
Edge / CDN Volumetric DDoS, Bots Cloudflare, AWS Shield
Reverse Proxy (L7) API Abuse, Brute Force, Crawlers NGINX limit_req, HAProxy stick-table
Application Business Logic Abuse Custom middleware (login attempts, password reset)

6. Health Checks & Circuit Breakers at the Proxy Level

A Reverse Proxy can also do passive health checks: - If a backend keeps returning 5xx or repeatedly hits a timeout → mark it DOWN. - Stop sending traffic to that node (circuit breaker). - Keep running active health checks periodically; bring it back once it recovers.

upstream backend {
    server 10.0.1.10:80 max_fails=3 fail_timeout=30s;
    server 10.0.1.11:80 max_fails=3 fail_timeout=30s;
}

7. Summary: The Reverse Proxy as the "Gate"

Function Done at the Proxy (Cheap) Not in the App (Expensive)
TLS Termination ✅ ❌
Request Buffering (Anti Slow Client) ✅ ❌
Response Buffering (Free Up Workers) ✅ ❌
Static File Serving ✅ ❌
Gzip / Brotli Compression ✅ ❌
Rate Limiting (Token Bucket) ✅ ❌
Path Routing to Microservices ✅ ❌
Access Log & Metrics ✅ ❌

Philosophy: The backend application should only run business logic. Everything infrastructure-related (network, protocol, protection) is handled at the front layer.


8. Key Terms

  • Reverse Proxy: A server that receives requests from clients and forwards them to the backend (one direction: in → out).
  • Forward Proxy: A proxy that clients use to go out to the internet (out → in); not the topic here.
  • Request Buffering: Holding a slow client's entire request before forwarding it to the backend.
  • Slow Client / Slowloris: A client that, intentionally or not, sends data very slowly and drains server resources.
  • Static Offload: Serving static files (CSS, JS, img) directly from the proxy/CDN.
  • Rate Limiting: Limiting the number of requests per unit of time per client (IP / User ID).
  • Circuit Breaker: Automatically stopping traffic to a backend that keeps failing.