UNDER PRESSURE

Level 0 · Reading the Pressure

Session 3: The Queue Never Lies

Session 3 / 343 min read

1. Imagine if

Imagine you run a small clinic with one general practitioner.
Patients arrive on average 1 person every 10 minutes (\lambda = 6 patients/hour).
The doctor examines 1 patient in 8 minutes (T_s = 8 minutes \rightarrow \mu = 7.5 patients/hour).

The question is simple: how long do patients wait in the waiting room?


2. What actually happens

Let's calculate Utilization (U):

U = \frac{\lambda}{\mu} = \frac{6}{7.5} = 0.8 = 80\%

Using the Average Response Time (T_r) formula for an M/M/1 system:

T_r = \frac{T_s}{1 - U} = \frac{8}{1 - 0.8} = \frac{8}{0.2} = 40 \text{ minutes}

Service time: 8 minutes.
Total time in the system: 40 minutes.
Waiting time (queue): 32 minutes.

80% utilization \rightarrow customers spend 80% of their time just waiting.


3. What if we try raising the load?

Say arrivals increase to 1 patient every 9 minutes (\lambda = 6.67, U = 89\%).

T_r = \frac{8}{1 - 0.89} = \frac{8}{0.11} \approx 73 \text{ minutes}

Utilization went up 9 points (80% → 89%), but response time almost doubled (40 → 73 minutes).

This is the exponential queue curve. The closer you get to 100%, the more the curve shoots upward without limit.


4. The official name

Term Definition
Utilization (U) The fraction of time the server is busy (U = \lambda / \mu).
Saturation The point where U \approx 100\%; the queue grows without limit.
Queue Depth (L) The average number of requests in the system (queued + being processed).
Little's Law (L = \lambda W) L = throughput \times time in the system. Holds universally.
Knee Point The utilization at which latency starts to spike sharply (usually 70–80%).
Buckle Point The point of peak throughput before thrash / collapse.

5. In our world (Software)

Clinic Software
Doctor CPU / Worker Thread / DB Connection
Patient Request / Query / Transaction
Waiting room Socket Queue / Thread Pool Queue / Connection Pool Queue
Examination time Service Time (T_s)
Patient's total time Response Time (R)

A real example:
DB Connection Pool = 100 connections (servers).
Query Service Time = 5 ms.
Target throughput = 15,000 req/s \rightarrow needs 75 simultaneous connections (Little's Law).

If a traffic spike pushes it to 20,000 req/s \rightarrow you need 100 connections (100% utilization).
The queue explodes. Latency climbs from 5 ms → hundreds of ms → timeout.


6. The performance tester's lens

  1. Don't trust averages. Find the Knee Point with a step-load test (raise the load in steps, record p95/p99 latency).
  2. Safe utilization target: 60–70% for interactive systems (APIs, web). 80%+ only for batch/throughput-oriented ones.
  3. Little's Law is a diagnostic tool:
    L = \lambda \times R \rightarrow if R rises but \lambda stays flat, it means L (the queue) is swelling.
  4. Saturation ≠ Error. A system can be healthy (HTTP 200) yet unresponsive (latency 10× normal).

7. Question for the next round

Session 4: "Three Numbers That Often Get Mixed Up"
Why do teams so often struggle to tell Concurrency, Throughput, and Response Time apart?
And how does Little's Law (N = X \times R) bind all three into a single whole that can't be pulled apart?