Level 0 · Reading the Pressure
Session 3: The Queue Never Lies
1. Imagine if
Imagine you run a small clinic with one general practitioner.
Patients arrive on average 1 person every 10 minutes (\lambda = 6 patients/hour).
The doctor examines 1 patient in 8 minutes (T_s = 8 minutes \rightarrow \mu = 7.5 patients/hour).
The question is simple: how long do patients wait in the waiting room?
2. What actually happens
Let's calculate Utilization (U):
Using the Average Response Time (T_r) formula for an M/M/1 system:
Service time: 8 minutes.
Total time in the system: 40 minutes.
Waiting time (queue): 32 minutes.
80% utilization \rightarrow customers spend 80% of their time just waiting.
3. What if we try raising the load?
Say arrivals increase to 1 patient every 9 minutes (\lambda = 6.67, U = 89\%).
Utilization went up 9 points (80% → 89%), but response time almost doubled (40 → 73 minutes).
This is the exponential queue curve. The closer you get to 100%, the more the curve shoots upward without limit.
4. The official name
| Term | Definition |
|---|---|
| Utilization (U) | The fraction of time the server is busy (U = \lambda / \mu). |
| Saturation | The point where U \approx 100\%; the queue grows without limit. |
| Queue Depth (L) | The average number of requests in the system (queued + being processed). |
| Little's Law (L = \lambda W) | L = throughput \times time in the system. Holds universally. |
| Knee Point | The utilization at which latency starts to spike sharply (usually 70–80%). |
| Buckle Point | The point of peak throughput before thrash / collapse. |
5. In our world (Software)
| Clinic | Software |
|---|---|
| Doctor | CPU / Worker Thread / DB Connection |
| Patient | Request / Query / Transaction |
| Waiting room | Socket Queue / Thread Pool Queue / Connection Pool Queue |
| Examination time | Service Time (T_s) |
| Patient's total time | Response Time (R) |
A real example:
DB Connection Pool = 100 connections (servers).
Query Service Time = 5 ms.
Target throughput = 15,000 req/s \rightarrow needs 75 simultaneous connections (Little's Law).
If a traffic spike pushes it to 20,000 req/s \rightarrow you need 100 connections (100% utilization).
The queue explodes. Latency climbs from 5 ms → hundreds of ms → timeout.
6. The performance tester's lens
- Don't trust averages. Find the Knee Point with a step-load test (raise the load in steps, record p95/p99 latency).
- Safe utilization target: 60–70% for interactive systems (APIs, web). 80%+ only for batch/throughput-oriented ones.
- Little's Law is a diagnostic tool:
L = \lambda \times R \rightarrow if R rises but \lambda stays flat, it means L (the queue) is swelling. - Saturation ≠ Error. A system can be healthy (HTTP 200) yet unresponsive (latency 10× normal).
7. Question for the next round
Session 4: "Three Numbers That Often Get Mixed Up"
Why do teams so often struggle to tell Concurrency, Throughput, and Response Time apart?
And how does Little's Law (N = X \times R) bind all three into a single whole that can't be pulled apart?