Level 0 · Reading the Pressure
Session 5: What Actually Runs Out
1. Imagine if
Imagine you have a work cabinet with six different drawers. - Drawer 1: Form paper (CPU) - Drawer 2: Writing desk (RAM) - Drawer 3: Filing cabinet (Disk I/O) - Drawer 4: Outgoing phone lines (Network) - Drawer 5: Room door keys (File Descriptor / Socket) - Drawer 6: Number of staff chairs (Thread Pool / DB Connection)
The operating rule is brutal: the office is declared completely paralyzed the moment ONE drawer is bone empty, even if the other five are still 90% full.
This is every systems engineer's nightmare: watching a CPU dashboard sitting calmly at 15%, while customers outside are shouting because their transactions are hitting mass timeouts.
2. What actually happens
Most beginner teams only watch two main charts: CPU Usage and RAM Usage. If both are still green, they assume the server is healthy.
But inside a Linux/Unix operating system, a server handles thousands of connections using many different resource abstractions:
- File Descriptor (FD): Every incoming TCP/HTTP connection, every open log file, and every database socket counts as 1 File Descriptor. The OS default limit is often just 1,024. The moment you hit that limit, the error
Too many open filesshows up instantly. - Database Connection Pool: Java/Node/Go applications cap the number of active connections to the DB (say, 50 connections). If 51 requests each need a 2-second query, the 51st request goes into the pool's waiting queue.
- Thread Pool: The number of worker threads in a web server (e.g. Tomcat/Puma). When every thread is stuck waiting on an external response, new requests are held back in the kernel socket.
- Disk I/O & IOPS: The CPU is fast, RAM has plenty of room, but the queue of transaction log writes makes the disk back up badly (high I/O wait).
- Garbage Collection (GC) Pause: In runtimes like the Java JVM, when memory fills up, the GC performs a Stop-the-World pause. CPU spikes to 100%, not to serve users, but to clean up memory garbage.
3. What if we try to find the leaking drawer?
Say we run a load test on an API Gateway: - Traffic: 3,000 RPS. - CPU: 22% (Very safe). - RAM: 4.2 GB / 16 GB (Plenty of room). - Response Time: Jumps from 20 ms \rightarrow 15,000 ms (Timeout).
After inspecting down at the kernel level:
It turns out the application's File Descriptor limit was set to 1024. At 3,000 RPS, incoming connections are rejected right at the socket's front door. The CPU stays cool because the code's instructions never even got a chance to run.
4. The official name
| Term | Definition |
|---|---|
| Resource Saturation | A condition where one specific resource has reached 100% of its capacity and starts queuing work. |
| Thread Pool Exhaustion | Every worker thread is busy; new requests are held in the application's internal queue. |
| Connection Pool Saturation | Every active database connection is in use; new queries wait for a connection to be released. |
| File Descriptor (FD) Limit | The maximum number of open files/network sockets per process (ulimit -n). |
| I/O Wait (%iowait) | The percentage of time the CPU sits idle waiting for data transfer to/from disk to finish. |
| GC Pause (Stop-The-World) | A runtime pause where all application execution is temporarily halted for memory cleanup. |
5. In our world (Production Systems)
In modern microservices architectures, drawer saturation often moves around quietly:
- The Ephemeral Port Exhaustion Case:
Microservice A calls Microservice B thousands of times per second without HTTP Keep-Alive. Local ports (range 32768–60999) run out because they get stuck in theTIME_WAITstate. The service fails to connect to its upstream. - The Connection Pool Leaks Case:
The code forgets to close the database connection in thefinallyblock. Within 10 minutes, all 100 connections are used up and the server locks up permanently.
6. The performance tester's lens
- Watch All 6 Drawers at Once: Never load test by looking only at CPU/RAM. Monitor FD count, Thread count, DB Active Connections, Network PPS, and Disk IOPS.
- Check the File Descriptor Limit: Make sure
ulimit -non the OS and the container has been raised (e.g. to65535or1048576) before testing begins. - Watch Out for %iowait: If the CPU looks like it's at 100% but most of it is iowait, don't add CPU; switch storage to SSD/NVMe or turn off excessive log writing.
- Analyze the GC Log: Record JVM pause durations. If GC pauses are > 500 ms at p99, that's your main source of latency spikes.
7. Question for the next round
Session 6: "The Law That Limits Everything"
What if we try doubling our servers from 10 to 20 machines, but capacity only goes up by 40%?
And why, when we go up to 40 machines, does the system's capacity actually drop?
The answer lies in Amdahl's Law & the Universal Scalability Law (USL).