UNDER PRESSURE

Level 0 · Reading the Pressure

Session 2: Where the Time Goes

Session 2 / 347 min read

1. Imagine if

Imagine you walk into a popular coffee shop at eight in the morning.

Behind the counter is a barista named Budi. Budi is very skilled. From the moment you give your order until the receipt prints and your coffee is ready, Budi needs only 50 seconds to serve one person.

But as you step through the door, there are 25 people lined up in front of the counter. You have to stand behind them, shuffling forward one slow step at a time, for 5 minutes (300 seconds) before it's finally your turn in front of Budi.

Total time from walking through the door to holding your cup of coffee: 350 seconds (almost 6 minutes).

Now imagine the owner sees that snaking line and sends Budi to a crash barista course. After the course, Budi is twice as fast — he can process an order in 25 seconds instead of 50.

The result? You walk in, wait behind 25 people for 300 seconds, then Budi serves you in 25 seconds. Your total wait is now: 325 seconds.

From 350 seconds down to 325 seconds. The line at the door still snakes around. Customers are still angry because they're late for work.

In the software development world, this story plays out every day. The database query is logged at 20ms. The app processes business logic in 30ms. But the user on their phone waits 3 seconds (3,000ms).

Where did the other 2,950 milliseconds go?

This session takes apart the biggest myth in performance measurement: assuming the app's processing time is the same as the user's waiting time.


2. What actually happens

When a user sends a request, the total time they experience — the Response Time — isn't a single number. It's the sum of three components working in different places:

\text{Response Time} = \text{Network Time} + \text{Queue Time} + \text{Service Time}

Let's take them apart one by one.

Component 1 — Network Time (Travel Time)

This is the time data spends crossing cables, radio waves, routers, and the TLS handshake we dissected in Session 1. Under normal conditions, this number is fairly stable as long as the user's signal and location don't change.

Component 2 — Service Time (Pure Work Time)

This is the time when the server's CPU is actually executing code instructions, or when the database's disk/memory is reading your rows of data. This is "Budi serving your order".

In most modern web apps, the service time for one simple request ranges from only 5 to 50 milliseconds.

Component 3 — Queue Time (Waiting Time)

This is the time when your request has already reached the server, but there's no idle CPU thread or database connection to work on it. Your request is placed in a memory queue (queue buffer) and forced to wait.

This is "the 25 people standing in front of the counter".

[ Request Arrives ] ────► ░░░ QUEUE (Queue Time) ░░░ ────► [ SERVER WORKS (Service Time) ] ────► [ Response Out ]
                            2,920 ms                                    30 ms

Why Does Queue Time Hide Itself?

The main reason queue time so often goes undetected is where it gets recorded.

Many teams measure app performance by placing timing code (execution timer) inside the app's functions:

start_time = now()
result = process_order(request) # <--- Service time only
log_duration(now() - start_time)

The timer above records only service time (30ms). That timer never knows the request had already been sitting in the web server's queue (for example NGINX or the Tomcat thread pool) for 2,920ms before the process_order() function was called.

On the internal APM dashboard, the graph shows "average response time 30ms". At the same moment, users are yelling because the app froze for 3 seconds. Both numbers are correct — they're just measuring different things.


3. What if we try...

What if we optimize the app code from 30ms down to 10ms?

Let's calculate the impact while the queue is backed up by 2,920ms:

  • Before optimization: 50\text{ms (network)} + 2{,}920\text{ms (queue)} + 30\text{ms (service)} = 3{,}000\text{ms}
  • After optimization: 50\text{ms (network)} + 2{,}920\text{ms (queue)} + 10\text{ms (service)} = 2{,}980\text{ms}

Saving 20ms out of 3,000ms is an improvement of 0.66%. Not a single user will notice.

What if we upgrade the server's CPU specs (Vertical Scaling)?

A faster CPU makes service time quicker. But if the bottleneck is the size of the thread pool or the number of locked database connections (contention), a bigger CPU doesn't reduce the waiting time at the door. The queue still piles up because the entrance is limited.


4. The official name

In architecture design documents and APM reports, these terms show up under standard names:

  • Service Time (T_s) — The pure time a component needs to finish one unit of work without any queuing delay.
  • Queue Time / Wait Time (T_q) — The duration a request spends in a queue buffer before a worker thread starts processing it.
  • Response Time / Latency (T_r) — The total duration from the caller's point of view (T_r = T_s + T_q + \text{network}).
  • Tail Latency (p95 / p99) — A measure of response time at the 95th or 99th percentile. If p99 is 3 seconds, it means 1% of all requests (say 10,000 out of 1 million users) wait 3 seconds or more.
  • Queue Depth / Backlog — The number of requests currently queued, waiting for their turn to be processed.

5. In our world

How does this queue time phenomenon show up in real production systems?

  1. Tomcat / NGINX Thread Exhaustion:
    When every worker thread (say 200 threads) is busy processing slow requests, request number 201 is held in the OS socket backlog buffer. The Java app doesn't record this latency because request 201 hasn't even reached the servlet yet.

  2. Database Connection Pool Bottleneck:
    The app has 100 worker threads, but the connection pool to PostgreSQL only opens 10 connections. Once those 10 connections are in use, the other 90 threads stop and queue, waiting for a DB connection to be released. Here, queue time happens inside the app's memory before the query is ever sent to the database.

  3. Peak Hours (Flash Sale / Payday):
    Under normal traffic, queue time is close to 0ms. But at peak hours, the moment processing capacity is exceeded even slightly, the queue explodes exponentially within seconds.


6. The performance tester's lens

As a performance tester, Session 02 changes how we read load test graphs:

  1. Don't Trust the Average (Average Is a Lie):
    Averages hide tail latency. An average of 200ms could mean 90% of users wait 50ms and 10% of users wait 5,000ms in the queue. Always ask for and monitor p95 and p99.

  2. Measure Latency from the Client Side (Load Generator):
    The key metric for testing how the queue holds up is the Response Time recorded by the testing tool (JMeter/K6/Gatling), not just the server's APM metrics. The difference between the JMeter number and the APM number is purely Queue Time + Network Time.

  3. Monitor Saturation Metrics (Queue Depth):
    During a load test, don't just monitor CPU & Memory. Monitor internal queue metrics: http_req_queue_depth, db_pool_waiting_threads, and tcp_listen_overflows.


7. Question for the next round

We've seen that when the system is full, the user's waiting time explodes from 30ms to 3 seconds because of the queue (queue time).

But there's one strange thing:

Why, when server CPU usage is only at 70%, does the queue already start snaking and response time start swelling? Isn't there still 30% of CPU free?

Why doesn't the queue curve rise in a straight line, but suddenly explode upward like a vertical wall?

The answer lies in queueing math and Little's Law.

That's what Session 3 is about: The Queue Never Lies.


SYSTEM UNDER PRESSURE · Session 2 of 34