UNDER PRESSURE

Level 0 · Reading the Pressure

Session 4: Three Numbers That Often Get Mixed Up

Session 4 / 346 min read

1. Imagine if

Imagine you're watching a three-lane toll road.

At the toll entrance, 1,000 cars are lining up to get in at the same time. The toll's operations manager stands next to you and says proudly: "Look, we've got 1,000 vehicles on our toll road at once!"

But when you count the cars making it through the exit gate at the far end, the number stays consistent: exactly 10 cars per second.

An hour later, another 1,000 cars force their way in. Now there are 2,000 cars on the toll road. It's completely jammed, cars crawling bumper-to-bumper.

You count the exit gate again. The number has actually dropped to 4 cars per second.

The toll manager is confused: "How is it possible that twice as many vehicles came in, but fewer vehicles are making it out?"

In the performance testing world, misunderstanding these three numbers is the number one trap that most often wrecks load test analysis.

People often mix up how many people are trying (Virtual User / Concurrency) with how many transactions actually finish (Throughput / TPS).


2. What actually happens

In performance engineering, there are three main metrics tightly bound together in a triangle:

                  [ CONCURRENCY (N) ]
                   /               \
                  /                 \
                 /                   \
    [ THROUGHPUT (X) ] ────────── [ RESPONSE TIME (R) ]

Let's break down the exact definition of each:

1. Throughput (TPS / RPS)

Throughput is the amount of work finished per unit of time (usually Transactions Per Second / TPS, or Requests Per Second / RPS). It's the rate of water actually coming out of the tap.

If your system produces 500 TPS, it means 500 transactions were completed from start to finish within one second.

2. Concurrency (N / Concurrently Active Users)

Concurrency is the number of requests or users being processed at the same time inside the system at a single point in time. It's the number of cars currently on the toll road itself.

3. Response Time (R / Latency)

Response Time is how long one transaction takes from entering to finishing. It's how many minutes one car spends on the toll road from the entrance to the exit.

The Triangle Relationship (Little's Law, System Version)

These three numbers are tied together by a simple formula:

\text{Concurrency } (N) = \text{Throughput } (X) \times \text{Response Time } (R)

Or, if you account for the pause while users type (Think Time Z):

N = X \times (R + Z) \quad \Rightarrow \quad X = \frac{N}{R + Z}

The Three Phases of a Load Test Curve

When you gradually increase the number of Virtual Users (N) during a load test, the results graph will always pass through 3 main phases:

 Throughput
   (TPS)
    ▲             PHASE 2: SATURATION (KNEE)
    │           ┌──────────────────────┐
    │          /                        \  PHASE 3: BUCKLE (CRASH)
    │         /                          \
    │        /                            \
    │       /  PHASE 1: LINEAR             \
    │      /
    └─────┴──────────────────────────────────► Virtual Users (VU)
  1. Phase 1 — Linear Phase:
    The load is still light. Every time you add 100 Virtual Users, Throughput (TPS) rises proportionally in a straight line. Response Time stays low and constant.

  2. Phase 2 — Saturation / Knee Phase:
    The system reaches its maximum processing capacity (Max TPS). When you add more Virtual Users, going from 500 to 1,000, TPS doesn't rise at all (it flattens out). Why? Because all the CPU/threads are already 100% in use. Adding VUs only lengthens the queue (Response Time swells).

  3. Phase 3 — Buckle Phase (System Collapse):
    You keep forcing the VUs up to 2,000. The queue memory overflows, the CPU spends all its time on context switching and garbage collection, threads fight over locks. Throughput (TPS) suddenly plummets, while Response Time rockets sky-high. The system buckles (breaks).


3. What if we try...

What if we claim: "Our app can handle 10,000 concurrent users"?

That sentence is an ambiguous and often misleading claim when it doesn't state TPS and Response Time.

10,000 concurrent users who are mostly sitting still reading an article (think time of 30 seconds) need only 333 TPS of processing power.

But 10,000 concurrent users hitting the checkout button at the same time with no think time and a response time of 100ms need 10,000 TPS of processing power!

The exact same sentence can describe a very relaxed system or a world-class giant.

What if we set up a JMeter script with 1,000 Virtual Users and no Think Time?

A load test script that fires requests with no pause (zero think time) isn't simulating 1,000 humans. It's a Denial of Service (DoS) attack bombarding the server as fast as the code can loop.


4. The official name

In official performance test reports, these terms are written under standard names:

  • Throughput (TPS / RPS) — The number of business transactions or HTTP requests successfully completed per second.
  • Concurrency / Concurrent Users — The number of active user sessions interacting with the app within one time window.
  • Simultaneous Users — The number of users pressing a button in the exact same millisecond (a small subset of concurrent users).
  • Virtual User (VU) — A simulated thread (thread/coroutine) inside the load testing tool (JMeter/K6) that executes the steps of a test scenario.
  • Think Time (Z) — A simulated pause while the user reads the screen before taking the next action.
  • Pacing — A deliberate pause between the end of one scenario iteration and the start of the next, to keep the target TPS stable.
  • Buckle Point — The load point where the system's processing collapses and throughput drops drastically due to extreme queuing overhead.

5. In our world

How does misreading these 3 numbers trigger problems in the real world?

  1. The "1 VU = 1 TPS" Myth:
    Many beginner testers assume that if they set 500 Virtual Users in JMeter, the server is being tested at 500 TPS. That's wrong. If the app's response time is 2 seconds, then 500 VUs with zero think time produce only 250 TPS (500 / 2).

  2. Ticketing / Flash Sale Incidents (Simultaneous Spike):
    On a normal day, 50,000 concurrent users are spread out randomly with a think time of 10 seconds. Incoming TPS is only 5,000. But at exactly 12:00 when the promo opens, those 50,000 users turn into simultaneous users pressing the button in the same second. TPS jumps 10x instantly, throwing the system straight from Phase 1 to Phase 3 (Buckle).


6. The performance tester's lens

As a performance tester, Session 04 changes how we put together test result reports:

  1. Always Report the Full Trichotomy:
    Never report a single number. A test report must present all three numbers together: "At a load of 1,000 VUs (Concurrency), the system peaked at 450 TPS with a p95 Response Time of 1.2 seconds."

  2. Distinguish Concurrency from the Throughput Target:
    If the business asks that "the system must handle 2,000 TPS", use Little's Law to calculate how many VUs your test script needs based on estimated response time and think time.

  3. Find the Buckle Point to Know the Breaking Limit:
    A Stress Test deliberately pushes the system past the Knee Phase until it finds the Buckle Point. Knowing where the system breaks tells the operations team when rate limiting or a circuit breaker should kick in.


7. Question for the next round

We now know that when load is pushed past the knee point, the system enters the saturation phase and eventually buckles (TPS drops).

But when the system breaks and starts rejecting requests, what actually runs out inside the server?

Is it the CPU hitting 100%? Is the memory exhausted by an OutOfMemoryError? Or is there a secret cupboard inside the OS — like file descriptors or the socket backlog — that empties out with nobody watching?

That's what Session 5 is about: What Actually Runs Out.


SYSTEM UNDER PRESSURE · Session 4 of 34