UNDER PRESSURE

Level 0 · Reading the Pressure

Session 1: One User, One Server

Session 1 / 3410 min read

1. Imagine if

Imagine if there were only one person in all of Indonesia using this app.

His name is Andi. At nine in the morning, he's sitting at a coffee stall, opens the app on his phone, and taps the Check Status button. No other users. No queue. The server is completely idle — CPU at 2%, plenty of free memory, the database carrying no load at all.

The most ideal conditions a system could ever experience.

The screen spins for a moment. His status appears. It took 1.4 seconds.

The question is: what were those 1.4 seconds spent on?

Nobody was queuing. Nobody was competing. There wasn't a single "the system is busy" excuse. And still, there are 1,400 milliseconds that have to be accounted for.

This session is a journey following that one request — like following a single drop of water from the tap all the way to the sea, and back again. Because if we don't understand where the time goes when the system is empty, we'll never understand what breaks when the system is full.


2. What actually happens

When Andi's finger touches the screen, his request passes through roughly ten stages. Most of those stages are not "our app".

Stage 1 — Looking up the address (DNS lookup). Andi's phone knows the server's domain name, but not its IP address. It has to ask a DNS resolver first. If the answer is already stored in the phone's cache: zero milliseconds. If not: 20–100ms just to ask "where does it live?".

Stage 2 — Knocking on the door (TCP handshake). Before a single byte of data is sent, the phone and the server have to greet each other three times: "hello" → "hello to you too" → "okay". These three back-and-forth trips cost one full round trip. On an Indonesian 4G network, a realistic round trip is 40–80ms.

Stage 3 — Exchanging keys (TLS handshake). This is personal data, so everything has to be encrypted. The TLS handshake needs one to two more round trips, plus cryptographic work on both sides. Add 80–150ms.

Notice: we've already spent about 200ms, and our app hasn't seen a single letter of Andi's request.

Stage 4 — The request sets off. The HTTP request travels across the cellular network, into the carrier's network, out to the internet, and into the data center network. 40–80ms.

Stage 5 — The front door receives it (web server / reverse proxy). A process sitting in front of the app accepts the connection, decrypts it, reads the headers, and decides where this request should be sent. Fast — 1–5ms — but this stage is worth remembering, because in Session 10 we'll see how this fast stage can become the cause of a system's death.

Stage 6 — A thread picks up the work. The app has a number of workers — let's call them threads — and one of them takes this request. When the system is empty, there's always an idle worker, so there's no waiting time. Remember that sentence well. That's the sentence that will change completely in Session 3.

Stage 7 — The app thinks. Validate the token, check authorization, build the query. This is "our code". Realistically 5–20ms.

Stage 8 — Asking the database. The app borrows a connection from the connection pool, sends the query, the database reads from memory or disk, and returns the rows. A simple, well-indexed query: 2–15ms. Plus the network trip inside the data center: 1–2ms.

Stage 9 — Building the answer. The query result is turned into JSON, encrypted again, and sent out. A few milliseconds.

Stage 10 — Going home. Across the same network, back to the phone. 40–80ms. Then the phone has to draw the screen — 50–200ms depending on how old the phone is and how heavy the page is.

The time budget

Stage Estimate Whose is it?
DNS lookup 0–150 ms Network / carrier
TCP handshake 40–120 ms Network
TLS handshake 80–260 ms Network + crypto
Request sets off 40–120 ms Network
Web server / proxy 1–5 ms Our infrastructure
App thinks 5–20 ms Our code
Database query 3–17 ms Our data
Building the response 2–5 ms Our code
Response goes home 40–320 ms Network
Phone draws the screen 50–400 ms User's device

The low end of the range is a good day: strong signal, a warm connection, a new phone. The high end is a morning at a coffee stall with patchy signal — and that's where Andi's 1.4 seconds live.

Add up the "our code" and "our data" columns: roughly 10 to 42 milliseconds.

Out of the 1.4 seconds Andi felt, the part we actually control might be only 2%.


3. What if we try...

What if we optimize the query?

Say we work hard and manage to cut the query from 15ms to 3ms. We save 12ms. Andi now waits 1.388 seconds instead of 1.4 seconds.

He won't feel a thing.

What if we double the server's CPU?

Stages 7 and 9 might get twice as fast. We save about 10ms. Andi still waits 1.39 seconds. We've just doubled our server bill for an improvement nobody can see.

This is the first lesson of the series, and one of the most expensive to forget:

Optimizing something that isn't the biggest contributor achieves nothing — no matter how impressive the optimization is.

If a system is slow when it's quiet, the answer is almost never inside the app. It's in the number of round trips, in the payload size, in handshakes being repeated over and over, or in the user's phone itself.

What actually matters in Andi's scenario are the things that sound unglamorous: reusing connections that are already open so handshakes aren't repeated, shrinking the payload so fewer packets are sent, and moving the entry point closer to the user so the round trip is shorter.

But — and this is the key — this story changes completely the moment Andi is no longer alone. With 5,000 people at once, that 15ms query can turn into 4 seconds without a single line of code changing. Why that happens is what the next three sessions are about.


4. The official name

What we just traced has names that will show up in design documents:

Request lifecycle — the complete journey of one request from the client into the system and back. When someone asks "where's the bottleneck?", the real question is: "at which stage of this lifecycle?"

Round trip time (RTT) — the time for one there-and-back trip between client and server. This number is the floor. A system can't be faster than the number of RTTs it needs, no matter how powerful its servers are.

Handshake — the introduction process before data can flow. The TCP handshake to open a connection, the TLS handshake to secure it.

Keep-alive / connection reuse — keeping a connection open so the next request skips Stages 2 and 3 entirely. In design documents this shows up as a small attribute like keepalive_timeout — and it's often a difference of hundreds of milliseconds.

Thread — a single worker unit in the app that handles one request at a time. There's a limited number of them. Right now this feels unimportant; in Session 5 it will become one of the most common causes of a system's death.

Connection pool — a set of database connections that are already open and used in turns, because opening a database connection is expensive. In Session 13, these two words will be at the center of the whole story.

TLS termination — the point where encryption is unwrapped. In this session it happens at the web server. In Session 9 we'll see why moving this point changes a lot of things.


5. In our world

The diagram above — one client, one server, one database — is the most simplified version. In a real production system, Stage 8 is rarely that simple.

When Andi checks his order status, the app doesn't always read from its own database. Often it has to ask another service: through an internal API gateway, then to the inventory system, maybe through one or two more layers. Every one of those hops is a new round trip — with its own handshake, its own timeout, and its own chance of failing.

That means "one user request" in a system that's been split into many services can actually mean four to six chained internal requests. If each hop adds 20ms, four hops add 80ms — before any useful work has been done.

The consequence to hold onto from now on: in a layered system, latency stacks up, and failure is contagious. Both of those are the theme of Level 3.


6. The performance tester's lens

If you test systems, this session has four direct implications:

Always measure a single-user baseline first. Before running a 5,000 virtual user scenario, run one user while the system is empty. That number is the floor — the system will never be faster than that under load. If the baseline is already 1.4 seconds, there's no point targeting 800ms at peak load. And if the baseline number rises from sprint to sprint, that's a pure regression with nothing to hide behind.

Know exactly where your load generator stands. If the generator sits inside the data center, you cut out almost all of the network cost — Stages 1, 2, 3, 4, and 10. The results will look far better than the reality users experience. That's not a wrong test, but its results answer a different question. Make sure everyone knows which question is being answered.

Check whether your script reuses connections. A script that opens a new TCP + TLS connection for every request is measuring something the real app never does. The difference can be 200ms per request — big enough to make a healthy system look like it's failing, or to overload the load generator until it becomes the bottleneck itself.

Separate client numbers from server numbers. The response time recorded by the load generator includes the network. Server-side metrics don't. When the two differ a lot, that gap isn't noise — that gap is the data. It tells you the problem is in the network or in a queue before the app, not inside the app.

This week's exercise: Take the one transaction your team tests most often. Draw its ten stages on a sheet of paper and fill in an estimated time for each stage. Mark which stages the development team can change, and which they can't. Bring the drawing to your weekly discussion — it's almost certain two people will draw it differently, and that difference is the most valuable conversation in this session.


7. Question for the next round

Andi got his answer in 1.4 seconds, and we now know where every millisecond went.

Now imagine it's not just Andi.

Imagine 500 people tapping the same button in the same second. The network is just as fast. The server isn't full yet. The query is still the same query, still 15ms, still using the same index.

But Andi now waits 4 seconds.

Not one of those ten stages got any slower. Everything still works exactly as before.

So where did those extra 2.6 seconds come from?

The answer isn't in that list of ten stages. The answer is in the space between those stages — a place we haven't looked at all yet.

That's what Session 2 is about: Where the Time Goes.


SYSTEM UNDER PRESSURE · Session 1 of 34