UNDER PRESSURE

Level 3 · Systems That Depend on Each Other

Session 31: Reading a Design Doc Like an Insider

Session 31 / 346 min read

1. Imagine if —

You've just joined the team. There's already a document on your desk — the architecture of a system launching in three weeks. Fifteen minutes for the presentation. The rest is up to you: which point is most likely to break, and how much capacity is really needed.

People who read a design doc like its designers know where to look. They don't read every section equally — they go straight to the pressure points. And they bring the same list of questions every time, no matter what system is in front of them.


2. What Actually Happens

What's in a Design Doc

A good design doc usually contains: an overall architecture picture, the components and their dependencies, the expected traffic volume, the promised SLOs, and the technical decisions along with their reasoning.

What's not always there: capacity numbers calculated from real needs (not guesses), failure scenarios that have been accounted for, and load limits that have been tested.

That's the gap a performance tester needs to fill.

How to Read It: Five Steps

1. Read the SLOs first. Before understanding the architecture, understand the promises made to users: availability, p99 latency, maximum throughput. These determine when the system is considered to "pass."

2. Find the traffic convergence points. Where in the diagram do all the paths meet? Usually the API Gateway, the main load balancer, or the primary database. That's the first candidate to break when load rises.

3. Calculate capacity from business numbers. Not from "the servers we have" — from "transactions per hour at the campaign peak" times "processing time per transaction" times "safety factor."

4. Look for hidden dependencies. External services without an SLA, legacy systems without a circuit breaker, a shared database that isn't on the diagram because it's "always been there."

5. Check the promised observability. How will the team know the system is dying before users call? If the answer is "we'll check the logs," then Sessions 27–30 haven't been implemented yet.

Capacity Planning: From Business Numbers to Server Numbers

The basic formula:

Capacity = (Peak TPS × Avg Response Time) / 1000 × Safety Factor

A concrete example: - Flash sale campaign: 50,000 transactions per hour at peak = ~14 TPS on average, but with a 10× spike = 140 TPS peak - Avg database response time: 80ms - Connections needed: 140 × 0.08 = ~12 active connections - With a 3× safety factor and connection overhead: you need a pool of at least 40 connections - With 100 pods, each pod needs 1 connection to the pool — but the total pool must be able to serve 100 × max concurrent per pod

This is why "add more pods" doesn't automatically help: if the database connection pool is already full (Session 16), more pods just means more of them waiting in line.

The Standard Question List

This is the checklist a performance tester brings to every design doc review:

Capacity: - What peak TPS is promised? What data does that number come from? - Is there a safety factor? How much? - If traffic rises 5× above the prediction, which component collapses first?

Dependencies: - Which external services are called? What are their SLAs? - Is there a circuit breaker for every external dependency? - Which database is used? Has the connection pool been configured?

Failover: - If one zone goes down, where is traffic redirected? How long does it take? - Are health checks configured on the load balancer? - How does the system behave at a 100% cache miss? (Session 19 — cache stampede)

Observability: - What metrics are monitored to detect degradation earlier than users do? - Is there an alert for p99 latency — not just availability? - What MTTD and MTTR targets have been tested?

Load Testing: - Has a load test been run before go-live? - Which chaos scenarios have been run? - Where is the document with the latest load test results?


3. What If We Try...

"Let's just trust the numbers already in the document." The numbers in a design doc often come from assumptions made three months ago, before the business case changed. Capacity gets written based on "we have 4 servers" — not "we'll have 40,000 users per hour." The document is a starting point for asking questions, not the final answer.

"We'll test when we're close to go-live." By then, architecture changes are very expensive. Capacity questions have to be answered in the design phase — that's where changing things is still free.


4. The Official Name

Concept Short Definition
Capacity Planning Determining how many resources are needed to serve the expected load with a safety margin
Headroom The difference between installed capacity and capacity already in use — how much room is left before the system collapses
Cell-Based Architecture Splitting infrastructure into independent units (cells) so a failure's blast is contained, instead of one incident killing everything
Blast Radius How wide the impact of one component's failure is — what percentage of users are affected if one node or one zone goes down
Peak Factor The ratio between peak traffic and average traffic — usually 5–20× for consumer systems
Chaos Engineering The practice of introducing controlled failures into a production system to verify that redundancy actually works

5. In Our World

In a mature system, the design doc review is a performance tester's most important moment. Because this is where architecture decisions can still be changed cheaply — before code is written, before infrastructure is configured.

When a big campaign is coming: review the design doc three weeks before go-live, and discover that the database connection pool is configured for a 200 TPS peak, while the new business forecast says 800 TPS. The configuration change: one hour. If it's found one week after go-live: an incident, a rollback, overtime, a postmortem.


6. The Performance Tester's Lens

Four documents you need to ask for before you can evaluate a design doc:

  1. Historical traffic data — not predicted numbers, but numbers that actually happened. This is where the peak factor gets calculated, not guessed.

  2. Current resource limit configuration — how many database connections per instance, how big the application thread pool is, what the pod memory limit is.

  3. The latest load test results — when it was run, with what scenario, and what was found.

  4. The official SLO document — not the "usually fine" one, but the one that's been communicated to users and has an SLA attached.

With those four documents, the standard question checklist in Section 2 can be filled in with real numbers — and the result can become an actionable recommendation.


7. Question for the Next Level

Level 3 is done. You now know how to break a system apart, how services talk to each other, how to guard the traffic paths, how to hold back waves of load, and how to read an architecture you've never seen before.

One question remains: once all those services are running, how do you know the logs the application writes actually reach a place where they can be read at three in the morning — and don't quietly disappear when the pod that wrote them dies?

That's where Session 32: The Invisible Pipeline — Fluent Bit, Fluentd, and Sidecar Log Collectors begins.