UNDER PRESSURE

34 sessions · 5 level

System Under Pressure

A learning series on large-scale systems for performance testing teams

From one user on one server to architectures that carry millions of transactions. Every component shows up as the answer to a problem you have already felt — not as a term to memorize.

Start with Session 1

0 / 34 sessions done

Level 0 · Session 1–6

Reading the Pressure

What actually runs out when a system is “slow”?

  1. 01 One User, One Server 10 min read
  2. 02 Where the Time Goes 7 min read
  3. 03 The Queue Never Lies 3 min read
  4. 04 Three Numbers That Often Get Mixed Up 6 min read
  5. 05 What Actually Runs Out 4 min read
  6. 06 The Law That Limits Everything 4 min read

Level 1 · Session 7–15

Multiplying Machines

If one server isn’t enough, what is the price of “many servers”?

  1. 07 The Server You Ordered, Not the One You Got 6 min read
  2. 08 Bigger or More (Vertical vs Horizontal Scaling) 4 min read
  3. 09 If There Are Two Servers, Who Decides? (Load Balancer & HAProxy) 4 min read
  4. 10 Layer 4 and Layer 7: The Blind Guard vs the Smart Guard 4 min read
  5. 11 One Door for Everyone: Reverse Proxy, Buffering, & Rate Limiting 5 min read
  6. 12 Stateless Is a Requirement: Why Redis Is More Than a Cache 4 min read
  7. 13 Every Server Is Up, the System Is Still Down: SPOF, High Availability, & Split Brain 4 min read
  8. 14 When One Building Isn't Enough: GTM, Multi-DC, & Anycast DNS 4 min read
  9. 15 The Robot That Keeps Pods Alive: Kubernetes, HPA, & Probes 4 min read

Level 2 · Session 16–21

The Database Is the Bottleneck

Why does adding pods end up killing the system?

  1. 16 Pods Up, Database Down: The Arithmetic of Connection Storms & Lock Contention 3 min read
  2. 17 The Database Gatekeeper: PgBouncer, ProxySQL, & Transaction Pooling 4 min read
  3. 18 Reading Far More Often Than Writing: Read Replica & Replication Lag 4 min read
  4. 19 Holding Questions Back from the Database: Redis Cache, Hit Ratio, & Thundering Herd 4 min read
  5. 20 When One Database Isn't Enough: Sharding, Shard Keys, & Cross-Shard Queries 4 min read
  6. 21 Separating the Read Path and the Write Path: CQRS & Materialized Views 4 min read

Level 3 · Session 22–31

Systems That Depend on Each Other

How do hundreds of services fail together — and how do you prevent it?

  1. 22 When Services Are Split Apart: From Monolith to Microservices & the Danger of Fan-Out 4 min read
  2. 23 The Languages Services Use to Talk 6 min read
  3. 24 Traffic Between Services 6 min read
  4. 25 The Domino Effect: Cascading Failure, Timeout, & Circuit Breaker 4 min read
  5. 26 When Everyone Tries Again: Retry Storms, Jitter, & Backoff 3 min read
  6. 27 Absorbing Load Without Waiting: Message Queue, Backpressure, & DLQ 4 min read
  7. 28 Throttling the Flow: Rate Limiting, Leaky Bucket, & Token Bucket 4 min read
  8. 29 Shooting Our Own Servers: Load Testing, Spike Testing, & Chaos Engineering 3 min read
  9. 30 Seeing What Happens Inside: The Three Pillars of Observability (Logs, Metrics, & Traces) 4 min read
  10. 31 Reading a Design Doc Like an Insider 6 min read

Level 4 · Session 32–34

Eyes That Never Sleep

Who carries, stores, and shows the logs when an incident hits?

  1. 32 The Invisible Pipes (Sidecar, Fluent Bit, & Fluentd) 9 min read
  2. 33 The Price of an Index (Elasticsearch, ILM, & Kibana) 9 min read
  3. 34 Buy the Eyes or Build Your Own (Dynatrace vs the DIY Path) 9 min read

How to use this series