Level 3 · Systems That Depend on Each Other
Session 29: Shooting Our Own Servers: Load Testing, Spike Testing, & Chaos Engineering
What if the best way to make sure our system doesn't die in users' hands during a midnight flash sale is to deliberately shoot and kill our own servers during working hours in broad daylight?
1. Imagine If
A team of NASA astronauts never flies straight to the moon after just reading a paper manual. They step into a zero-gravity simulation room, cut off the simulated oxygen supply, sabotage the navigation system, and drill their panic responses until they become automatic reflexes under pressure (GameDay & Chaos Simulation).
If we only find out our servers crash during a product launch in front of 50,000 real buyers, that isn't testing — that's reputational suicide for the business.
2. Why Is Testing in a Dev Environment So Often Misleading?
An application that runs smoothly on a developer's laptop (localhost) or in staging often falls apart completely in production because of 3 classic lies:
1. Empty Data Volume: The staging database holds only 100 rows of dummy data (every query is lightning fast), while production holds 50 million rows (a full table scan kills memory).
2. No Network Latency: On localhost, latency between services is 0 ms. In cross-region production, 30 ms of latency amplifies blocking threads.
3. Zero Concurrency: Tested by 1 QA person vs. stormed by 10,000 simultaneous connections fighting over database locks.
3. A Taxonomy of Load: 5 Types of Load Tests
- Smoke Test: A minimal load test (1–5 users) to verify the basic functionality of the test script.
- Load Test: Tests system performance at the expected normal peak load (for example: 2,000 QPS for 30 minutes).
- Stress Test: Gradually raises the load beyond normal capacity to find the system's breaking point.
- Spike Test: Fires an extreme, instant traffic surge (from 100 QPS jumping to 10,000 QPS in 5 seconds) to test how the auto-scaler and rate limiter respond.
- Soak Test (Endurance): Runs a constant load for 12–48 hours nonstop to detect hidden Memory Leaks or resource handle leaks.
4. Chaos Engineering: Injecting Destruction
Pioneered by Netflix through Chaos Monkey: - Core principle: Assume that anything that can break will break in production. - Randomly kill Kubernetes pods, cut the network links between data centers, inject 5 seconds of latency into the database, or fill a disk to 100% during a normal working day. - The goal is to prove that the Failover, Circuit Breaker, and Self-Healing architecture works automatically without human intervention.
5. The Official Name
- Load Testing: Testing to measure how a system responds under a given load.
- Breaking Point: The load level at which the system starts failing to meet its SLA (latency skyrockets or error rate >1%).
- Soak Testing: Long-duration load testing to look for gradual performance degradation (memory leaks).
- Chaos Engineering: The discipline of experimenting on a system to build confidence in its ability to withstand turbulent conditions in production.
- GameDay: A scheduled incident simulation exercise where the engineering team practices its emergency response.
6. In Our World
- k6 (Grafana): A modern JavaScript/Go-based tool for defining load test scenarios as code, with high-precision p95/p99 metrics.
- Chaos Mesh / LitmusChaos: Open-source Chaos Engineering platforms for the Kubernetes ecosystem.
7. The Performance Tester's Lens
Crucial metrics and tests:
1. Threshold Pass/Fail Criteria: Define strict tolerance limits in the CI/CD pipeline (example: http_req_duration: ['p(95)<200'], http_req_failed: ['rate<0.01']).
2. Memory Leak Slope: Analyze the RAM allocation graph over a 24-hour soak test: make sure it forms a flat plateau, not a permanently rising slope (staircase memory leak).
3. Chaos Recovery RTO: Measure how many seconds it takes the cluster to return to 100% healthy after 1 worker node is suddenly killed under a 3,000 QPS load.
8. Question for the Next Round
Once our system has been tested with thousands of requests and is ready to fight in the real world, how does the engineering team know what's happening inside the belly of hundreds of microservices when something strange happens in the middle of the night?
The answer is in the curriculum's grand finale session: Session 30: Seeing What Happens Inside: The Three Pillars of Observability (Logs, Metrics, & Traces).