Level 3 · Systems That Depend on Each Other
Session 30: Seeing What Happens Inside: The Three Pillars of Observability (Logs, Metrics, & Traces)
What if the CPU and RAM on every server show a green 15%, yet users all over the world are screaming on social media because their checkout button spins forever, without a single error message in the terminal?
1. Imagine If
A modern commercial jet is flying at 40,000 feet in the middle of a pitch-black night storm. - Traditional Monitoring: The pilot has only two analog needles: "There's still fuel" and "Engine temperature normal". If the controls suddenly start shaking strangely, the pilot has no idea whether a wing has broken, the nose sensor has frozen, or the navigation computer has failed. - Full Observability: The cockpit is equipped with an advanced glass digital panel, a black box flight recorder, weather radar sensors, and real-time telemetry that immediately shows: "Hydraulic line number 3 on the left wing has lost 12% pressure due to a micro-leak at a valve joint".
2. Why Does Traditional Monitoring Fail?
In a distributed microservices architecture: - Server CPU and RAM can look very relaxed (10–20%), while the system is completely paralyzed by thread starvation, connection pool exhaustion, or lock contention. - Monitoring only answers: "Is the system alive or dead?" (Binary / Black-box). - Observability answers: "Why is the system behaving strangely, and where is the root cause among 50 microservices calling each other?" (Internal state inference).
3. The Three Pillars of Modern Observability
-
Metrics (Aggregate Metrics — "Something Is Wrong"): - Time-series numeric data that is cheap to store and ideal for visualization dashboards and real-time alerting. - The RED Method (Microservices Standard):
- Rate: How many requests per second are coming in (QPS).
- Errors: What percentage of requests fail (HTTP 5xx rate).
- Duration: The distribution of response latency (p50, p95, p99 histogram).
-
Distributed Tracing ("Where the Problem Happens"): - Industry standard: OpenTelemetry (OTel) & Jaeger/Tempo. - Every request entering the API Gateway is given a unique Trace ID. - This Trace ID is propagated through the HTTP/gRPC header (
traceparent) to every microservice hop. - It produces a Waterfall Span Chart visualization that precisely breaks down execution time:- Gateway: 1,200 ms
- Auth: 10 ms
- Cart: 15 ms
- Payment RPC: 1,150 ms (BOTTLENECK FOUND!)
- MySQL Query
SELECT FOR UPDATE: 1,130 ms (LOCK CONTENTION!)
- MySQL Query
-
Logs (Structured Event Records — "What the Detailed Cause Is"): - Leave plain text logs behind (
console.log("error disini")). - Use Structured JSON Logging that includestrace_id,user_id, anderror_stack. - Just search for the problematictrace_idin Grafana Loki / ElasticSearch to see the exact line of code that caused the failure within 1 second!
4. The 60-Second Incident Investigation Flow (The Golden Triad)
- Seconds 0–10 (Metric Alert): Prometheus fires an alert to Slack/PagerDuty: Payment p99 Duration > 1,000ms.
- Seconds 11–30 (Trace Analysis): The engineer opens Grafana Tempo / Jaeger, clicks one of the Trace IDs in the p99 spike, and sees a red Span stretch at
MySQL Lock. - Seconds 31–60 (Log Deep-Dive): The engineer clicks the log link for that Trace ID in Loki, reads the error
Lock wait timeout exceeded on row id=9821, then executes a rollback or kill-query mitigation.
5. Interim Recap: Where We Are Now
Four sessions left to the finish line. So far, the whole SYSTEM UNDER PRESSURE journey has covered: - Level 0 (Sessions 01–06): The anatomy of a single machine, CPU/RAM/Disk latency, queues, and the limits of Amdahl's law. - Level 1 (Sessions 07–13): Virtualization, horizontal scaling, L4/L7 Load Balancers, Reverse Proxy, Stateless Architecture, and High Availability. - Level 2 (Sessions 14–21): Multi-DC Anycast, Kubernetes Pods, Database Connection Storm, PgBouncer, Read Replica, Redis Cache, Sharding, and CQRS. - Level 3 (Sessions 22–31): Microservices fan-out, inter-service protocols (REST/gRPC/GraphQL), service mesh, cascading failure, retry storm, message queue, rate limiting, load & chaos testing, and the three pillars of observability. One session remains at this level: the synthesis of reading a design doc. - Level 4 (Sessions 32–34): Not started yet — log pipelines, Elasticsearch & Kibana, then Dynatrace vs. the DIY route.
6. The Performance Tester's Lens
Crucial metrics and tests:
1. Trace Propagation Completeness: Verify that 100% of downstream microservices preserve the traceparent header with no broken spans.
2. Observability Overhead Under Load: Make sure the OpenTelemetry collector agent consumes <2% CPU and <50MB RAM at 20,000 QPS of traffic.
3. MTTD & MTTR Reduction: Measure incident detection time (Mean Time to Detect) and recovery time (Mean Time to Resolve) before vs. after implementing distributed tracing.
7. Question for the Next Round
The tools are complete: metrics, traces, and logs all light up together. But complete tools don't mean the team is ready.
The remaining question is no longer about tools, but about how to read: you're handed an architecture diagram you've never seen before, and fifteen minutes to find where it will break. How much capacity is really needed — calculated from business numbers, not from guesses? And what standard questions should a performance tester ask every time a new design doc lands on their desk?
The answer is in Session 31: Reading a Design Doc Like an Insider — Capacity Planning & the Standard Question List.