Level 3 · Systems That Depend on Each Other
Session 24: Traffic Between Services
1. Imagine if —
Thirty microservices calling each other. Every service has its own driver who knows: which roads are jammed, when to wait patiently for a moment before trying again, when to stop completely before getting dragged under too, and whether the party it's talking to really is a legitimate service and not an intruder in disguise.
Without that driver, one slow service can take out everything above it in the chain — within seconds. This is no longer about the language being spoken (Session 23), but about who guards the lanes and makes the decisions at every intersection.
2. What actually happens
The Problem Without a Service Mesh
In a microservice architecture without a service mesh, every service is on its own for: - Retry — trying again when downstream is slow - Timeout — giving up waiting after a certain time limit - Circuit Breaker — stopping calls to a service that's already sick - mTLS — making sure communication between services is encrypted and authenticated
That means this infrastructure logic gets rewritten in every language, in every service, by every team. A bug in one implementation goes undetected until an incident hits production. And when an incident does happen, there's no central visibility into what's actually going on between those services.
Envoy: The Proxy That Thinks
Envoy is an L7 proxy designed for communication between services. Unlike Nginx, which faces end users (north-south), Envoy sits on the east-west path — between services.
What makes Envoy different: it understands HTTP/2, gRPC, and modern protocols natively, and it exposes everything that happens through metrics, logs, and traces — without changing a single line of application code.
The Sidecar Pattern: One Proxy per Pod
In Kubernetes, Envoy usually runs as a sidecar — an extra container living in the same pod as the application container. They share a network namespace: all traffic going in and out of the pod passes through Envoy, not directly from the application.
The result: the application knows nothing about retry, timeout, mTLS, or circuit breakers. All of that becomes Envoy's job. The application just speaks plain HTTP to localhost, and Envoy handles the rest.
Control Plane: The One Giving Orders
Envoy itself is the data plane — it's what moves the packets. But who tells it how to behave? That's the control plane.
In Istio (the most popular service mesh), the control plane is Istiod: it distributes configuration to every Envoy sidecar in the cluster — defining routing rules, timeout policies, circuit breaker thresholds, and the mTLS certificates to use.
Without a central control plane, you'd have to configure thousands of Envoy instances by hand.
Four Key Behaviors
1. Retry with Limits
Envoy can be configured to retry failed requests, with a cap on the number of attempts and clear conditions (for example: only retry on 503, at most 3 times). This prevents an uncontrolled retry storm.
2. Hierarchical Timeouts Timeouts can be set per request and per connection. If downstream doesn't respond within 200ms, Envoy cuts it off — it doesn't wait until the application's threads run out.
3. Circuit Breaker & Outlier Detection Envoy monitors the error rate of every upstream host. If one host produces more than X% errors within Y seconds, Envoy ejects that host from rotation (outlier detection) and opens the circuit — stopping all traffic to that host to give it time to recover.
4. mTLS: Zero-Trust on East-West Every Envoy sidecar holds a certificate issued by the control plane. When two services talk, both Envoys negotiate TLS — making sure the traffic is encrypted and both parties' identities are verified. No more "just trust it because it's in the same cluster."
3. What if we try...
"Write retry and circuit breakers ourselves in each service." This is what happens without a service mesh — and the result is inconsistent. Java implements it with Resilience4j, Python with tenacity, Node.js with opossum. Different versions, different behavior, different bugs. And when you want to change the timeout policy cluster-wide, you have to change code in every service.
"Use the same library in every service." Better, but there's still overhead: every language has its own library, every library upgrade requires redeploying every service, and a library can't control traffic at the TCP level.
4. The official name
| Concept | Short Definition | In the Design Doc |
|---|---|---|
| Envoy | Open-source L7 proxy, the default data plane for many service meshes | An "Envoy Proxy" or "Sidecar" box in the pod diagram |
| Sidecar | Deployment pattern: one proxy per pod, sharing the network namespace | An extra container in the Kubernetes pod spec |
| Service Mesh | Communication-layer infrastructure between services | Istio, Linkerd, Consul Connect |
| Data Plane | The layer that moves data (Envoy) | Sidecar instances |
| Control Plane | The layer that distributes configuration (Istiod) | Istio components in a dedicated namespace |
| Circuit Breaker | Automatically cuts the connection to a sick upstream | outlierDetection in an Istio DestinationRule |
| Outlier Detection | Envoy's mechanism for detecting and ejecting problematic hosts | consecutiveGatewayErrors, interval, baseEjectionTime |
| mTLS | Mutual TLS: both parties prove their identity | PeerAuthentication policy in Istio |
5. In our world
In large-scale microservice systems (100+ services), a service mesh isn't optional — it's mandatory infrastructure. Without it, an incident in one service cascades for minutes because nothing breaks the chain.
During a flash sale: traffic climbs to 40× normal. One inventory service starts timing out. Without a circuit breaker: every service calling inventory waits along with it, thread pools run out, and the system dies all at once. With an Envoy circuit breaker: after 3 consecutive errors, Envoy opens the circuit and returns a 503 error immediately without waiting — the other services stay alive, the team gets an alert, and inventory recovers within two minutes.
For performance teams: the observability that comes with a service mesh (Envoy metrics in Prometheus, traces in Jaeger/Tempo) is a trustworthy data source for answering "where is that latency actually being spent?"
6. The performance tester's lens
Four numbers that must be in the Envoy configuration before a load test:
-
Timeout per upstream call — what's the maximum time a single request to downstream is allowed to wait? Target: consistent with that service's own SLO (don't set 30 seconds if the SLO is 500ms).
-
Retry budget — how many total retries are allowed as a percentage of concurrent requests? Envoy's default of 20% is already good; don't raise it without thinking through the cascade.
-
Circuit breaker threshold —
consecutiveGatewayErrors: how many consecutive errors before a host is ejected? Too low = false positives, too high = slow to react. -
Envoy sidecar resource limits — every sidecar consumes CPU and memory. At 1,000 pods, that's 1,000 sidecars. Measure the overhead before committing to a service mesh.
Test scenarios: - Fault injection (Istio FaultInjection): force 50% of requests to one service to return 503, and verify the circuit breaker opens exactly at the threshold - Measure the extra latency from the sidecar (usually 1–5ms per hop within the same cluster) — make sure it fits in the p99 budget - Chaos: kill one service instance abruptly, and measure how long it takes before outlier detection ejects it from rotation
7. Question for the next round
The service mesh is already guarding the traffic: controlled retries, active circuit breakers, mTLS switched on. But there's a scenario not yet covered: not one slow service, but every service trying to reach the same service at the same moment — and all of them retrying together when they fail.
That's what we'll cover in Session 25: The Domino Effect — Cascading Failure, Retry Storm, and Jitter.