UNDER PRESSURE

Level 3 · Systems That Depend on Each Other

Session 22: When Services Are Split Apart: From Monolith to Microservices & the Danger of Fan-Out

What if splitting one big application into 30 small independent services actually means a single user click triggers 30 hidden network calls that multiply latency and cripple the system?

Session 22 / 344 min read

1. Imagine If

In a traditional restaurant (Monolith), a waiter jots down the order, shouts it straight to the cook in the kitchen next door, grabs the plate, and serves it at the table within 2 minutes. Everything happens in one room with no middlemen.

Then the restaurant switches concepts and becomes a "Distributed Microservices Restaurant": - The waiter has to phone the Meat Station on the 2nd floor. - The Meat Station phones the Spice Station in the building across the street. - The Spice Station phones the Vegetable Station in a warehouse in another city. - If the phone line to the Vegetable Station drops for 5 seconds, the guests in the restaurant sit waiting at an empty table. A single lunch order triggers 30 internal phone calls.


2. What Actually Happens

When a team splits a Monolith architecture into Microservices, the computing boundaries shift: 1. From In-Memory Function Call to Network RPC: - In the Monolith: orderService.getInventory() is a local memory call that takes 0.0001 milliseconds (sub-microsecond) with 100% reliability. - In Microservices: That same call becomes an HTTP/REST or gRPC call across the physical network that takes 5–50 milliseconds and can fail at any moment due to packet drops or congestion.

  1. The Request Amplification Phenomenon (Fan-Out Explosion): - One simple request from the user's browser (for example: GET /dashboard) arrives at the API Gateway / Backend For Frontend (BFF). - The API Gateway fans out by calling 30 internal microservices: Auth, Profile, Cart, Recommendations, Loyalty, Notification, Inventory, Payment, Review, Promo, and so on.

  2. The Arithmetic of Tail Latency Amplification (p99 Trap): - Suppose every microservice has 99% reliability (only 1 in 100 requests is slow/fails). - If 1 user request has to call N = 30 services in parallel:

    \text{Overall Success Probability} = 0.99^{30} \approx 73.97\%
    - Which means: 26% of all user requests will hit severe delays or failures! The more services called in a single chain, the more mathematically prone the system becomes to slowing down.


3. Architectural Solutions

  1. API Gateway & Backend For Frontend (BFF): - Combines (aggregates) responses from various internal services and delivers them to the client in one compact payload.
  2. Asynchronous Parallel Fan-Out: - Use Promise.all() / worker threads so the 30 calls run at the same time instead of one after another sequentially (latency = \max(\text{service}) rather than \sum \text{service}).
  3. Graceful Degradation (Partial Degradation): - If the recommendation service times out, don't cancel the whole dashboard page! Return the main dashboard with an empty recommendations section or a static fallback.

4. The Official Name

  • Microservices Architecture: A modular architecture where each service has its own database and deployment process.
  • Fan-Out: The pattern of splitting one incoming request into several parallel requests to downstream services.
  • Tail Latency Amplification: The increased probability of delays at high percentiles caused by accumulated parallel dependencies.
  • BFF (Backend For Frontend): A dedicated API Gateway layer designed around the needs of a specific client (Mobile App vs Desktop Web).

5. In Our World

  • Uber / Grab Booking Screen: Showing the map, promo prices, driver estimates, and payment options triggers a fan-out to dozens of internal services within <500 ms.
  • Netflix Homepage: Every movie carousel row is assembled by Zuul / the API Gateway from dozens of personalization microservices.

6. The Performance Tester's Lens

Crucial metrics and tests: 1. Fan-Out Factor Under Load: How many internal sub-requests are created per 1,000 public ingress QPS? (Example: 1,000 external QPS = 30,000 internal QPS). 2. Tail Latency p99/p99.9 Curve: Measure the drastic gap between p50 (median) and p99 as the internal network load rises. 3. Partial Failure Resilience: Shut down 3 non-critical downstream services and verify whether the API Gateway still responds with a 200 OK status and a degraded payload.


7. Question for the Next Round

Thirty internal calls per request is already an everyday reality. But there's a deeper question than "how many calls": what language do these services use to talk to each other?

Look at one small example. A service only needs one value — the customer's name — but the service next door sends back the entire data row: address, transaction history, 40 columns it doesn't need. Another service has the opposite problem: it has to make three calls just to gather data that should have been requested all at once.

Those two symptoms have names: over-fetching and under-fetching. And the choice of language — REST, gRPC, or GraphQL — decides how many bytes travel over the wire every second, how many calls get created, and how easy that pipeline is to trace at three in the morning.

The wrong language makes the fan-out we calculated in this session twice as expensive without anyone noticing.

The answer is in Session 23: The Language Services Use to Talk — REST, gRPC, & GraphQL.