Level 3 · Systems That Depend on Each Other
Session 23: The Languages Services Use to Talk
1. Imagine if —
You have an Order microservice. Every time an order comes in, it asks three services: User (get the name and tier), Product (get the price and stock), and Voucher (check whether the voucher is still valid).
The problem: User sends back the entire user profile — including address, phone number, login history, and 40 other fields you don't need. All you want is one field: loyalty_tier.
On the other side, Voucher only gives you an answer after three separate calls: first to check the code, second to check the expiry, third to check the remaining quota — even though all three could be answered at once if you could spell out your question in detail.
Two problems pointing in opposite directions, but one root cause: a conversation format that doesn't fit the real need.
2. What actually happens
REST — The Language Everyone Understands
REST (Representational State Transfer) works on top of plain HTTP. Every resource has a URL, every operation has a verb (GET, POST, PUT, DELETE). It's easy to debug (open a browser, type the URL, see the result) and easy for anyone who has ever used the web to understand.
But REST speaks in units of "resources" — you get one full row of data, not just the fields you asked for. This is called over-fetching: bandwidth and parsing time wasted on data that gets thrown away immediately.
The opposite problem is under-fetching: one endpoint isn't enough, and you have to call several different endpoints to gather everything you need — known as the N+1 problem.
gRPC — Strict Contracts, High Speed
gRPC uses Protocol Buffers (protobuf) as its contract schema: you define the data structures and methods up front in a .proto file, and the client and server code is generated automatically.
The payoff: data travels in a binary format (not JSON text), avoiding the overhead of text serialization/deserialization. Communication runs over HTTP/2, which supports multiplexing (one TCP connection handles many requests at once) and bidirectional streaming.
The price: it's much harder for humans to debug (binary data can't be read in a browser), it needs HTTP/2 along the whole path, and schema evolution has to be handled very carefully so you don't break backward compatibility.
GraphQL — The Client Chooses
GraphQL flips the paradigm: it's not the server that decides the shape of the data, but the client that defines the query. The Order client can ask:
query {
user(id: "U-991") { loyalty_tier }
voucher(code: "HEMAT20") { valid expires_at remaining_quota }
}
One request, and the answer is exactly what was asked for. No more, no less.
The price: every query lands at a single point — the resolver — which processes it. A poorly optimized resolver can become very expensive when queries are deep, nested, or allowed to combine anything arbitrarily. This is what's called the resolver N+1 problem in GraphQL: one field at the top level that triggers thousands of separate database queries at the level below.
3. What if we try...
"Just build one endpoint per use case." Feels reasonable — but it doesn't scale. Six teams, twenty use cases, and you end up with twenty endpoints that are each a variation of the same data. Changing the schema of a single field turns into a coordination opera.
"Just use gRPC for everything." It's true that gRPC is more efficient. But external services consumed by browsers and mobile apps can't talk gRPC directly — HTTP/2 in the browser has limitations, and the team's debugging toolchain gets longer. No protocol can be the universal answer.
4. The official name
| Protocol | Wire Format | Transport | Good For | Watch Out |
|---|---|---|---|---|
| REST | JSON/XML (text) | HTTP/1.1+ | Public APIs, broad ecosystem | Over-fetching, N+1 |
| gRPC | Protobuf (binary) | HTTP/2 | Internal communication, streaming, low latency | Hard to debug manually, browsers need a proxy |
| GraphQL | JSON (text) | HTTP/1.1+ | Diverse clients, heterogeneous data | Unoptimized resolvers, query abuse |
In a design doc, you'll see:
- REST: GET /users/{id}, POST /orders
- gRPC: a block like service OrderService { rpc CreateOrder (OrderRequest) returns (OrderResponse); }
- GraphQL: a single /graphql endpoint with a type Query { ... } schema
5. In our world
In large-scale production systems, these three protocols live side by side rather than competing:
- REST at the northbound API: endpoints consumed by the frontend and external partners, because its ecosystem and tooling are the broadest.
- gRPC for internal east-west traffic: communication between microservices inside the cluster, where low latency and strict schema typing are the priority.
- GraphQL at the aggregation layer: a BFF (Backend For Frontend) that gathers data from many services and serves it to the mobile app or web with a single query.
During a flash sale campaign: thousands of orders come in through external REST, each one triggering dozens of internal gRPC calls between microservices. GraphQL at the BFF layer aggregates data for the order confirmation page. The audit log of all these interactions is what we'll cover in Level 4.
6. The performance tester's lens
Three numbers you should ask about when you see protocol choices in a design doc:
-
Payload ratio: How many bytes are sent vs how many bytes are actually used? A bad REST or GraphQL setup can have a 10:1 ratio — only 10% of the data is relevant.
-
Connection overhead: For gRPC, has the connection pool been configured? The cost of setting up an HTTP/2 connection isn't zero, and a connection storm can happen here just like it does with databases (Session 16).
-
Resolver depth / join cost: For GraphQL, ask the team to show you the deepest query in the schema. Tier queries by depth and cap the maximum query depth clients can send.
Test scenarios: - Load test a REST endpoint with 1,000 VUs fetching the same resource — measure payload size and parsing time - Compare gRPC vs REST latency for same-size payloads at 10,000 QPS - In GraphQL, send a deeply nested query that goes past 5 levels — make sure there's a query complexity limiter that rejects it before it touches the database
7. Question for the next round
REST, gRPC, and GraphQL settle the question of "in what format" services talk to each other. But there's a more important next question: when one service asks three other services at once, who's responsible for making sure that if a fourth service is slow — nobody else gets dragged down with it?
That's what we'll cover in Session 24: Traffic Between Services — Envoy, Service Mesh, and Trust Boundaries.