UNDER PRESSURE

Level 1 · Multiplying Machines

Session 8: Bigger or More (Vertical vs Horizontal Scaling)

What if we just buy a server that's twice as big? How long can that trick keep working?

Session 8 / 344 min read

1. Imagine If: One Giant Counter vs a Row of Counters

Imagine a post office with one very busy clerk serving 500 people per hour. Because the line is snaking out the door, management makes the first, most intuitive decision: replace that clerk with a superhuman. This super clerk can type 4× faster and read 4× quicker.

For a while, the line clears up. But next year, visitors grow to 5,000 people per hour. Now management goes looking for a god-tier superhuman who can type 50× faster.

This is where two big walls come crashing down: 1. The Physical Wall: A superhuman who works 50× faster doesn't exist in the real world. Silicon components (CPU clock speed, memory bus bandwidth, socket limits) have physical limits on heat and signal propagation. 2. The Economic Wall (Exponential Cost): A 128-core server doesn't cost just 4× what a 32-core server costs; it can cost 12× to 20× as much, because of the very complex and expensive multi-socket NUMA bus architecture.

The second option is adding more counters: instead of hunting for one superhuman, open 10 ordinary counters staffed by 10 normal clerks.


2. The Mechanism Behind the System: Scale Up vs Scale Out

VERTICAL SCALING (SCALE UP)                HORIZONTAL SCALING (SCALE OUT)
┌─────────────────────────────────┐        ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│  SINGLE GIANT SERVER            │        │NODE 1│ │NODE 2│ │NODE 3│ │NODE 4│
│  64 Core → 128 Core → 256 Core  │        └──────┘ └──────┘ └──────┘ └──────┘
│  512 GB RAM                     │           ▲        ▲        ▲        ▲
│  Single Failure Point           │           └────────┴──┬─────┴────────┘
└─────────────────────────────────┘                       │
                                                    LOAD BALANCER

2.1 Vertical Scaling (Scale Up)

  • Definition: Adding resource capacity to a single physical machine/instance (CPU, RAM, NVMe disk I/O).
  • Pros:
  • Zero changes to the application architecture (no shared session needed, no distributed locking needed).
  • Absolute data consistency (ACID in local memory/disk).
  • Very low latency between components (communication over the internal bus/RAM, not network cables).
  • Cons:
  • There's an absolute upper limit (hardware ceiling).
  • Downtime during physical hardware upgrades.
  • Single Point of Failure (SPOF): If the motherboard or RAM of that giant machine crashes, 100% of the system goes down.
  • Exponential Cost Curve: Enterprise-class hardware costs climb steeply and exponentially at the top end of capacity.

2.2 Horizontal Scaling (Scale Out)

  • Definition: Adding more machines (commodity nodes / replicas / pods) that work in parallel to share the workload.
  • Pros:
  • Nearly unlimited theoretical capacity (elastic scalability).
  • High Availability: If 1 of 10 servers burns down, the other 9 keep serving the remaining 90% of traffic.
  • Cost Efficiency: Uses standard, mass-produced machines (commodity hardware).
  • Cons:
  • Needs a traffic distribution layer (Load Balancer).
  • The application must be stateless (it can't store login sessions in the server's local memory).
  • Has to deal with network latency between servers and the complexity of distributed data concurrency.

3. The Cost Curve and Physical Limits (The Law of Hardware Economics)

Metric Scale Up (Vertical) Scale Out (Horizontal)
Maximum Limit Bounded by physical socket & chipset limits Elastic, thousands of commodity nodes
Cost Curve Exponential (
O(e^k)
)
Linear (
O(N)
)
Application Readiness Works right away (Legacy friendly) Must be Stateless + Central Store
Disaster Resilience Fragile (Single Point of Failure) Resilient (Tolerates N-1 failure)
Ops Complexity Low at first, high when you hit the ceiling Needs orchestration, CI/CD, monitoring

4. The Performance Tester's Lens

For a performance tester, the difference between vertical and horizontal determines the failure pattern during a stress test:

  1. On a Vertical System (Scale Up): - As load climbs from 500 to 5,000 RPS, CPU/RAM metrics rise in a straight line up to 90–95%. - Once the limit is hit, the CPU's internal queue (run queue latency) explodes instantly, and system latency shoots up sharply (hockey stick curve). - All users experience the slowdown at the same time.

  2. On a Horizontal System (Scale Out): - Load is distributed across N nodes. - If one node suffers a memory leak or CPU throttling, only a small share of requests fail (partial degradation), as long as the load balancer has good health check and auto-drain mechanisms.


5. Question for the Next Round

"If we decide to scale out to 5 identical machines... who stands at the door deciding who goes to which server? And what happens if the one deciding splits the load badly?"

(Continue to Session 9: If There Are Two Servers, Who Decides?)