Level 1 · Multiplying Machines
Session 7: The Server You Ordered, Not the One You Got
1. Imagine if
Imagine you rent 1 server with 8 vCPU and 16GB of RAM from a cloud provider to run a payment application.
On Monday morning, you run a load test with 1,000 virtual users. The results look great: - Average response time: 120ms - p99 response time: 180ms - Throughput: 1,500 TPS with zero errors.
You record these numbers as your official performance baseline.
Then on Thursday night, you run the exact same test script, with the same data and a server configuration that hasn't changed by a single line. But the results fall apart: - Average response time: 450ms - p99 response time: 3,200ms (up 17 times!) - Your server's internal CPU is only at 35%, yet transactions are piling up and timeout errors start showing up.
The application didn't change. The test script didn't change. The server specs on the cloud dashboard read exactly the same.
So why can the same server deliver performance that's worlds apart?
2. What actually happens
The answer: The server you rented on the dashboard isn't its own physical computer.
You're sharing one giant physical server with dozens of other renters (tenants) you've never met.
In modern infrastructure, a server comes in three forms:
+-------------------------------------------------------------------------+
| 1. BARE METAL |
| [ Application ] ➔ [ Host OS Kernel ] ➔ [ Single Physical Hardware ] |
+-------------------------------------------------------------------------+
| 2. VIRTUAL MACHINE |
| [ App A | Guest OS ] [ App B | Guest OS ] [ App C | Guest OS ] |
| --------------------------------------------------------------------- |
| [ Hypervisor (KVM / VMware ESXi / Xen) ] |
| --------------------------------------------------------------------- |
| [ Physical Hardware ] |
+-------------------------------------------------------------------------+
| 3. CONTAINER |
| [ App A (Bin/Libs) ] [ App B (Bin/Libs) ] [ App C (Bin/Libs) ] |
| --------------------------------------------------------------------- |
| [ Container Engine (Docker / CRI-O / Containerd) ] |
| --------------------------------------------------------------------- |
| [ One Host OS Kernel (cgroups + namespace) ] |
| --------------------------------------------------------------------- |
| [ Physical Hardware ] |
+-------------------------------------------------------------------------+
A. Bare Metal: A Whole Physical Machine
- Every CPU core, the RAM, the motherboard bus, the network card (NIC), and the disk controller are used exclusively by your application.
- Pros: Deterministic, consistent latency, with no virtualization overhead.
- Cons: Expensive, slow to provision (takes hours or days), and hard to resize dynamically.
B. Virtual Machine (VM): Hypervisor Isolation
- A hypervisor slices one large physical machine into many independent virtual machines. Each VM carries its own copy of an OS kernel (Guest OS).
- Pros: Strong security isolation, and it can run many different OSes on one physical server.
- Cons: There's virtualization overhead (5–15%), slow boot times (minutes), and it's vulnerable to overcommit.
C. Container: Sharing the OS Kernel
- Not a virtual machine, but an ordinary process on the host OS isolated using Linux kernel features:
1.
Namespaces: Give the process the illusion that it's running alone (PID, mount, network, and user isolation). 2.cgroups(Control Groups): Cap the maximum share of CPU, memory, and I/O. - Pros: Very lightweight, starts in milliseconds, minimal memory footprint.
- Cons: Shares the same single host kernel. If the kernel hits a kernel panic, every container on that host dies with it.
3. The Cloud Provider's Secret: Overcommit & Noisy Neighbor
Why can renting cloud vCPUs be so cheap?
Because providers rely on a principle of probability: the Overcommit Ratio.
If one physical server has 64 physical cores, the provider doesn't just sell 64 vCPUs. They can sell 128 to 256 vCPUs to different customers (an overcommit ratio of 1:2 up to 1:4), assuming not every customer uses 100% CPU in the same second.
[ PHYSICAL SERVER: 64 PHYSICAL CORES ]
├── Customer A vCPU (Fintech Company - Running Load Test) ➔ 32 vCPU
├── Customer B vCPU (Heavy Batch Processing App) ➔ 32 vCPU
├── Customer C vCPU (Machine Learning Training) ➔ 32 vCPU
└── Your vCPU (Payment Application) ➔ 32 vCPU
-------------------------------------------------------------------------
TOTAL vCPU SOLD = 128 vCPU (200% Overcommit on 64 Physical Cores)
The Noisy Neighbor Phenomenon
When Neighbor B and Neighbor C suddenly start running heavy compute jobs: 1. The hypervisor is forced to queue up slices of physical CPU execution time (CPU scheduling latency). 2. Your application's network packets get held in the physical network card's ring buffer because the neighbors are flooding the bandwidth. 3. The storage disk heads fight over the IOPS queue.
Your application doesn't crash and your internal CPU only reads 35%, but your threads are stalled waiting for their turn on a physical CPU core. In Linux metrics, this shows up as %steal time (CPU Steal).
4. The official name
| Term | Definition in Design & Operations Documents |
|---|---|
| Bare Metal | A single physical compute server with no hypervisor layer, rented exclusively (single-tenant). |
| Hypervisor (Type-1 / Type-2) | The software/firmware layer (like KVM, ESXi) that creates and manages virtual machines. |
| Guest OS vs Host OS | The OS running inside a VM (Guest) versus the main OS running directly on the hardware (Host). |
| Containerization | An OS-level virtualization method using Linux Namespaces and cgroups without running a second kernel. |
| Overcommit Ratio | The ratio of total virtual vCPU/RAM allocated to tenants compared to the actual physical capacity. |
| Noisy Neighbor | A situation where one tenant's workload on a shared server drains shared resources and hurts other tenants' performance. |
| CPU Steal (%steal) | The percentage of time a virtual CPU waits for the physical hypervisor to allocate execution time while the physical cores are busy. |
5. In our world (Real Production Systems)
Real Scenario: Mysterious Latency Spikes at Peak Hours
A digital bank launched its transaction microservices on a public cloud using standard VM instances (Shared Multi-Tenant).
- Symptom: Every day from 12:00–13:00, transaction p99 latency jumped from 150ms to 4,500ms. The developer team checked the application logs, database queries, and garbage collection, but everything came back clean.
- Investigation: APM monitoring showed the VM's
%stealmetric spiking as high as 42%. That means 42% of the processor time that should have been running application threads was being grabbed by other tenants' VMs on the same physical host. - Architectural Solution: 1. Moved the database and core payment engine to a Dedicated Instance / Bare Metal (zero overcommit). 2. Enabled CPU pinning on the Kubernetes nodes so threads don't hop between physical cores.
6. The performance tester's lens
As a performance tester, understanding what a server physically is keeps you from misdiagnosing bottlenecks:
-
Watch the
%stealMetric When Load Testing in the Cloud: - Always monitor%stealintop,vmstat 1, or your APM agent (Datadog/Prometheus). - If%steal> 5%, your load test results are not valid, because you're measuring how crowded your neighbors are, not the limits of your application. -
Run Load Tests at Different Times: - Run load tests in the morning, the afternoon, and at midnight. If throughput differs significantly on the same environment, noisy neighbor interference at the I/O or network layer is the likely cause.
-
Test Disk IOPS and Network in Isolation: - Use tools like
fioto benchmark disk IOPS andiperf3for network throughput before running the application load test. Make sure the bandwidth you get matches the instance SLA. -
Pick the Right Test Environment: - The Performance Testing environment (Staging/Perf) must have an overcommit ratio and instance type identical to Production (e.g. a Compute-Optimized Dedicated Instance, not a Burstable Shared T-Series).
7. Summary & Bridge to the Next Session
- Servers today come as Bare Metal (exclusive & consistent), Virtual Machine (flexible, but with a hypervisor & overcommit), and Container (very fast & sharing one kernel).
- Cheap cloud pricing is driven by overcommit; the risk is the Noisy Neighbor phenomenon and CPU Steal.
- Now we understand what one server really is. The next question: when one server can no longer handle the load, should we buy a much bigger server or buy ten small servers?
Continue to Session 8 — Bigger or More (Vertical vs Horizontal Scaling).