Level 1 · Multiplying Machines
Session 15: The Robot That Keeps Pods Alive: Kubernetes, HPA, & Probes
What if a container server died suddenly at three in the morning, and within 800 milliseconds its replacement was already serving users before a single engineer had a chance to wake up?
1. Imagine If
A factory has 10 coffee-making machines. Above those machines sits one automatic supervisor robot. The robot holds a checklist: "There must always be 10 machines running."
If machine number 4 chokes and dies, the robot doesn't wake the security guard or sound an alarm. Within seconds, it throws the broken machine in the trash, fires up a new machine number 11 from the spare warehouse, tastes the coffee to make sure the temperature is right, and then opens the door for customers.
2. What Actually Happens
This is the heart of Container Orchestration (Kubernetes / OpenShift): it's not humans managing servers one by one, but a state reconciliation Control Loop that constantly matches the Desired State against the Current State.
- Pod: - The smallest schedulable unit in Kubernetes. It contains one or more containers that share the same IP address and network namespace.
- ReplicaSet & Deployment:
- Makes sure the number of running pod replicas always matches the declaration (
replicas: 10). If 2 pods crash from OOM (Out Of Memory), the ReplicaSet immediately tells the Scheduler to create 2 new pods on available Nodes. - Liveness Probe vs Readiness Probe: - Liveness Probe: "Is this pod still alive?" If it fails (for example: a deadlocked thread that no longer responds), Kubernetes immediately kills (restarts) the pod. - Readiness Probe: "Is this pod ready to receive traffic yet?" While startup / model loading isn't finished, or if the CPU is at 100%, the pod is not killed; it's just removed from the Load Balancer / Service target list so users don't get 502/503 errors.
- HPA (Horizontal Pod Autoscaler): - Adds or removes pods dynamically based on CPU, Memory, or custom RPS metrics (for example: if average CPU > 70%, scale pods from 10 up to 30).
- Requests vs Limits:
- Requests: The guaranteed minimum capacity a worker Node must provide when scheduling the pod.
- Limits: The absolute cap on CPU/RAM usage. If a pod breaks through its Memory Limit, the Linux kernel immediately sends an
OOMKilledsignal (Exit Code 137).
3. What If We Try...
"Let's set the HPA to auto-scale all the way to 200 Pods the moment a flash sale starts!"
Why does this so often end in disaster at Level 2?
- Database Connection Storm: If each pod opens a pool of 20 database connections, 200 pods means 4,000 simultaneous connections slamming into a single primary database. The database locks up instantly (CPU 100%, disk I/O saturated), and the whole system goes completely down.
- Node Capacity Choke: If the physical worker nodes run out of RAM to fit 200 pods, the new pods get stuck in Pending status because the scheduler has run out of room.
- Adding pods just moves the bottleneck from the application layer to the data layer!
4. The Official Name
- Desired State vs Current State: Kubernetes' declarative reconciliation principle.
- HPA (Horizontal Pod Autoscaler): The controller that scales pod replicas automatically.
- Kube-Scheduler: The control plane component that places pods on the optimal worker node.
- Liveness & Readiness Probes: Pod health checking mechanisms at the container runtime level.
- OOMKilled (Exit Code 137): A pod forcibly killed for breaching its memory limit.
5. In Our World
- Zero-Downtime Rolling Update: Kubernetes starts 1 pod of the new version, waits for its Readiness Probe to succeed, and only then shuts down 1 pod of the old version, step by step.
- Cold Start Delay: A Java/Spring Boot or Python application that takes 45 seconds to start up can make the HPA too late to respond to a sudden traffic surge (spike traffic).
- Graceful Termination (
preStophook): Gives a pod time to finish the requests it's currently handling before it's shut down (avoiding 502 errors on in-flight connections).
6. The Performance Tester's Lens
Critical metrics and tests:
1. Autoscaling Reaction Time: How many seconds pass between a traffic surge and new pods actually receiving requests (Ready)?
2. Pod Startup Latency: The time a pod spends in the initialization phase before passing its Readiness Probe.
3. OOM Threshold Testing: Load tests that trigger memory leaks until pods get OOMKilled at their limit.
4. Connection Storm Impact: How big is the spike in DB connections when the HPA multiplies pods from minimum to maximum?
7. Question for the Next Round
If growing from 10 pods to 100 is as easy as changing one number in Kubernetes, why does our system so often become much slower after pods are added?
The answer opens the gate to LEVEL 2: THE DATABASE IS THE BOTTLENECK in Session 16: Pods Up, Database Down.