Level 5 · Life and Death of a Pod on OpenShift
Session 38: All Eggs in One Basket: topologySpreadConstraints, Anti-Affinity, & PodDisruptionBudget
What if we were already running 6 pod replicas, HPA enabled, clean probes, perfect graceful shutdown — but 5 of those 6 pods turned out to be sleeping on the same worker node, and that node just died?
1. Imagine If
A farmer has 60 eggs and 3 baskets. He's proud: "I have plenty of eggs. If one breaks, I still have 59." But the first basket sits closest to the door and has the most room, so every time he adds a new egg, his hand drifts into that basket. In the end there are 50 eggs in the first basket, 7 in the second, and 3 in the third.
One morning the first basket falls off the table. The farmer doesn't lose one egg. He loses 83% of his eggs in a single second.
The number of eggs was never the problem. The problem was where they were placed. In OpenShift, the eggs are pods, the baskets are worker nodes or zones, and the hand that automatically picks the roomiest basket is the kube-scheduler.
2. What Actually Happens
In Session 13 we learned that fake redundancy is a SPOF in disguise. In Session 14 we went up a floor, to the level of buildings and data centers. This session brings it back down to the place most often forgotten: pod placement inside a single cluster.
-
The scheduler promises no spread: - The kube-scheduler's default scoring does lean toward spreading (plugins like the default
PodTopologySpreadandNodeResourcesBalancedAllocation), but that is only a scoring preference, not a guarantee. The node with the most free resources often wins again and again. - The result: 6 replicas can land 5 on one node and 1 on another. On paperreplicas: 6is satisfied. In terms of resilience, we have almost no redundancy. -
The scheduler never moves running pods: - Imagine worker-2 is drained for maintenance. All its pods move to worker-1 and worker-3. An hour later worker-2 is back
Ready, but empty. Pods don't drift back on their own. - The scheduler only acts when a new pod is created. Rebalancing running pods requires a descheduler (in OpenShift: the Kube Descheduler Operator) or a fresh rollout.
BEFORE DRAIN AFTER worker-2 RETURNS
worker-1 [P][P] worker-1 [P][P][P]
worker-2 [P][P] --> worker-2 [ ][ ][ ] <- empty forever
worker-3 [P][P] worker-3 [P][P][P]
-
topologySpreadConstraints: explicit spread rules:
yaml apiVersion: apps/v1 kind: Deployment metadata: name: checkout-api spec: replicas: 6 selector: matchLabels: app: checkout-api template: metadata: labels: app: checkout-api spec: topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: DoNotSchedule # hard: zones must be balanced labelSelector: matchLabels: app: checkout-api matchLabelKeys: - pod-template-hash # count per ReplicaSet during rolling updates - maxSkew: 1 topologyKey: kubernetes.io/hostname whenUnsatisfiable: ScheduleAnyway # soft: try to balance across nodes labelSelector: matchLabels: app: checkout-api-maxSkew: The maximum allowed difference in matching pod count between the most loaded and the least loaded topology domain. -topologyKey: The node label that defines a "basket".kubernetes.io/hostnamemeans per node,topology.kubernetes.io/zonemeans per zone. -whenUnsatisfiable:DoNotSchedule= hard rule, the pod staysPendingif it can't be satisfied.ScheduleAnyway= soft rule, only a scoring preference. -labelSelector: Must match the pods' own labels. One typo and the constraint counts zero pods and does nothing, with no error. - Optional:minDomains(minimum number of domains assumed to exist),matchLabelKeys(e.g.pod-template-hashso spread is counted per ReplicaSet during a rolling update, not a mix of old and new pods),nodeAffinityPolicy/nodeTaintsPolicy(whether nodes that fail affinity/taints count as domains). -
maxSkew math, a concrete example (3 zones):
| Replicas | Spread | Allowed with maxSkew 1? | Reason |
|---|---|---|---|
| 6 | 2 / 2 / 2 | Yes | difference 0 |
| 7 | 3 / 2 / 2 | Yes | difference 1 |
| 7 | 4 / 2 / 1 | No | difference 3 |
| 6 | 4 / 1 / 1 | No | difference 3 |
Important: skew is evaluated at scheduling time only. If a node dies afterward and the spread becomes lopsided, the constraint doesn't fix it.
-
podAntiAffinity: "at most one":
yaml affinity: podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchLabels: app: checkout-api topologyKey: kubernetes.io/hostname- Therequiredform means at most 1 pod per node. If there are more replicas than nodes, the rest stayPending. -preferredDuringSchedulingIgnoredDuringExecutionis the soft form. - The difference from spread: anti-affinity says "never together", spread says "balanced". Anti-affinity is also computationally expensive for the scheduler in large clusters with thousands of pods. -
PodDisruptionBudget (PDB): a brake on intentional disruption:
yaml apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: checkout-api-pdb spec: maxUnavailable: 1 # or minAvailable: 5 selector: matchLabels: app: checkout-api- Protects against voluntary disruption:oc adm drain, cluster upgrades, and OpenShift MachineConfig rollouts that reboot nodes one by one. - Does not protect against node crashes, kernel panics, or a zone going dark. A PDB is only honored by processes that ask permission through the Eviction API.
3. What If We Try...
"Just bump replicas to 10, surely that's safe!"
Without spread rules, all 10 replicas can land on the 2 nodes that happen to have the most room. We pay for 10 pods but only own 2 baskets. One node dies, half the capacity disappears at once.
"Fine, put maxSkew: 1 with DoNotSchedule everywhere!"
Now the problem flips. During a flash sale, HPA wants to go from 6 to 9 pods. Zone C turns out to be full (resource requests exhausted). Zones A and B can each take only one extra pod (3/3/2). The next pod for A or B would push skew to 2, so it's held back too, even though A and B still have room. The result:
zone-a [P][P][P] zone-b [P][P][P] zone-c [P][P] FULL
new pod: Pending
Events: 0/9 nodes are available: 3 node(s) didn't match pod topology
spread constraints, 3 Insufficient cpu ...
HPA reports desired replicas 9, the business dashboard sees latency climbing, and nobody notices that scale-up silently failed because pods are Pending. A rule that's too rigid turns resilience into a capacity ceiling.
That's why the common combination is zone hard + hostname soft (when each zone has generous capacity), or zone soft + hostname hard (when there are few zones but many nodes). Capacity decides, not taste.
4. The Official Name
- Pod Topology Spread Constraints: Rules for spreading pods across topology domains (
maxSkew,topologyKey,whenUnsatisfiable,labelSelector). - Topology Domain: One value of a topology label, such as one node or one zone.
- Inter-Pod Anti-Affinity: A rule that keeps certain pods from being placed together in one domain.
- PodDisruptionBudget (PDB): A limit on how many pods may be unavailable due to voluntary disruption.
- Voluntary vs Involuntary Disruption: Intentional disruption (drain, upgrade) vs unintentional (crash, power loss).
- Descheduler: A component that evicts pods so they get rescheduled into a healthier spread.
- N+1 Capacity: Enough capacity to survive peak load when one domain is lost.
5. In Our World
- Failure math: 6 pods, 4 on node A. Node A dies, 67% of capacity is gone at once. The 2 remaining pods carry 3x load, latency spikes, readiness probes fail, then liveness probes (Session 35) start killing pods that are merely exhausted. Restarted pods come back to 3x load and die again. That's a cascading failure.
| Spread | Domain lost | Capacity lost | Load per surviving pod |
|---|---|---|---|
| 4 / 1 / 1 (per node) | node A | 67% | 3x |
| 2 / 2 / 2 (per zone) | zone A | 33% | 1.5x |
- N+1 per zone: With 3 zones, when one zone is lost the two remaining zones must hold 100% of peak. That means every zone must be able to handle 50% of peak load, not 33%. A tidy spread without headroom still falls over.
- Data needs separate baskets too: A Redis primary and replica (Session 37), or a database primary and standby, sleeping on the same node is the classic Session 13 SPOF in a new costume. StatefulSets need spread rules just as strict.
- A PDB that's too strict:
maxUnavailable: 0, orminAvailableequal to the replica count, makesoc adm drainhang forever. The OpenShift cluster upgrade stalls halfway, the MachineConfigPool sits inUpdating, and the platform team has to delete the PDB by hand. - Cross-zone cost: Spreading pods across zones means some requests hop zones, adding network latency and, in the cloud, data transfer cost. Topology-aware routing (Topology Aware Hints) can keep traffic in the same zone while local capacity is sufficient.
6. The Performance Tester's Lens
Before and during testing:
1. Check placement before every test: Don't trust replicas: 6. Look at where the pods actually live.
bash
oc get pods -l app=checkout-api -o wide --no-headers | awk '{print $7}' | sort | uniq -c
oc get nodes -L topology.kubernetes.io/zone
2. Node/Zone Loss Test: While a load test runs at steady load, oc adm drain one node (or cordon and stop every node in one zone). Measure capacity lost, the p95/p99 spike, error rate, and recovery time until throughput is back to normal.
3. Scale-up Under Constraint: Trigger HPA scale-up at peak. Count Pending pods and look for didn't match pod topology spread constraints or didn't match pod anti-affinity rules events.
bash
oc get events --field-selector reason=FailedScheduling
4. PDB Drain Test: Confirm a drain can still proceed under load without breaching the SLO, and confirm the PDB doesn't block the drain forever.
5. N+1 Validation: Run peak load with one zone deliberately taken out. If the two remaining zones can't hold it, the tidy spread is cosmetic.
6. Rebalance Check: After a node returns Ready, check the spread again. If it's lopsided, log it as a risk until a descheduler or rollout fixes it.
7. Question for the Next Round
Level 5 has followed the full life of a pod: where it's placed (Session 38), declared alive and ready (Session 35), carrying connections to Redis and the database (Session 37), and dying gracefully (Session 36). So here's the question: in the system you're testing right now, which of those four stages is most likely to break first when one node disappears in the middle of peak load, and how will you prove it instead of guessing?
This is the final session of the SYSTEM UNDER PRESSURE series. The next round isn't a new video. It's your own system. Go back to Session 31 to read the design doc and mark every single basket, then to Session 29 to design the load test and chaos test that knock that basket over on purpose, before production does it for you.