UNDER PRESSURE

Level 5 · Life and Death of a Pod on OpenShift

Session 35: The Probe That Kills Healthy Pods: Liveness, Readiness, & Startup Probes

What if the pod that killed itself in the middle of a load test wasn't sick at all, and the one pulling the trigger was the health check we wrote ourselves?

Session 35 / 3810 min read

1. Imagine If

A hospital has a duty nurse who taps a patient on the shoulder every 20 seconds and asks, "Are you still conscious?" The rule is simple: if the patient doesn't answer within 5 seconds three times in a row, the nurse calls the resuscitation team and the patient is put under again from scratch.

The problem is that the patient just came out of major surgery. He needs a full minute to wake up from anesthesia. The nurse starts asking at second 15, then second 35, then second 55. Three times, no answer. The resuscitation team arrives, the patient is sedated again, and the cycle repeats. The patient was actually healthy. He just wasn't given time.

Worse still: one day the nurse gets a new rule, "Also ask whether the hospital elevator is working." When the elevator jams, every patient on every floor is considered unconscious and sedated again at the same time. The elevator doesn't get any faster. The hospital is paralyzed instead.

That is a probe in OpenShift: a lifesaving tool that, when misconfigured, becomes a machine for killing healthy pods.


2. What Actually Happens

In Session 15 we met liveness and readiness probes briefly. Now we open up the machine. Probes are run by the kubelet on the node where the pod lives, not by the control plane. The kubelet calls the container's endpoint periodically and decides the pod's fate based on the result.

Three probe types, three different consequences:

Probe Question If it fails
livenessProbe "Is this process stuck/deadlocked?" Container is restarted by the kubelet
readinessProbe "Is this pod ready to receive traffic?" Pod is removed from Service endpoints, not restarted
startupProbe "Has the application finished booting?" Liveness & readiness are held back until it succeeds; if the budget runs out, the container is restarted

Check mechanisms (handlers): - httpGet: an HTTP request to a path/port; status 200-399 counts as success. - tcpSocket: success if a TCP connection can be opened. - exec: runs a command inside the container; exit code 0 = success. Every check forks a new process, so it gets expensive with short periods and many pods. - grpc: calls the gRPC Health Checking Protocol.

Parameters that control timing:

Parameter Meaning
initialDelaySeconds Wait after container start before the first probe
periodSeconds Interval between probes
timeoutSeconds How long a single probe waits for an answer; exceeding it = failure
failureThreshold Number of consecutive failures before it counts as failed
successThreshold Consecutive successes needed to count as recovered; must be 1 for liveness and startup

Case study. This is a real-world configuration we often find in Deployments:

livenessProbe:
  httpGet:
    path: /healthz
    port: 8080
  initialDelaySeconds: 15
  periodSeconds: 20
  timeoutSeconds: 5
  failureThreshold: 3
readinessProbe:
  httpGet:
    path: /ready
    port: 8080
  initialDelaySeconds: 5
  periodSeconds: 10
  timeoutSeconds: 3
  failureThreshold: 3

Looks reasonable at a glance. Now do the math.

Liveness timeline when the app has never answered (slow startup):

Time Event
t=0s Container starts, JVM begins booting
t=15s Liveness probe #1 → fail (1/3)
t=35s Liveness probe #2 → fail (2/3)
t=55s Liveness probe #3 → fail (3/3) → kubelet kills the container
t=55s+ Restart #1; in the worst case (jitter, 5s timeout per probe) it lands closer to ~75s

In other words: if the app needs more than ~55 seconds to become ready, it gets killed before it ever gets to live. On a developer laptop Spring Boot starts in 25 seconds. On a busy node, with a 500m CPU limit and CFS throttling, the same startup can take 70-90 seconds. Result: kill, restart, kill again. The kubelet waits longer and longer between restarts (exponential backoff: 10s, 20s, 40s, 80s, ... capped at 5 minutes). The pod status becomes CrashLoopBackOff, even though the application has no bug.

Liveness mid-life. If a running app suddenly deadlocks, detection takes between (failureThreshold-1) × periodSeconds and failureThreshold × periodSeconds + timeoutSeconds, which is roughly 40-65 seconds. During that time a zombie pod is still considered alive.

Readiness mid-life. When /ready starts failing, the pod is only removed from endpoints after 3 failures × 10 seconds, which is roughly 20-30 seconds (plus a 3-second timeout). Throughout that window, the Service keeps sending traffic to a pod that can no longer cope. In load test results, this shows up as a clump of errors or a latency spike lasting about as long as that window.


3. What If We Try...

"Easy, we'll make /healthz thorough: check Redis, check the database, check the payment gateway. If any of them is down, the pod is unhealthy."

Sounds diligent. It's the fastest way to turn a small glitch into a total outage.

  1. Redis slows down, say because of a single KEYS * command or a replica failover. Latency rises from 1ms to 6 seconds.
  2. /healthz on every pod waits on Redis and blows past timeoutSeconds: 5. Every pod fails liveness at almost the same time.
  3. After 3 failures, the kubelet restarts every pod. Application capacity drops to zero.
  4. Restarting does nothing to fix Redis. The problem lives outside the pod.
  5. When Redis recovers, dozens of pods boot together and open new connections all at once: a connection storm (see Session 16). Redis is under pressure again, probes fail again. This is the cascading failure pattern we covered in Session 25.

"Fine, then we'll move the dependency checks to /ready."

Better, but still dangerous. If every pod goes unready because one shared dependency is down, the Service has zero endpoints. The OpenShift router immediately returns 503 for every request, including endpoints that don't need Redis at all. A more sensible approach: /ready checks pod-local things (thread pool available, cache warm-up finished, this pod's own connection pool), and the application degrades gracefully (fallback, circuit breaker) when a dependency is in trouble.

"If startup is slow, just bump initialDelaySeconds to 120."

That patches the symptom. Every time the pod starts, including restarts after a real deadlock, liveness does nothing for 2 minutes. And when the node is so busy that startup takes 130 seconds, the problem is back.

Extra trap: a 5-second timeout during a load test. At peak load, the container hits its CPU limit and gets CFS throttled. The HTTP thread for /healthz queues behind business requests. The app is healthy, just slow, but /healthz answers in 6 seconds. Liveness fails, and the pod is restarted right at the peak. Capacity drops, the remaining pods carry more load, they get throttled too, and they fail probes too. This is a death spiral triggered by a health check.


4. The Official Name

  • Liveness Probe: Detects a stuck process; failure = container restarted.
  • Readiness Probe: The traffic gate; failure = pod leaves Service endpoints.
  • Startup Probe: Protects the boot phase; holds back liveness & readiness until the app has started.
  • CrashLoopBackOff: The status when a container keeps dying and the kubelet delays restarts with exponential backoff (max 5 minutes).
  • CFS Throttling: CPU limiting by the Linux Completely Fair Scheduler when a container uses up its limits.cpu quota.
  • Cascading Failure / Correlated Restart: Many pods failing together because of one shared cause.
  • Endpoints / EndpointSlice: The list of Ready pod IPs that a Service targets.

A safer revised configuration:

startupProbe:
  httpGet:
    path: /healthz
    port: 8080
  periodSeconds: 10
  failureThreshold: 30      # startup budget of up to 300 seconds
livenessProbe:
  httpGet:
    path: /healthz          # local checks only: event loop/threads alive, no Redis/DB
    port: 8080
  periodSeconds: 10
  timeoutSeconds: 5         # leave room for jitter under CPU throttling
  failureThreshold: 6       # ~60 seconds of continuous failure before restart
readinessProbe:
  httpGet:
    path: /ready            # pod-local readiness; dependencies handled gracefully
    port: 8080
  periodSeconds: 5
  timeoutSeconds: 3
  failureThreshold: 3       # leaves endpoints within ~10-15 seconds

Why this is better: - startupProbe gives a 300-second boot budget, but stops as soon as it succeeds. Fast startups don't wait, slow startups don't get killed. - initialDelaySeconds is no longer needed; liveness only kicks in after the startup probe succeeds. - Liveness is more patient (period 10 × threshold 6), so one or two slow probes from throttling don't trigger a restart, but a real deadlock is still caught in about a minute. - Readiness is more responsive (period 5), so an overwhelmed pod leaves rotation quickly, and comes back quickly once it recovers.


5. In Our World

  • Java/Spring Boot on a crowded node: A 30-second startup in staging becomes 80 seconds in production when 20 pods roll out together and all fight for CPU. Without a startup probe, the rolling update stalls in CrashLoopBackOff.
  • Spring Boot Actuator: /actuator/health/liveness and /actuator/health/readiness separate the health groups. A common trap: Redis/DB health indicators ending up in the liveness group by default in older configurations.
  • Checking on OpenShift:
  • oc get pods → a RESTARTS column that keeps climbing.
  • oc describe pod <name> → the Events section: Liveness probe failed: Get "http://10.128.2.14:8080/healthz": context deadline exceeded.
  • oc get events --field-selector reason=Unhealthy → every probe failure in the namespace.
  • oc get endpoints <service> → how many pods are actually receiving traffic right now.
  • oc logs <pod> --previous → logs from the container before it was killed.
  • Rolling updates: A readiness probe that passes too early (e.g. only checking that the port is open) lets new pods receive traffic before caches and connection pools are ready, producing a latency spike on every deploy.

6. The Performance Tester's Lens

Probes are part of the system under test, not just infrastructure config. A wrong probe can be the main source of errors in a load test result.

Metrics to watch during the test:

Metric Source Danger signal
Restart count oc get pods (RESTARTS column) Rises during the load test with no OOMKilled
Probe events oc get events (Liveness probe failed, Readiness probe failed) Appear in clusters at peak load
Time-to-ready Container start → Ready Much longer on a busy node than an idle one
Endpoint count oc get endpoints / kube_endpoint_address_available metric Flapping up and down, or dropping to 0
Error rate per time window Load test tool results Clumps of errors lasting ~20-30 seconds that line up with readiness events

Must-run test scenarios: 1. Load test at the CPU limit: Push load until pods hit limits.cpu. Watch for liveness-triggered restarts. Restarts with no OOMKilled at peak load almost certainly mean the probe is too strict. 2. Chaos on a dependency: Slow Redis down (latency injection) or stop it briefly. Verify the pods are not restarted; the only acceptable outcomes are degradation or controlled errors. 3. Cold start on a busy node: Scale from 2 to 20 pods while load is high. Measure time-to-ready and make sure no pod lands in CrashLoopBackOff. 4. A real deadlock: Simulate a stuck thread. Measure how many seconds until liveness restarts the pod, and how many errors happen in that window. 5. Time correlation: Overlay probe event timestamps on the error rate graph. If the patterns match, the errors come from the probes, not the code.

A question to bring to the team during review: "What does our /healthz check? If Redis is 10 seconds slow, do our pods get restarted?"


7. Question for the Next Round

We've tamed the probes: pods no longer get killed for no reason. But pods will still die, whether from a rolling update, a scale down, or a node drain. When OpenShift decides a pod has to stop, what happens to the requests it is currently processing? And why does our load test always show a row of 502 errors exactly at every deploy, even though every probe is green?

The answer is in Session 36: The Last Seconds of a Pod, covering SIGTERM, terminationGracePeriodSeconds, and the preStop hook.