Level 5 · Life and Death of a Pod on OpenShift
Session 36: The Last Seconds of a Pod: Graceful Shutdown, SIGTERM, & the preStop Hook
What if every time our team deployed a new version, 300 users got a 502 error, even though every new pod was healthy and every probe was green?
1. Imagine If
A restaurant is closing at 10 PM. The waiters inside have been told: "From now on, no new guests. Finish serving the people who are eating, then go home."
The problem is that the sign out front still says OPEN. The person who has to flip that sign is at the end of the street and won't get there for another five minutes. During those five minutes, guests keep arriving, push the door, and find it already locked from the inside. They leave annoyed.
Nothing was wrong with the kitchen. Nothing was wrong with the waiters. What went wrong was the order: the door was locked before the sign was changed.
That is exactly what happens to pods on every rolling update, scale-down, and node drain when nobody manages their last seconds.
2. What Actually Happens
In Session 35 we covered how a pod is born and declared healthy through startup, readiness, and liveness probes. This session covers the other side: how a pod dies.
Pods get deleted for many reasons: oc rollout restart, a rolling update to a new image, an HPA scale-down, oc adm drain during node maintenance, eviction because a node is short on memory, or a manual oc delete pod. Whatever the trigger, the sequence is the same:
- The API server marks the pod
Terminatingand setsdeletionTimestamp. TheterminationGracePeriodSecondscountdown (default 30 seconds) starts right now. - Two processes run IN PARALLEL, not in sequence:
- (a) The kubelet path: the kubelet on the node runs the
preStophook (if one exists). Once the hook finishes, the kubelet sends SIGTERM to PID 1 inside the container. - (b) The network path: the endpoints controller removes the pod IP from the EndpointSlice. That change then propagates asynchronously to kube-proxy / iptables on every node, to the OpenShift router (HAProxy) that serves the Route, to any other ingress controller, and to the service mesh (Envoy sidecars) if you run one. - If the app finishes early, the container exits with code 0 and the pod disappears cleanly.
- If the countdown runs out, the kubelet sends SIGKILL. The process dies instantly with exit code 137, and every in-flight request is cut off mid-flight.
The key word is parallel. Path (a) can finish in milliseconds: SIGTERM arrives, the app immediately closes its listening socket. Path (b) takes time: one to several seconds on a small cluster, longer on a large cluster with hundreds of nodes and thousands of endpoints. The OpenShift HAProxy router also has its own reload interval.
During that gap, the router still believes the pod is alive and keeps sending it new requests. The result: 502 Bad Gateway, 503, or connection refused. This is a race condition, not an application bug.
Timeline without a preStop hook
| Time | Kubelet path | Network path | Impact |
|---|---|---|---|
| t=0 | Pod Terminating, SIGTERM sent |
EndpointSlice update begins | - |
| t=0.05 | App closes listener | kube-proxy on node A updated | - |
| t=0.5 | App finishing in-flight requests | HAProxy router not yet reloaded | New requests to pod: connection refused |
| t=1.5 | App exits 0 | kube-proxy on node C just updated | Requests via node C: 502 |
| t=3 | Pod gone | Every component updated | Errors stop |
Three seconds of errors per pod. Multiply by 20 pods in a single rolling update, and every deploy leaves a trail of 502s on the dashboard.
The fix: the preStop hook as a "polite pause"
The preStop hook runs before SIGTERM. If all it does is sleep 10, the app keeps serving normally for 10 seconds while the network path finishes propagating. By the time SIGTERM finally arrives, no router is sending new requests anymore.
apiVersion: apps/v1
kind: Deployment
metadata:
name: checkout-api
spec:
template:
spec:
terminationGracePeriodSeconds: 45 # sleep 10 + drain 20 + cleanup + margin
containers:
- name: app
image: registry.example.com/checkout-api:2.4.1
lifecycle:
preStop:
exec:
command: ["sleep", "10"]
# Native action (introduced in Kubernetes 1.29, on by default since 1.30):
# sleep:
# seconds: 10
Timeline with a preStop hook
| Time | Kubelet path | Network path | Impact |
|---|---|---|---|
| t=0 | Pod Terminating, preStop sleep 10 starts |
EndpointSlice update begins | App still serving |
| t=0.5 | Still sleeping | kube-proxy on some nodes updated | Existing requests still served |
| t=3 | Still sleeping | HAProxy router reloaded, all updated | No new requests arriving |
| t=10 | SIGTERM to PID 1 | - | App starts draining |
| t=10–25 | Finish in-flight, close pools | - | 0 errors |
| t=26 | Exit 0 | - | Pod gone cleanly |
| t=45 | (SIGKILL deadline, never reached) | - | - |
Note: the 45-second countdown already includes the preStop duration. If preStop takes 10 seconds and the app needs 20 seconds to drain, the default 30-second grace period is not enough. The app gets SIGKILLed at second 30.
Sizing rule
preStop sleep >= endpoint propagation time (measure it; usually 5-10 seconds)
sleep + longest request drain + cleanup < terminationGracePeriodSeconds
What the app must do on SIGTERM
- Stop accepting new connections.
- Finish in-flight requests.
- Close keep-alive connections: send a
Connection: closeheader on the next response, or close idle connections. Without this, clients and proxies using HTTP/1.1 keep-alive, HTTP/2, or gRPC stay "pinned" to the old pod. - Close DB/Redis connection pools cleanly, so the server isn't left holding zombie connections.
- Flush log and metric buffers.
- Exit 0.
# Spring Boot application.yaml
server:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 20s
Node.js does not do this automatically. You have to register process.on('SIGTERM', ...) and call server.close(). Gunicorn has graceful_timeout.
The PID 1 trap
If the Dockerfile uses shell form, for example ENTRYPOINT java -jar app.jar or command: ["sh", "-c", "java -jar app.jar"], PID 1 is sh, not Java. The shell does not forward SIGTERM to its child process. The app never learns it has been asked to stop, waits until the grace period runs out, and is killed by SIGKILL. Always exit 137, always 30 seconds late.
The fix: use exec form (ENTRYPOINT ["java", "-jar", "app.jar"]), add exec in the script (exec java -jar app.jar), or use a tiny init such as tini / dumb-init.
What about the readiness probe?
Many teams believe readiness must be made to fail on SIGTERM so the pod leaves the Service. That isn't needed for that purpose: a Terminating pod is automatically removed from the ready endpoints, whatever its probe says. What does not disappear automatically are long-lived connections that are already open. Those are the ones the app must close itself.
3. What If We Try...
"Easy, just set terminationGracePeriodSeconds: 300 to be safe!"
A long grace period does not make shutdown any cleaner. It only extends the deadline:
- Rollouts become painfully slow. If the app doesn't handle SIGTERM (the PID 1 trap), every pod waits the full 300 seconds. 20 pods with maxUnavailable: 1 means a rollout can take around 100 minutes.
- Node drains stall. oc adm drain waits for every pod on that node to finish. An OpenShift cluster upgrade that rolls node by node can drag on for hours.
- The autoscaler is slow to release capacity. Scaled-down pods linger in Terminating and keep holding resources.
- The race condition is still there. Without preStop, SIGTERM still arrives at second 0, and errors still happen between second 0 and second 3.
"Then skip preStop. Our app already handles SIGTERM quickly."
The faster the app stops, the wider the error window, because the listener closes before the router knows. The result is a small 502 spike that shows up on every deploy, too short to trigger an alert, but enough to fail real user transactions.
4. The Official Name
- Pod Termination Lifecycle: The sequence
Terminating→ preStop → SIGTERM → grace period → SIGKILL. - terminationGracePeriodSeconds: The total deadline (default 30 seconds) from the start of termination, including preStop.
- preStop Hook: A container lifecycle hook that runs before SIGTERM (
exec,httpGet, or the nativesleepaction since Kubernetes 1.29). - SIGTERM / SIGKILL: A stop request the app can handle, and a forced kill it cannot.
- Exit Code 137: 128 + 9 (SIGKILL). It can mean OOMKilled (Session 15) or an exceeded grace period. Check the
reasonfield to tell them apart. - EndpointSlice Propagation Delay: The gap between a pod being removed from endpoints and every kube-proxy / router knowing about it.
- Connection Draining: Finishing existing connections without accepting new ones.
5. In Our World
- Rolling updates on OpenShift:
oc rollout restart deployment/checkout-apishuts down old pods one by one. Without preStop, every dying pod contributes a few seconds of errors. This often shows up as "a small 502 spike every day at 2 PM", which is the deploy schedule. - The OpenShift HAProxy router: The router reads endpoint changes and then updates its configuration. On clusters with many Routes, that update can lag several seconds behind the EndpointSlice.
- How this differs from Session 9: In Session 9, graceful drain happened at the load balancer. The LB decided when a backend stopped receiving traffic. In Kubernetes there is no single central LB. There are dozens of components (kube-proxy on every node, the router, the mesh), each learning on its own, asynchronously. That's why the pod has to actively "wait".
- gRPC and HTTP/2: One connection carries thousands of requests. If the server doesn't send GOAWAY or close the connection, the client keeps using the old pod until it actually dies, and then every stream on that connection fails at once.
- Node drain during a cluster upgrade: Every pod on the node gets evicted at the same time. Without a
PodDisruptionBudget, several replicas of the same Deployment can die together.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: checkout-api-pdb
spec:
minAvailable: 80%
selector:
matchLabels:
app: checkout-api
6. The Performance Tester's Lens
Graceful shutdown almost never gets tested, because load tests usually run against a cluster that's standing still. Change that.
- Deploy-under-load test: Run steady load (for example 500 RPS for 15 minutes), and at minute 5 run
oc rollout restart deployment/checkout-api. Count 502s, 503s, and connection resets per deploy. Target: 0. - Scale-down test: Let the HPA reduce replicas when load drops mid-test. Errors during scale-down are just as harmful as errors during a deploy.
- Node drain test: Run
oc adm drain <node> --ignore-daemonsets --delete-emptydir-datawhile load is running. Measure errors and whether the PDB is respected. - Exit code audit:
oc get pod <pod> -o jsonpath='{.status.containerStatuses[0].lastState.terminated}'or check the events. Exit 137 with a reason other thanOOMKilledmeans the grace period was exceeded: likely the PID 1 trap or a drain that takes too long. - Measure durations: Time from SIGTERM to exit (from app logs), and total time the pod spends in
Terminating. If it's always exactly 30 seconds, the app is almost certainly not receiving SIGTERM. - Measure endpoint propagation: Send continuous requests through the Route, delete one pod, and record when the last request lands on that pod. That number is the basis for sizing the preStop
sleep. - Long-lived connections: For keep-alive or gRPC clients, check whether connections move to new pods, or whether all the errors pile up in the old pod's last second.
A useful report doesn't say "p95 = 180 ms". It says "every deploy produced 0 errors, 137 never appeared, pods spent an average of 18 seconds in Terminating".
7. Question for the Next Round
The old pod has left politely. But during a rolling update or an HPA scale-up, dozens of new pods are born almost at once, and each one immediately opens its own connection pool to Redis and the database. One pod opens 10 connections. What about 200 pods?
The answer is in Session 37: One Pod 10 Connections, 200 Pods 2,000 Connections.