Org Skills

Probes and graceful shutdown

Supporting material for service-operability. Agents load it on demand; it ships inside the skill folder.

RawSource

Verified 2026-09 against https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/.

The three probes

Probe Question On failure Must be
Liveness Is this process wedged? Container restarted Local, cheap, no dependencies
Readiness Should it get traffic now? Removed from endpoints Cheap, cached, dependency-aware
Startup Has it finished booting? Container restarted Generous; disables the other two until it passes

Example configuration:

startupProbe:
  httpGet: { path: /health/startup, port: 8080 }
  periodSeconds: 5
  failureThreshold: 30          # allows 150s to boot
livenessProbe:
  httpGet: { path: /health/live, port: 8080 }
  periodSeconds: 10
  failureThreshold: 3
readinessProbe:
  httpGet: { path: /health/ready, port: 8080 }
  periodSeconds: 5
  failureThreshold: 2

What each dependency check should do

Dependency Readiness check Not this
Relational database Acquire a pooled connection and run a trivial statement, with a short timeout A real query, or anything that touches application tables
Cache A ping, only if the service cannot serve without it Treating an optional cache as required
Message broker Connection is established and the consumer is subscribed Publishing a test message
Downstream HTTP API Usually nothing Calling it — you inherit its outage and can create a cascade

The guiding rule: readiness reflects whether this instance can serve, not whether the whole system is healthy. A service that degrades gracefully without a dependency should stay ready. Otherwise every instance goes unready at once and a partial outage becomes total.

Cache the result for one or two seconds. With a 5-second period across many replicas, an uncached readiness check becomes meaningful load on the database it is checking.

Shutdown sequence with timings

Assume a 30-second termination grace period:

t Step
0s SIGTERM received
0s Readiness starts failing; requests continue to be served
0–5s Load balancer propagation window — new requests still arrive, and must succeed
5s Stop accepting new connections
5–25s Finish in-flight requests; stop consumer prefetch, finish or nack the current message
25s Close database pools, flush telemetry
~26s Exit 0

The propagation window at 0–5s is the step most often missed. Endpoint removal is eventually consistent: closing the listener the moment SIGTERM arrives guarantees a burst of connection errors on every deploy.

Sizing: the grace period must exceed the longest drain, which is bounded by your slowest normal request. If a request can take 60 seconds, either the grace period exceeds that, or long work moves to a queue where it can be resumed instead of finished.

Consumers and background work

  • Stop prefetching first, so no new messages are claimed.
  • Finish the in-flight message, or nack it for redelivery — do not exit holding an unacknowledged message with a long visibility timeout, or it stays invisible until that expires.
  • A scheduled job mid-run either completes within the window or records progress and stops at a resumable point.

Verifying it

The only test that counts: run a steady request load, do a rolling deploy, and confirm zero failed requests. Then repeat with a consumer running and check no message is lost or processed twice. Reading the code is not sufficient — the propagation window and the grace period only show up under real traffic.