skip to content

How do Spring Boot's availability states behave across the application lifecycle, and how does this interact with graceful shutdown in Kubernetes?

level: seniorimportance: should knowfreq 35%

answer

  1. liveness CORRECT early; readiness ACCEPTING only when fully started
  2. shutdown -> readiness REFUSING_TRAFFIC (drain)
  3. server.shutdown=graceful + timeout-per-shutdown-phase
  4. SIGTERM vs endpoint removal race -> preStop sleep
  5. liveness stays CORRECT during planned stop

basics

~20 s

At startup liveness becomes CORRECT early; readiness becomes ACCEPTING_TRAFFIC only when the app is fully ready. On shutdown Spring flips readiness to REFUSING_TRAFFIC so Kubernetes drains traffic before the process stops, which pairs with graceful shutdown.

solid answer

~40 s

Spring emits availability changes tied to lifecycle events. Early in boot, `LivenessState` is set to `CORRECT`. `ReadinessState` reaches `ACCEPTING_TRAFFIC` only once the context is fully started — around `ApplicationReadyEvent`. On a SIGTERM-driven shutdown, Spring publishes `ReadinessState.REFUSING_TRAFFIC` at the start of graceful shutdown, so the readiness probe goes OUT_OF_SERVICE and Kubernetes stops routing new requests while in-flight ones drain. To make this work you enable `server.shutdown=graceful` and set a `spring.lifecycle.timeout-per-shutdown-phase`. Crucially, Kubernetes sends SIGTERM and removes the endpoint asynchronously, so you also configure a `preStop` sleep (or terminationGracePeriodSeconds) so the readiness change propagates before the container dies. Liveness stays CORRECT throughout normal shutdown — you don't want a restart during a planned stop.

code

yaml · 24 lines
yaml
server:
  shutdown: graceful
spring:
  lifecycle:
    timeout-per-shutdown-phase: 30s
management:
  endpoint:
    health:
      probes:
        enabled: true

# Kubernetes pod spec
# spec:
#   terminationGracePeriodSeconds: 45
#   containers:
#     - name: app
#       lifecycle:
#         preStop:
#           exec:
#             command: ["sh", "-c", "sleep 10"]
#       readinessProbe:
#         httpGet: { path: /actuator/health/readiness, port: 8080 }
#         periodSeconds: 5
#         failureThreshold: 1

go deeper

for a junior

Know that readiness gates startup and drains on shutdown.

for a middle

Enable graceful shutdown and map the readiness probe; know the states across lifecycle.

for a senior

Explain the SIGTERM/endpoint-removal race and the preStop + timeout tuning that mitigates it.

for a principal

Standardize graceful-shutdown timing across services and reason about zero-downtime rollout guarantees end to end.

## Lifecycle timeline of availability states **Startup** - `LivenessState.CORRECT` is published early — the application process is alive and its internal state is intact well before it can serve traffic. - `ReadinessState.ACCEPTING_TRAFFIC` is published only when the application is fully initialized and ready to serve — corresponding to the `ApplicationReadyEvent` / the web server having started. Until then readiness reports OUT_OF_SERVICE, so Kubernetes withholds traffic during warm-up. **Normal running** - Both remain CORRECT / ACCEPTING_TRAFFIC unless your code (or a failure) publishes a change. **Shutdown (SIGTERM)** - When graceful shutdown begins, Spring publishes `ReadinessState.REFUSING_TRAFFIC`. The readiness probe now reports OUT_OF_SERVICE. Kubernetes' endpoint controller removes the pod from Service endpoints so no new requests are routed. - The web server then stops accepting new requests and lets in-flight requests complete, bounded by `spring.lifecycle.timeout-per-shutdown-phase`. - Liveness remains `CORRECT` — a planned shutdown is not a broken state, and you don't want a restart. ## Enabling graceful shutdown ```yaml server: shutdown: graceful # default is 'immediate' spring: lifecycle: timeout-per-shutdown-phase: 30s ``` ## The Kubernetes race condition (key gotcha) When Kubernetes terminates a pod it does two things roughly concurrently: (1) sends SIGTERM to your container, and (2) tells the endpoint controller to remove the pod from Service endpoints. These are **not synchronized**, and endpoint propagation to every kube-proxy takes time. If your app stops immediately on SIGTERM, some clients may still be routed to it and get connection errors. Spring flipping readiness to REFUSING_TRAFFIC helps, but the probe interval + endpoint propagation still lags. The standard mitigation is a **preStop hook** that sleeps a few seconds so the readiness change and endpoint removal propagate before the process exits: ```yaml lifecycle: preStop: exec: command: ["sh", "-c", "sleep 10"] terminationGracePeriodSeconds: 45 # must exceed preStop + shutdown timeout ``` ## Why liveness must not flip during shutdown If liveness went BROKEN on shutdown, Kubernetes might restart the container mid-drain, defeating the graceful stop. Spring keeps it CORRECT. ## Gotchas summary - Readiness gates warm-up on startup automatically — don't add manual sleeps for that. - `terminationGracePeriodSeconds` must be larger than preStop sleep + shutdown timeout, or the pod is SIGKILLed mid-drain. - The readiness OUT_OF_SERVICE only helps if your readiness probe is actually mapped and its `periodSeconds`/`failureThreshold` are tight enough to notice quickly.

  • Why add a preStop sleep if Spring already flips readiness to REFUSING_TRAFFIC on shutdown?
    Because Kubernetes sends SIGTERM and removes the pod from Service endpoints asynchronously, and endpoint removal must propagate to every kube-proxy. A preStop sleep gives that propagation (and the readiness flip) time to take effect before the process actually exits, avoiding requests routed to a dying pod.
  • Why doesn't Spring flip liveness to BROKEN during shutdown?
    A planned shutdown isn't an unrecoverable fault. Flipping liveness would let Kubernetes restart the container mid-drain, defeating graceful shutdown. Liveness stays CORRECT.

saying these in an interview costs you the question

  • Saying readiness is ACCEPTING_TRAFFIC from the very first moment of startup (it waits for full readiness)
  • Claiming graceful shutdown alone fully solves the traffic-routing race without preStop/endpoint propagation
  • Setting terminationGracePeriodSeconds smaller than preStop + shutdown timeout
  • Expecting liveness to change during a normal shutdown

context