How do Spring Boot's availability states behave across the application lifecycle, and how does this interact with graceful shutdown in Kubernetes?
answer
- liveness CORRECT early; readiness ACCEPTING only when fully started
- shutdown -> readiness REFUSING_TRAFFIC (drain)
- server.shutdown=graceful + timeout-per-shutdown-phase
- SIGTERM vs endpoint removal race -> preStop sleep
- liveness stays CORRECT during planned stop
basics
~20 sAt startup liveness becomes CORRECT early; readiness becomes ACCEPTING_TRAFFIC only when the app is fully ready. On shutdown Spring flips readiness to REFUSING_TRAFFIC so Kubernetes drains traffic before the process stops, which pairs with graceful shutdown.
solid answer
~40 sSpring emits availability changes tied to lifecycle events. Early in boot, `LivenessState` is set to `CORRECT`. `ReadinessState` reaches `ACCEPTING_TRAFFIC` only once the context is fully started — around `ApplicationReadyEvent`. On a SIGTERM-driven shutdown, Spring publishes `ReadinessState.REFUSING_TRAFFIC` at the start of graceful shutdown, so the readiness probe goes OUT_OF_SERVICE and Kubernetes stops routing new requests while in-flight ones drain. To make this work you enable `server.shutdown=graceful` and set a `spring.lifecycle.timeout-per-shutdown-phase`. Crucially, Kubernetes sends SIGTERM and removes the endpoint asynchronously, so you also configure a `preStop` sleep (or terminationGracePeriodSeconds) so the readiness change propagates before the container dies. Liveness stays CORRECT throughout normal shutdown — you don't want a restart during a planned stop.
code
yaml · 24 linesserver:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 30s
management:
endpoint:
health:
probes:
enabled: true
# Kubernetes pod spec
# spec:
# terminationGracePeriodSeconds: 45
# containers:
# - name: app
# lifecycle:
# preStop:
# exec:
# command: ["sh", "-c", "sleep 10"]
# readinessProbe:
# httpGet: { path: /actuator/health/readiness, port: 8080 }
# periodSeconds: 5
# failureThreshold: 1go deeper
Know that readiness gates startup and drains on shutdown.
Enable graceful shutdown and map the readiness probe; know the states across lifecycle.
Explain the SIGTERM/endpoint-removal race and the preStop + timeout tuning that mitigates it.
Standardize graceful-shutdown timing across services and reason about zero-downtime rollout guarantees end to end.
## Lifecycle timeline of availability states **Startup** - `LivenessState.CORRECT` is published early — the application process is alive and its internal state is intact well before it can serve traffic. - `ReadinessState.ACCEPTING_TRAFFIC` is published only when the application is fully initialized and ready to serve — corresponding to the `ApplicationReadyEvent` / the web server having started. Until then readiness reports OUT_OF_SERVICE, so Kubernetes withholds traffic during warm-up. **Normal running** - Both remain CORRECT / ACCEPTING_TRAFFIC unless your code (or a failure) publishes a change. **Shutdown (SIGTERM)** - When graceful shutdown begins, Spring publishes `ReadinessState.REFUSING_TRAFFIC`. The readiness probe now reports OUT_OF_SERVICE. Kubernetes' endpoint controller removes the pod from Service endpoints so no new requests are routed. - The web server then stops accepting new requests and lets in-flight requests complete, bounded by `spring.lifecycle.timeout-per-shutdown-phase`. - Liveness remains `CORRECT` — a planned shutdown is not a broken state, and you don't want a restart. ## Enabling graceful shutdown ```yaml server: shutdown: graceful # default is 'immediate' spring: lifecycle: timeout-per-shutdown-phase: 30s ``` ## The Kubernetes race condition (key gotcha) When Kubernetes terminates a pod it does two things roughly concurrently: (1) sends SIGTERM to your container, and (2) tells the endpoint controller to remove the pod from Service endpoints. These are **not synchronized**, and endpoint propagation to every kube-proxy takes time. If your app stops immediately on SIGTERM, some clients may still be routed to it and get connection errors. Spring flipping readiness to REFUSING_TRAFFIC helps, but the probe interval + endpoint propagation still lags. The standard mitigation is a **preStop hook** that sleeps a few seconds so the readiness change and endpoint removal propagate before the process exits: ```yaml lifecycle: preStop: exec: command: ["sh", "-c", "sleep 10"] terminationGracePeriodSeconds: 45 # must exceed preStop + shutdown timeout ``` ## Why liveness must not flip during shutdown If liveness went BROKEN on shutdown, Kubernetes might restart the container mid-drain, defeating the graceful stop. Spring keeps it CORRECT. ## Gotchas summary - Readiness gates warm-up on startup automatically — don't add manual sleeps for that. - `terminationGracePeriodSeconds` must be larger than preStop sleep + shutdown timeout, or the pod is SIGKILLed mid-drain. - The readiness OUT_OF_SERVICE only helps if your readiness probe is actually mapped and its `periodSeconds`/`failureThreshold` are tight enough to notice quickly.
- Why add a preStop sleep if Spring already flips readiness to REFUSING_TRAFFIC on shutdown?Because Kubernetes sends SIGTERM and removes the pod from Service endpoints asynchronously, and endpoint removal must propagate to every kube-proxy. A preStop sleep gives that propagation (and the readiness flip) time to take effect before the process actually exits, avoiding requests routed to a dying pod.
- Why doesn't Spring flip liveness to BROKEN during shutdown?A planned shutdown isn't an unrecoverable fault. Flipping liveness would let Kubernetes restart the container mid-drain, defeating graceful shutdown. Liveness stays CORRECT.
saying these in an interview costs you the question
- Saying readiness is ACCEPTING_TRAFFIC from the very first moment of startup (it waits for full readiness)
- Claiming graceful shutdown alone fully solves the traffic-routing race without preStop/endpoint propagation
- Setting terminationGracePeriodSeconds smaller than preStop + shutdown timeout
- Expecting liveness to change during a normal shutdown