What actually triggers Spring Boot's graceful shutdown, and how does the SmartLifecycle phase system enforce the drain timeout?
answer
- registerShutdownHook → context.close()
- SIGTERM fires hook; SIGKILL doesn't
- SmartLifecycle: highest phase stops first
- WebServerGracefulShutdownLifecycle drains
- timeout is per phase, not global
basics
~20 sGraceful shutdown runs when the ApplicationContext closes — usually because a SIGTERM fires the JVM shutdown hook. The embedded server is a SmartLifecycle bean in the shutdown phase, and timeout-per-shutdown-phase caps how long that phase waits for in-flight requests.
solid answer
~40 sSpring Boot registers a JVM shutdown hook that calls `context.close()`. On close, the `Lifecycle`/`SmartLifecycle` beans are stopped in **descending phase order**. The embedded web server is wrapped by `WebServerGracefulShutdownLifecycle`, which on stop tells the web server to begin graceful shutdown: reject new requests, drain in-flight ones, then invoke the stop callback. `spring.lifecycle.timeout-per-shutdown-phase` bounds how long the `DefaultLifecycleProcessor` waits for all beans in a given phase to report stopped before moving on — so it's effectively the drain deadline. Triggers include SIGTERM (kill, `docker stop`, k8s termination), Ctrl-C, the Actuator `/shutdown` endpoint, and explicit `context.close()`. SIGKILL and OOM kills skip it entirely because the hook never runs. Because it hangs off context close, the same mechanism also runs `@PreDestroy`/`DisposableBean` cleanup, ordered by lifecycle phase.
code
java · 21 lines// A custom SmartLifecycle that must drain before the web server stops.
// Higher phase => stopped EARLIER during shutdown.
@Component
class MessageConsumerLifecycle implements SmartLifecycle {
private volatile boolean running;
@Override public void start() { running = true; /* subscribe */ }
@Override public void stop(Runnable callback) {
// stop pulling new work, let in-flight finish, then:
running = false;
callback.run(); // signals this phase complete
}
@Override public void stop() { stop(() -> {}); }
@Override public boolean isRunning() { return running; }
// Stop before the default phase (web server sits near DEFAULT_PHASE for graceful shutdown)
@Override public int getPhase() { return SmartLifecycle.DEFAULT_PHASE + 1; }
}
// Drain of each phase is bounded by spring.lifecycle.timeout-per-shutdown-phase (default 30s).go deeper
Know shutdown happens on SIGTERM/context close; details are advanced.
Understand the JVM shutdown hook and that the timeout bounds draining.
Explain SmartLifecycle phase ordering and how the web-server lifecycle drains within a phase.
Reason about per-phase timeout accumulation, phase placement of custom lifecycle beans, readiness coordination, and orchestrator grace-period sizing for true zero-downtime.
### The trigger chain 1. **Shutdown hook**: `SpringApplication` calls `context.registerShutdownHook()` by default, registering a JVM shutdown hook (`Runtime.addShutdownHook`). When the JVM receives a terminating signal it runs registered hooks. 2. **Signals that fire hooks**: `SIGTERM` (the polite `kill`, `docker stop`, Kubernetes pod deletion, systemd stop), `SIGINT` (Ctrl-C), and normal JVM exit. **`SIGKILL` (kill -9) and hard OOM kills do NOT** — the JVM dies without running hooks, so there is no graceful shutdown. 3. Other triggers: the Actuator **`POST /actuator/shutdown`** endpoint (disabled by default; enable via `management.endpoint.shutdown.enabled=true`), and any explicit `applicationContext.close()`. ### What `context.close()` does `AbstractApplicationContext.close()` → `doClose()` → stops the `LifecycleProcessor`, publishes `ContextClosedEvent`, then destroys singletons (`@PreDestroy`, `DisposableBean`, `destroyMethod`). The **`DefaultLifecycleProcessor`** is what stops `Lifecycle`/`SmartLifecycle` beans. ### SmartLifecycle phases `SmartLifecycle` beans declare a `getPhase()`. On **startup**, lower phases start first; on **shutdown**, they stop in **reverse** (highest phase stops first). Beans in the same phase are stopped together, and the processor **waits up to `spring.lifecycle.timeout-per-shutdown-phase` (default 30s)** for all of them to signal completion via the `stop(Runnable callback)` variant before it proceeds to the next (lower) phase. ### Where the web server fits Spring Boot registers `WebServerStartStopLifecycle` and, for graceful shutdown, `WebServerGracefulShutdownLifecycle`. When `server.shutdown=graceful`, stopping this lifecycle bean calls the web server's `shutDownGracefully(callback)`: - New requests are refused (network-layer for Tomcat/Jetty/Reactor Netty; HTTP 503 for Undertow). - In-flight requests are allowed to finish. - When they drain, the callback fires and the phase completes — **or** the `timeout-per-shutdown-phase` elapses and the processor moves on, letting the server force-close remaining connections. Because the graceful-shutdown lifecycle sits at a phase that is stopped **before** the application's own beans are destroyed, requests can still use their dependencies while draining. ### Interactions and edge cases - **Ordering vs `@PreDestroy`**: request draining happens during the lifecycle-stop phase, which precedes singleton destruction — so beans a request needs are still alive during the drain. Custom `SmartLifecycle` beans should pick phases carefully relative to `SmartLifecycle.DEFAULT_PHASE`. - **Timeout is per phase, not global**: if you have multiple shutdown phases each can consume up to the timeout; total shutdown can exceed a single timeout value. Size your orchestrator grace period accordingly. - **Readiness coordination**: on shutdown Spring Boot flips availability to `ReadinessState.REFUSING_TRAFFIC` (`AvailabilityChangeEvent`); wire the Kubernetes readiness probe to `/actuator/health/readiness` so the load balancer stops routing before/while draining. - **`terminationGracePeriodSeconds`**: must exceed the sum of any preStop delay plus the drain timeout, or the platform SIGKILLs mid-drain, defeating the whole mechanism. - **Reactive vs servlet**: the mechanism works for both stacks (Reactor Netty on WebFlux, Tomcat/Jetty/Undertow on MVC). ### When to reason about this deeply Zero-downtime rolling deploys, long-running requests/streaming, message-listener draining, and any place where dropping an in-flight unit of work is costly. For a plain internal service the defaults (`immediate`) may be acceptable; for user-facing prod, enable graceful and coordinate with the platform.
- If you have two SmartLifecycle phases that each take 25s to drain, is your total shutdown bounded by 30s?No. timeout-per-shutdown-phase is applied per phase, not globally. Two phases can each wait up to the timeout, so total shutdown can approach 60s. Size terminationGracePeriodSeconds to cover the sum, not a single phase.
- Why does kill -9 make graceful shutdown irrelevant?SIGKILL terminates the JVM immediately without running shutdown hooks, so context.close() and the lifecycle drain never execute. Only signals that let hooks run (SIGTERM/SIGINT) or an explicit context close trigger graceful shutdown.
saying these in an interview costs you the question
- Thinking timeout-per-shutdown-phase is a single global shutdown budget
- Believing Actuator /shutdown is enabled by default
- Assuming @PreDestroy runs before in-flight requests finish draining
- Claiming SIGKILL triggers the JVM shutdown hook
- Forgetting readiness deregistration and treating graceful shutdown as sufficient alone