skip to content

During rolling restarts, a Spring Boot service on embedded Tomcat returns errors for requests that were already in flight. What does `server.shutdown=graceful` change about how the embedded container stops, which property bounds it, and why is that setting alone often not enough?

level: seniorimportance: should knowfreq 50%

answer

  1. the default cuts work in progress
  2. stop accepting, then drain
  3. the wait has a bound and a default
  4. in-flight is not the same as inbound
  5. the routing tier learns late

basics

~20 s

With server.shutdown=graceful the embedded container stops accepting new requests and lets in-flight ones finish, bounded by spring.lifecycle.timeout-per-shutdown-phase (default 30 seconds). It does not stop the load balancer routing new traffic, so a pre-shutdown delay is still needed.

solid answer

~50 s

The default is immediate: on shutdown the web server stops at once and any request still executing is cut, which the caller sees as a reset or a 502 from the proxy. `server.shutdown=graceful` changes that to a two-phase stop — the container refuses new requests first, then waits for the in-flight ones to complete before the context closes. The wait is bounded by `spring.lifecycle.timeout-per-shutdown-phase`, which defaults to 30 seconds; anything still running when that expires is terminated anyway. Two things still break it. First, the JVM must actually run its shutdown hook, so the process needs a SIGTERM and enough time before any SIGKILL. Second, and more commonly, the load balancer or service registry keeps sending *new* requests for however long its health check or deregistration takes to notice — and those arrive after the container has begun refusing. That gap is closed by a delay before shutdown starts, configured in the deployment platform, not by the Boot property.

code

properties · 2 lines
properties
server.shutdown=graceful
spring.lifecycle.timeout-per-shutdown-phase=25s

go deeper

for a junior

Know that Spring Boot stops the embedded server abruptly by default, and that one property switches it to letting current requests finish first.

for a middle

Explain the two phases — refuse new requests, then drain — and name the property that bounds the wait along with its 30-second default.

for a senior

Demonstrate the whole sequence: mark out of rotation, wait for the routing tier to notice, then signal; and be able to say why enabling the property alone leaves the error rate roughly where it was.

for a principal

Set the deploy contract across a fleet: consistent grace periods that fit inside the platform's termination window, bounded request durations that make draining possible, and a measured zero-error rollout rather than an assumed one.

## What the default does `server.shutdown` defaults to `immediate`. When the JVM receives a termination signal, Spring's shutdown hook closes the application context and the web server is stopped without regard for work in progress. A request halfway through a database transaction has its connection torn out; the client sees a connection reset, or, if a proxy sits in front, a 502. Under a rolling deploy of a busy service this is not a rare edge case — every instance replacement drops whatever it was holding. ## What graceful adds ```properties server.shutdown=graceful spring.lifecycle.timeout-per-shutdown-phase=30s ``` With `graceful`, shutdown becomes two phases. First the web server stops accepting: new connections and new requests are refused, so nothing else enters the application. Then the container waits for active requests to complete. Once they drain — or once the grace period expires — the context closes and the process exits. `spring.lifecycle.timeout-per-shutdown-phase` bounds that wait, defaulting to 30 seconds. Choosing it is a real trade. Too short and long requests are still killed, just later. Too long and a deploy crawls, or a wedged request holds the instance until the platform loses patience and sends a kill signal — at which point you are back to an abrupt stop with extra delay added. A sensible value sits just above the service's realistic worst-case request duration, which is itself bounded by the timeouts on its own downstream calls. If a request can legitimately run for two minutes, no shutdown setting will save it; shorten the request. ## Why the process must be asked nicely Graceful shutdown runs from the JVM shutdown hook, so it only happens if the process receives a signal it can handle. A forced kill bypasses everything. That has two practical implications: the process must be the signal recipient (a shell wrapper that does not forward signals will swallow it, leaving the JVM oblivious until it is killed), and the platform's termination grace period must be longer than the Boot shutdown timeout, or the kill lands mid-drain. ## The gap graceful shutdown cannot close This is the part candidates usually miss. Graceful shutdown protects requests that have *already arrived*. It does nothing about the requests still being sent. A load balancer or registry learns an instance is going away only when its health check fails or its deregistration propagates, and that takes seconds. During those seconds the instance is refusing — so the caller gets connection errors from an instance that is behaving exactly as configured. The fix is sequencing, and it lives outside the application: 1. Mark the instance out of rotation, or let its health signal go unhealthy. 2. **Wait** long enough for the routing tier to observe that and stop sending — typically a fixed delay in the deployment platform before the termination signal is sent. 3. Only then send the signal, letting graceful shutdown drain what remains. Step 2 is the one that is forgotten. Without it, graceful shutdown converts "errors on in-flight requests" into "errors on newly routed requests" and the error rate barely moves, which is exactly why the property alone often does not fix the reported symptom. ## What still gets cut Graceful shutdown drains requests, not connections. Long-lived streams — server-sent events, WebSocket sessions, a long-polling endpoint — will simply be running when the grace period expires and will be terminated. Those need their own application-level handling: a close frame or a signal that tells clients to reconnect elsewhere before the process goes. Background work started by a request but running on a separate executor is likewise not tracked as an in-flight request; its lifecycle is separate. ## Verifying it Do not take the property on trust. Put a slow endpoint under steady load, send the same termination signal the platform sends, and check three things: the in-flight requests completed with a normal status, the process exited within the grace period, and the error count during the window is zero rather than merely lower. If errors moved from in-flight requests to connection refusals, the missing piece is the delay in step 2, not the Boot configuration.

  • How would you choose a value for spring.lifecycle.timeout-per-shutdown-phase?
    Set it just above the service's realistic worst-case request duration, which should itself be bounded by the timeouts on its outbound calls. Too short and slow requests are still killed; too long and deploys drag while a wedged request pins the instance. It must also stay below the platform's termination grace period, or a forced kill lands mid-drain and the setting is meaningless.
  • Why can a wrapper script around the java process break graceful shutdown entirely?
    Graceful shutdown runs from the JVM shutdown hook, which fires only if the JVM itself receives the termination signal. A shell wrapper that does not exec the process or forward signals receives it instead, the JVM never learns it should stop, and the eventual forced kill terminates everything abruptly. The process must be the signal recipient.
  • Does graceful shutdown protect an open server-sent-events stream?
    No. It drains requests that will complete, and a long-lived stream will not complete within the grace period — it is terminated when the timeout expires. Streaming endpoints need application-level handling: signal clients to reconnect, or close the stream deliberately when shutdown begins, so they reconnect to a healthy instance rather than discovering the drop.

saying these in an interview costs you the question

  • Thinks graceful shutdown also stops the load balancer sending traffic
  • Believes the default already drains in-flight requests
  • Assumes a forced kill still runs the drain phase
  • Sets an unbounded or very long grace period to be safe
  • Expects long-lived streams to survive the shutdown window

context