Compare how a stateless HTTP service and one using in-memory server sessions behave during a rolling deployment, a sudden instance crash, and an autoscale scale-in event. What must you still get right in the stateless case?
answer
- Failure unit: request, not session
- Rolling deploy = staggered mass logout when stateful
- Drain: fail readiness, delay, finish in-flight, exit
- Connection: close / GOAWAY on shutdown
- Cold start: readiness gates warm-up
basics
~20 sStateless: killed instances only lose in-flight requests, so rolling deploys, crashes and scale-in are routine. In-memory sessions: every replaced or crashed instance logs its users out. You still need connection draining, retry-safe requests and readiness checks.
solid answer
~50 sWith in-memory sessions, a rolling deploy is a staggered mass logout: each replaced pod destroys the sessions it owned. A crash does the same abruptly, and scale-in deliberately deletes live user state. Teams then hesitate to deploy, over-provision, and disable scale-in. Stateless changes the failure unit from a user's session to a single in-flight request. Replacing an instance costs at most the requests currently on it. That makes blue-green and rolling deploys, spot instances and aggressive scale-in safe. What you still need: **connection draining** - stop new traffic at the load balancer, let in-flight requests finish, then exit within the termination grace period; **readiness and liveness probes** so traffic only reaches warm instances; **retry-safe semantics** so a client retry of an interrupted write does not duplicate it; and awareness that HTTP keep-alive and HTTP/2 connections outlive a request, so shutdown must send Connection: close or GOAWAY.
go deeper
Say that with stateless servers a restarted instance does not log anyone out, whereas in-memory sessions are lost when the instance goes away.
Add the mechanics: health checks remove the instance, clients retry failed in-flight requests, and new instances join the pool once ready.
Own the operational detail - draining order, termination grace period, keep-alive shutdown, retry idempotency, warm-up gating - and note the shared store's own failure modes.
Reason about deploy strategy and blast radius: how failure-unit size changes release cadence, whether to run spot or preemptible capacity, and how much headroom the shared state tier needs.
## Three events, two designs **Rolling deployment.** Instances are replaced a few at a time. In a stateless service this is invisible: traffic shifts to remaining instances, new ones join the pool once ready, and no user notices beyond a brief capacity dip. With in-memory sessions, each replaced instance takes its sessions with it, so a full rollout logs out effectively the entire logged-in population, in waves. The usual reaction is deploying only at night, which slows delivery and makes each release larger and riskier. **Instance crash.** Stateless: the load balancer's health check marks the node down, in-flight requests on it fail, clients retry, and the incident is a small error spike. Stateful: everyone routed to that node loses their session mid-task - abandoned checkouts, lost form input, forced re-login. **Scale-in.** This is where in-memory session state hurts most, because the platform is deliberately deleting healthy instances. Teams respond by disabling scale-in or setting a very high minimum, which removes most of autoscaling's value. A stateless service can scale in aggressively; the only cost is draining. ## What stateless does not give you for free Statelessness makes instances **disposable**, not requests **immortal**. Four things still need engineering. **Connection draining / graceful shutdown.** On termination the instance must be removed from the load-balancer pool first, then keep serving until in-flight requests finish, then exit. In practice: catch the termination signal, fail the readiness probe, wait long enough for the balancer to notice (propagation is not instant), stop accepting new work, wait for the in-flight set to empty, exit before the grace period ends and the platform force-kills. Skipping the initial delay is the most common bug: the process exits correctly but the balancer is still sending it traffic. **Keep-alive connections.** HTTP/1.1 persistent connections and HTTP/2 multiplexed connections are long-lived. A client or proxy holding an open connection to a dying instance will keep sending on it. The server should send `Connection: close` on HTTP/1.1 responses during shutdown, or an HTTP/2 `GOAWAY` frame, so peers re-resolve to a live instance. **Retry safety.** If a request is cut off mid-flight, the client cannot tell whether the server applied it. Safe and idempotent methods can be retried freely; a non-idempotent write needs an idempotency mechanism so a retry is deduplicated. Without that, disposable instances turns into duplicate charges under deploy churn. **Warm-up.** New instances start cold: empty caches, unwarmed runtime optimisation, empty connection pools. Readiness probes should not pass until the instance can actually serve at latency, otherwise scaling out momentarily makes latency worse and can trigger a feedback loop of further scaling. ## The shared-store caveat Moving session state to a shared store makes instances disposable but makes that store a shared dependency. During a scale-out burst, every new instance opens connection-pool capacity against it; during a store failover, all instances degrade together. Statelessness redistributes the fragility rather than deleting it, so the store needs its own capacity headroom, timeouts and defined failure behaviour. ## Interview framing State the contrast crisply (failure unit = one request, not one session), then immediately show operational depth by naming draining order, keep-alive shutdown signals, readiness gating and retry safety. That combination - the principle plus the four things it does not solve - is what separates a senior answer from a textbook one.
- Your pods shut down gracefully but users still see errors during deploys. What is the likely cause?The process almost certainly stops accepting connections before the load balancer has observed it as unready, so traffic is still being routed to a socket that is closing. Fix it by failing the readiness probe first, waiting past the balancer's detection interval, and only then closing listeners; also close keep-alive connections with Connection: close or an HTTP/2 GOAWAY.
- Does statelessness remove the need for sticky routing on a WebSocket or streaming endpoint?No. A long-lived connection is inherently bound to one instance for its lifetime, so that instance is special until the connection ends. You keep instances disposable by making reconnection cheap and resumable - the client reconnects to any instance and replays a position or subscription, rather than the server holding unrecoverable per-connection state.
saying these in an interview costs you the question
- Assuming stateless means zero user impact during deploys, ignoring in-flight requests
- Closing listeners immediately on the termination signal without failing readiness first
- Forgetting that keep-alive and HTTP/2 connections outlive a single request
- Claiming a shared session store makes deploys risk-free while ignoring its own failover behaviour
- Retrying non-idempotent writes after an interrupted request with no deduplication