Your edge proxy tier carries long-lived connections — WebSockets and gRPC streams that can last hours — and the tier has to be deployed weekly. How would you decide a drain window and a maximum connection lifetime, and what are you trading off?
answer
- a stream outlives the release cycle
- waiting for completion never terminates
- cap the lifetime, add jitter
- spread the reconnects, don't concentrate them
- old instances keep old config while draining
basics
~20 sConnections that never end cannot be waited out, so pick a bounded drain window and cap connection lifetime deliberately. Capping spreads reconnects continuously instead of concentrating them at deploy time; the cost is constant reconnect churn and a hard client-side reconnect requirement.
solid answer
~50 sWith hours-long connections and weekly deploys, waiting for natural completion is not a strategy — the session outlives the release cycle, so every deploy either cuts connections or never finishes. I would decide two numbers from the session-length distribution. The **drain window** covers the short tail: long enough that ordinary request/response and brief streams complete, short enough that a deploy is not held hostage; beyond it, force-close. The **maximum connection lifetime** is the more important knob: capping it — say to an hour, with jitter so reconnects do not synchronise — converts a rare, concentrated, deploy-shaped disruption into a continuous low rate of reconnects that the client path is exercised against every day. The trade is real: you are paying handshakes and client churn permanently to avoid a large disruption occasionally, and it only works if every client can reconnect and resume. That last point is a platform contract, not a per-team choice.
go deeper
Know that WebSocket and streaming connections stay open for a long time, so deploying the tier in front of them eventually has to close connections rather than wait for them.
Be able to explain why a bounded drain window is required and what a maximum connection lifetime does to the distribution of reconnects.
Show that you derive both numbers from measured session lengths, size for the overlap window, and prefer a polite go-away or close-frame signal over a reset.
Own the trade explicitly: continuous rehearsed churn against rare large disruption, the capacity cost of overlap, the fact that draining instances keep stale config, and reconnect-and-resume as a contract in the shared client library.
## Why this is a design decision rather than a setting For ordinary request/response traffic, replacing a proxy instance is a solved problem: stop accepting new connections, let the in-flight ones finish in a second or two, exit. The whole approach depends on connections ending on their own shortly after you stop feeding them. Streaming breaks that assumption completely. A WebSocket held open for four hours will not end because you started a deploy. So the question stops being "how long do we wait" and becomes "how do we want to distribute the disruption", which is a judgement about the product and the client population, not a proxy tuning exercise. ## The two numbers **Drain window.** Derive it from the session-length distribution, not from the maximum. Something around the p90–p95 of session duration lets most short sessions finish naturally while keeping the deploy bounded. Everything past it is force-closed. Setting it to cover the tail is a trap: with hours-long sessions the tail is longer than the deploy cadence, so the window would never close and the old and new tiers would coexist indefinitely. **Maximum connection lifetime.** This is the lever that actually changes the shape of the problem. If connections are capped at, say, an hour plus jitter, then at any moment the population is uniformly distributed across its lifetime, reconnects happen continuously at a predictable rate, and a deploy disrupts only the fraction still attached at the drain deadline. Without a cap, the reconnect storm arrives all at once at deploy time — thundering herd against the new instances, which are cold, plus whatever re-authentication and state rebuild each reconnect triggers. Jitter matters as much as the cap. A fixed lifetime applied to connections that were themselves established in a burst reproduces the burst one lifetime later. ## What you are trading - **Continuous churn vs concentrated disruption.** Capping lifetimes means paying handshakes, TLS setup, authentication and any per-session warm-up forever, in exchange for never facing a mass reconnect you did not schedule. For most platforms that is the right trade, because the continuous path is *exercised* and therefore known to work. - **Capacity for the overlap.** During the drain, the old and new instances both serve. The tier has to be sized for the sum, and the longer the drain window the larger the overlap. - **Stale behaviour on the old instances.** Connections established before the deploy keep the old code and old routing decisions for the whole drain window. A config change — including a rollback — does not reach them. Anything security-relevant, such as a revoked route or a tightened policy, needs an explicit mechanism to close connections rather than a config push. - **Client capability.** All of it assumes clients reconnect cleanly and resume: exponential backoff with jitter, resumption from a checkpoint or offset rather than restarting a stream from zero, and idempotence on whatever they replay. Where the client is a browser you control, that is engineering. Where it is a third-party SDK or an embedded device, it may simply be false, and then forced closes are user-visible failures and your lifetime cap has to be much longer or the tier much more stable. ## Signalling rather than severing A forced close is the last resort; ask first. HTTP/2 and HTTP/3 have a connection-level go-away signal that tells the peer to stop starting new streams and reconnect, allowing in-flight streams to finish — which is exactly the drain semantic, expressed to the client instead of imposed on it. For WebSocket, a close frame with an application-defined code lets a well-written client reconnect immediately and quietly, whereas a TCP reset presents to the client as a network error and often triggers a much less graceful path. Building the polite signal into the platform is what makes lifetime capping tolerable. ## How I would decide, concretely 1. Measure the session-length distribution, per client type. If p50 already exceeds the deploy interval, capping is mandatory rather than optional. 2. Set the lifetime cap so that the steady-state reconnect rate is comfortably below what the tier and the auth path can absorb, and add jitter of at least ten percent. 3. Set the drain window from p90–p95 of *session* length, capped by what the deploy pipeline will actually tolerate. 4. Size the tier for the overlap, and verify the capacity assumption during a real deploy rather than on paper. 5. Make reconnect-and-resume a platform contract implemented in the shared client library, so no team can opt out of it by accident. ## The failure mode to avoid The worst configuration is an unbounded lifetime plus an ambitious drain window: deploys become slow, scary and rare, so the tier's release path decays from disuse, and the eventual forced reconnect is both larger and less rehearsed than it would ever have been under a cap. Frequent small disruptions that the system is designed for beat rare large ones it is not.
- Why add jitter to a maximum connection lifetime?Because connections established together will expire together. A fixed lifetime applied to a population that connected in a burst — after an incident, or after the last deploy — recreates that burst exactly one lifetime later, and then again after that. Ten to twenty percent of randomness spreads the expiries into a steady rate the auth path and the tier can absorb.
- What does a draining proxy instance do about configuration changes during its drain window?Generally nothing useful for connections already established: they keep the routing and policy they started with until they close. That is fine for a deploy and dangerous for a security change, so revocations and policy tightenings need a mechanism that closes affected connections rather than a config push you assume propagated.
- When would you not cap connection lifetime?When the clients cannot reconnect gracefully — embedded devices, third-party SDKs, anything without resumable state — because then every expiry is a user-visible failure rather than an invisible reconnect. In that case you stabilise the tier instead: deploy it far less often, decouple its release from the services behind it, and accept a longer drain.
saying these in an interview costs you the question
- Plans to wait for long-lived connections to end naturally
- Caps lifetime with no jitter, recreating the burst later
- Forgets the tier serves double capacity during the drain
- Assumes a config rollback reaches established connections
- Treats client reconnect-and-resume as someone else's problem