As lead of an app that deploys several times a day, how would you decide what one release may change while clients from the previous release are still running?
answer
- measure session lifetime first
- the window is the policy
- additive inside it, free outside it
- expand now, contract one release later
- escape hatch before full isolation
basics
~20 sMeasure how long sessions really live and make that the compatibility window: inside it, changes to what a loaded client touches must be additive or retained. Add a cheap escape hatch so an out-of-date client reloads without losing input.
solid answer
~40 sStart from evidence: the distribution of tab and session lifetimes tells you how long a client from the previous release keeps talking to you, and error reports tagged with the client's build identifier tell you what it costs today. That window becomes policy. Inside it, changes to the surface an old client touches — asset filenames, write-handler addresses, route payload shapes — must be additive or the old form must be retained, which usually means expand in one release and contract in the next. Outside it, you change freely. Then buy an escape hatch: a client that detects it is out of date preserves pending input and hard-loads. Reserve heavyweight isolation, keeping whole previous deployments addressable, for flows where a mid-session failure destroys work. Review the policy against measured incidents, not against fear.
go deeper
Take away the principle: a release should not remove something a page loaded minutes earlier still depends on, and adding is safer than moving or renaming.
Be able to name the surfaces a policy has to cover — retained assets, write-handler addresses, payload shapes, persisted browser state — and the expand-then-contract sequence that keeps each safe.
Show how you would derive the window from measured session lifetimes and skew error rates, and implement detection, input preservation and a guarded one-time reload.
Own the economics: what the window costs in delivery speed and migration compatibility, which flows justify retained deployments, and how the policy lives in the release process rather than in one person's head.
## Frame it as a window, not a rule The question is not "can a release break old clients" — it always can — but **for how long you owe those clients anything**. That duration is measurable and should be: - the distribution of tab and session lifetimes, at the median and at the high percentiles; - how many of those sessions cross a deploy, which is session length multiplied by deploy frequency; - what actually fails today, from error reports tagged with the client build identifier alongside the serving release. A team that deploys hourly with five-minute sessions has almost no exposure. A team that deploys hourly into an app people leave open all day has a great deal, and the same policy would be either wasteful or negligent depending on which one you are. Measure first; the number is usually a surprise in one direction or the other. ## The compatibility contract inside the window Once there is a window, it applies to exactly one surface: **what an already-loaded client touches**. That is a short, enumerable list, which is what makes the policy cheap to enforce. | Surface | Rule inside the window | How it is enforced | |---|---|---| | Built asset files | Retain the previous builds' files; never delete on deploy | Deploy pipeline, not code review | | Write-handler addresses | Keep the previous mapping answerable, or pin the request to its build | Platform capability; verified per release | | Route payload shapes | Additive only; remove a field one release later | Code review, ideally a schema check in CI | | Persisted browser state | New client must read what the old one wrote | Code review | | Shared database and caches | Readable by the oldest client and version you keep alive | Migration review | Outside the window everything is free, and saying so matters: engineers who are told "never break compatibility" quietly hear "carry every field forever", and the codebase silts up. The instruction is expand-then-contract with a defined delay, not permanence. ## Buy the escape hatch before you buy isolation The cheap layer removes most of the pain and should come first: 1. **Detect.** The client compares the build it came from with the one the server reports, or treats an unroutable write as a version signal. 2. **Preserve.** Any pending input is stashed before anything drastic happens. 3. **Recover.** A full document load, one time, guarded so a genuinely broken release cannot produce a reload loop — a marker that says an attempt has already been made, after which the user is told rather than reloaded again. 4. **Nudge.** For long-lived sessions, an unobtrusive "a new version is available" affordance converts a forced interruption into a choice the user makes between tasks. Only then consider the heavyweight option of keeping whole previous deployments addressable and routing old clients back to them. It removes mid-session failure, and it costs running several versions at once plus a database-compatibility window as wide as your retention. Spend it where a mid-session failure destroys work — a checkout, an editor, a multi-step form — not across the whole app by default. ## Deciding by blast radius Rank the flows, not the releases: - **A skewed read** shows a broken view and is fully repaired by a reload. Cost: an annoyance, occasionally a support ticket. - **A skewed write** fails after effort has been invested and can lose typed data. Cost: real, and the one users tell you about. - **A silently wrong render** — old code falling into a default branch for data it does not understand — produces no error at all. Cost: the highest, because you learn about it late and from the wrong source. That ranking says where the budget goes: input preservation and explicit version detection on write paths first, payload discipline second, full pinning last and selectively. ## Make it a system, not heroics - Put the window and its rules in the release checklist, so "is this change safe for a client from the last release?" is asked by the process rather than remembered by the most experienced person present. - Emit the client build identifier with every error and every write, and chart skew errors against deploy times; a policy without that signal cannot be tuned and will drift toward superstition. - Re-derive the window when behaviour changes — a new long-session surface, a change in deploy cadence — rather than treating the first number as permanent. - Accept residual failures explicitly. Some clients will be arbitrarily old; the goal is that they fail visibly, recoverably and without losing work, not that they never fail.
- Why is a blanket "never break compatibility" rule a poor policy here?It has no end date, so retained fields and old handler mappings accumulate forever and nobody can tell which are still needed. A bounded window gives the same protection and a clear moment to delete: expand in one release, contract after the window has passed, with the removal scheduled rather than hoped for.
- What single piece of instrumentation makes this policy tunable?The client build identifier, reported with every error and ideally every write, next to the release that served the request. It turns skew from an untestable theory into a measurable rate that spikes at each deploy and decays, showing which surfaces actually break and how long exposure really lasts.
- How do you decide which flows deserve full deployment pinning rather than the cheap recovery path?By what a mid-session failure destroys. A browsing surface loses a view and a reload restores it. A checkout, an editor or a long form loses work the user cannot reconstruct, and there the cost of running retained deployments for that path is smaller than the cost of the failures it prevents.
saying these in an interview costs you the question
- Sets the compatibility window by instinct without measuring session lifetimes
- Requires permanent backward compatibility with no removal step
- Buys full deployment pinning before adding input preservation and reload recovery
- Reloads out-of-date clients automatically with no loop guard
- Assumes shipping client and server together makes the window unnecessary
- Treats a silently wrong render as less serious than a visible crash