What standing policy should an organisation set for how long a mixed-version window may stand and who may cross the abandon point?
answer
- cumulative damage across an estate
- finish or revert, no third option
- bounded skew is a tested property
- crossing needs a named owner
- reversible phase routine, irreversible not
basics
~20 sSet a policy limit on how long a cluster may sit part-rolled, require every roll to finish or revert inside it, and make crossing the abandon point a named decision with an owner and a written recovery plan rather than a step in a runbook.
solid answer
~50 sAn estate that upgrades continuously is permanently in some mixed-version window, so the interesting rules are organisational rather than per-cluster. Three are worth fixing. First, a **policy limit on how long** a cluster may stand part-rolled — long enough to soak on new binaries, short enough that the version skew stays inside what the release was tested at, with finish-or-revert as the only two exits. Second, **who may cross the abandon point**: since a reinstall stops working there, crossing needs a named owner and a written far-side plan, not a tick in a runbook. Third, **how far behind a cluster may fall at all**, because a cluster several releases back cannot be rolled in one hop and its upgrade becomes a project. The point of writing these down is that each roll then carries a deadline and a decision owner by default, instead of each team improvising both under pressure.
go deeper
Know that a part-rolled cluster is a temporary state with a deadline, not a resting place, and that someone must own finishing it.
Explain why the window has to be bounded: interoperability is only tested across a limited gap, and every incident inside it must be diagnosed against two releases.
Argue the split — a routine reversible phase and a separately approved crossing — and name the evidence and recovery plan you would demand before the crossing.
Set the estate rules: a stated window limit with finish-or-revert as the only exits, a named owner for each crossing, a floor on how far behind a cluster may fall, and an inventory that makes violations visible.
## Why this is a policy question, not a runbook question One cluster's upgrade is a procedure. An estate of clusters, upgraded by several teams on their own cadences, produces a different problem: at any moment some clusters are part-rolled, some are soaking between passes, and some were left half-finished when the engineer who started went on leave. Nobody is doing anything wrong at the level of a single step. The damage is cumulative, and only a standing rule addresses it. ## Rule one: a policy limit on how long a mixed-version window may stand The window is useful — soaking on new binaries under real traffic is the best evidence available, and it is still reversible. It is also a liability that grows: - **version skew is bounded by testing.** Releases are validated to interoperate across a limited gap. A window left open across another release's arrival drifts outside what anyone tested. - **two behaviours to debug.** Any incident during the window has to be diagnosed against two releases, and "which node served this request" becomes a question every investigation must ask. - **the way back decays.** Reverting one member is easy; reverting eleven, days later, with configuration and monitoring already adjusted for the new release, is a project. - **attention decays fastest of all.** A half-rolled cluster that is serving fine stops being anybody's task. So: a stated maximum — expressed in days, appropriate to the size of the cluster — and only two ways out of it. Finish the roll, or revert it. "Leave it and see" is not one of them. ## Rule two: crossing the abandon point is a named decision Up to the abandon point, going back is reinstalling the previous release. After it, going back is a restore or a cutover to another cluster. Those are different classes of act and deserve different governance. A workable standard: 1. the change record **names the abandon point** for this release — which step it is, since platforms tie those steps together differently; 2. the far-side recovery plan is **attached and current**: what would be restored, from when, what the restore does not bring back, and what the recovery targets are; 3. a **named owner** accepts the crossing, and it is scheduled as its own change rather than bundled with the last upgrade wave; 4. the crossing is **announced** to the teams whose services publish into or read from the cluster, because after it their incident options change too. This is not ceremony for its own sake. It converts an invisible irreversible step inside a long procedure into a visible one that someone has consciously taken. ## Rule three: a floor on how far behind a cluster may fall The cheapest upgrade is the one from the release just before. The rules that make this achievable are unglamorous: | rule | what it prevents | |---|---| | a maximum age for the release a cluster runs | a cluster that needs intermediate hops to be upgraded at all | | one supported release list for the whole estate | every cluster being a special case | | an owner recorded per cluster | a part-rolled cluster nobody is responsible for | | an inventory of releases actually running | discovering skew during an incident | ## What the policy should not try to do Two temptations are worth resisting. - **Mandating a fixed wave size or a fixed wait between members.** That is the restart discipline of one cluster and depends on its capacity and copy-set arithmetic; a central number is either too slow for small clusters or unsafe for loaded ones. - **Approving the whole roll as one change.** If the first pass and the crossing need the same approval, one of them is wrong: the reversible phase should be routine and the irreversible one should not. ## Where the estate shape changes the answer How much of this you actually control varies. Where clusters are self-operated, the whole policy is yours to set and enforce. Where clusters are rented, the provider decides when its releases roll and may cross its own equivalents of the abandon point on its own schedule — so the tenant's policy becomes a different one: knowing when the provider's windows fall, ensuring clients survive reconnects unattended, and deciding what evidence is required before a behaviour change lands without notice. Most large estates are a mixture, and the rule that spans both is the same: a change that removes the cheap way back should never happen as a side effect of something else.
- How would you choose the actual number of days for the limit?From two constraints rather than taste: the interoperability gap the release states, which caps it from above, and the time needed to observe a full business cycle on the new binaries, which sets the floor. If the floor exceeds the cap, the cluster is too large to roll in one window and needs upgrade waves with intermediate finishing points.
- What is the cheapest signal that the policy is being violated?An inventory of which release each broker node in the estate is actually running, refreshed automatically and compared against the supported list. Mixed releases inside one cluster, with a timestamp for when that started, is the single query that finds forgotten rolls, and it costs far less than any review process.
- Does this policy apply to a rented cluster?Not in the same form. The provider chooses when its releases roll and takes its own irreversible steps, so the tenant cannot set the window. What the tenant can set is the requirement that clients survive unannounced reconnects, that someone knows when the provider's windows fall, and that behaviour changes are detected rather than discovered during an incident.
saying these in an interview costs you the question
- Leaves a part-rolled cluster open indefinitely because it is serving fine
- Treats crossing the abandon point as a routine step in the runbook
- Mandates one wave size and wait time for every cluster centrally
- Assumes any release can interoperate with any other indefinitely
- Has no inventory of which release each broker node is running
- Writes the recovery plan only after deciding to go back