skip to content

Every broker node now runs the new release, but the capability the upgrade was for does nothing — why?

level: middleimportance: should knowfreq 42%

answer

  1. installed is not enabled
  2. two gates, not one
  3. the agreement must be raised too
  4. inert until the roll truly finishes
  5. enabling it is one-way

basics

~20 s

A capability that changes what members write or how they talk to each other arrives inert. It stays off until every broker node carries the new binaries and the agreed internal version has been raised, because a half-upgraded cluster could not tolerate it.

solid answer

~50 s

New binaries on every member is only the first of the two gates. A capability whose output other members — or the stored data — must understand cannot switch itself on the moment a node starts, because during the roll its neighbours may still be old. So it is a **held-off capability**: present in the release, inert until the cluster as a whole says it is safe, which normally means the agreed internal version has been raised and, for anything that changes what is written, the stored format version too. Until then the feature is installed and idle. The mistake is reading "all nodes upgraded" as "upgrade finished" — the roll is not complete until the second pass, and enabling the capability is a separate, deliberate and usually one-way decision rather than a side effect of deployment.

go deeper

for a junior

Remember that a release can contain a feature that is installed but deliberately idle, and that the upgrade is not finished when the last binary is replaced.

for a middle

Explain the two gates — all members new, and the agreed internal version raised — and why a capability whose output peers must parse cannot be live before both.

for a senior

Show the diagnostic order: a skipped member, a forgotten second pass, an unthrown switch, or a per-object override beating the cluster-wide default.

for a principal

Own the policy that a roll is not closed until the second pass and the enablement are done or explicitly deferred, so clusters do not accumulate half-finished upgrades nobody remembers.

## Installed is not enabled Two different things can be true of a feature in a new broker release: - **the code is present** on every member, because the binaries have been replaced; - **the cluster is entitled to use it**, because every member could cope with what it produces. A rolling upgrade delivers the first well before the second. That gap is not an oversight; it is the mechanism that makes a mixed-version window survivable. A **held-off capability** is one that stays inert until both gates are satisfied. ## Which capabilities need holding off Not every improvement in a release needs a gate. The distinction is whether anything *outside the one process* has to understand the result. | kind of change | needs holding off? | why | |---|---|---| | an internal performance fix on one member | no | nothing else observes it | | a new field in the messages members exchange | yes | an old peer cannot parse it | | a change to the layout written to disk | yes | an old binary cannot read it back | | a new cluster-wide coordination behaviour | yes | every member must participate | | a new metric or log line | usually not | it is emitted locally and read by humans or agents | So the honest rule is: capabilities that change **what members say to each other** or **what is written down** are gated; purely local improvements arrive live with the binary. ## The two gates, and why both 1. **Every broker node carries the new binaries.** Otherwise a member that does not implement the capability at all would be expected to participate in it. 2. **The agreed internal version has been raised.** The binaries being new is not the same as the cluster having *declared* that no old member will return. Until that declaration, the cluster is deliberately behaving as the previous release, so a capability that contradicts that behaviour must stay off. Some platforms bundle these as a single operator action; others expose the capability as a separate thing to turn on once the agreement is raised. The failure mode is the same in both cases: a team upgrades, sees no change, and concludes the release is broken or the feature was overstated. ## Why enabling it is a decision, not a formality Enabling a held-off capability is usually a **one-way step**, and that is the part interviews probe. Once the cluster starts producing state that presumes the capability — a record layout, a coordination record, a metadata entry — reinstalling the previous binaries does not undo it. The previous release is not able to read or participate in what has been written since. A workable sequence, therefore: 1. finish the first pass and let the cluster soak on new binaries with the old agreement; 2. confirm no member is unhealthy and nothing is being carried by a member you would hesitate to restart; 3. raise the agreed internal version; 4. **separately**, and having decided you will not go back, enable the capability; 5. watch the thing the capability was supposed to improve, so that you can attribute a regression to it rather than to the roll. Compressing steps 3 and 4 into one change is common and usually survivable, but it costs you the ability to say which of the two caused a problem. ## The diagnostic, when the feature does nothing The question "why is it inert?" almost always resolves to one of four answers: - one member is still on the old binaries — often a member that was skipped because it was unhealthy at the time; - every member is new but the agreed internal version was never raised, because the second pass was forgotten after a successful first one; - the agreement was raised but the capability is a separate switch nobody threw; - the capability applies per stream or per queue rather than cluster-wide, and the cluster-wide default was changed while the objects that matter carry an override. The last one is a settings-scope question rather than a version one, but it presents identically and is worth eliminating. ## Where platforms differ How visible any of this is varies. Some platforms make the agreed internal version an explicit operator-set value and list, per release, which capabilities it gates. Others infer the floor from the oldest member present and light capabilities up on their own, which is convenient and removes the deliberate decision point. On a rented cluster the provider may enable capabilities on its own schedule, so the tenant's version of this question becomes "which behaviour changed under me, and when" rather than "what have I not switched on". The invariant across all of them is the reason for the gate: no capability whose output others must understand can be live while a member that does not understand it might still be in the cluster.

  • Which changes in a release do not need to be held off?
    Anything whose effect never leaves the one process: an internal performance fix, a memory-handling change, a new log line or metric. Nothing outside that member has to parse the result, so it takes effect as soon as the binary starts. The gate exists only for behaviour other members or the stored data must be able to cope with.
  • A team upgraded months ago and only now notices the capability is off. What is the risk of just enabling it?
    Enabling it is normally one-way, so the risk is taking an irreversible step with no recent evidence that the cluster is healthy on the new binaries. Treat it as its own change: confirm every member is actually new, confirm the agreement is raised, enable it in isolation, and watch the behaviour it affects so a regression can be attributed.

A new rule is printed in every copy of the rulebook the day the books are delivered, but it does not govern play until the league declares the season started under the new edition.

saying these in an interview costs you the question

  • Reads all nodes upgraded as the upgrade being finished
  • Assumes every feature in a release is live the moment the binary starts
  • Thinks enabling a held-off capability can simply be switched back off
  • Concludes the release is broken rather than checking the second pass
  • Bundles the raise and the enablement and then cannot attribute a regression