skip to content

questions

5

Why is a new broker release rolled out one broker node at a time, and what runs side by side while it is?

level: juniorimportance: must knowfreq 70%

answer

  1. availability during the change
  2. one member out, the rest serve
  3. two releases live at once
  4. the mixed-version window
  5. held back until the last node

basics

~20 s

Taking one broker node at a time keeps every other member serving, so the cluster stays available through the change. The price is a mixed-version window: until the last node is done, two broker releases are live in the same cluster.

solid answer

~40 s

Stopping the whole cluster to replace binaries is an outage as long as the slowest machine's start-up. Rolling instead means one broker node is taken out, moved to the new broker release, brought back and confirmed healthy before the next one is touched, so at every instant the rest of the cluster is still accepting writes and serving readers. What that buys in availability it pays for in a **mixed-version window**: from the first node returning on the new release until the last one does, two releases are live together and must interoperate — replicating, electing and sharing membership through the metadata service. Anything the new release does differently has to stay held back until that window closes, which is why a version change is normally a two-pass shape rather than one sweep.

go deeper

for a junior

Recall the two facts: members are taken to the new broker release one at a time so the rest keep serving, and while that runs two releases are live in the same cluster.

for a middle

Explain what the mixed-version window demands — nodes of different releases replicating and electing together, and new behaviour held back until the last member is upgraded.

for a senior

Show the pre-flight judgment: capacity and durability headroom with a member absent, a stated interoperability gap, and a decided answer for a node that comes back unhealthy.

for a principal

Frame it as a standing cost. An estate that rolls constantly is permanently in some mixed-version window, so bounded skew and a finish-or-revert deadline become policy, not per-cluster improvisation.

## What "rolling" actually means A cluster runs the same broker software on several machines. Moving that software to a **new broker release** can be done two ways. - **All at once** — stop every broker node, replace every binary, start everything again. This is what a single-machine deployment does, and on a cluster it is a full outage whose length is the slowest machine's start-up time plus whatever recovery work a node does after being stopped. Writers either block or fail, and their own callers see it. - **Node by node** — take one member out of service, replace its binaries, bring it back, confirm it is healthy and carrying its share of work, then move to the next. At every instant, every member but the one being worked on is still serving. Production clusters use the second shape, and availability is the whole reason. The restart mechanics themselves — how long to wait between members, how you satisfy yourself that a returning node is back in sync, how many may be out at once — are the rolling-restart subject and are assumed here. This subject is the **version change** that the restart carries. ## The price: a mixed-version window From the moment the first node comes back on the new release until the moment the last one does, **two broker releases are live in the same cluster**. That interval is the *mixed-version window*, and nearly everything awkward about upgrading a cluster lives inside it: - nodes of different releases must still talk to one another — replicate records, elect leaders, agree membership and configuration through the metadata service — over a protocol both understand; - a client connects to whichever node it reaches, so one connection may be served by a new node and the next by an old one, and the application must not be able to tell; - anything the new release does *differently on the wire or on disk* must be held back, or an old node meets something it cannot handle. That third point is the reason a version change is normally a **two-pass roll**: the binaries move first while the cluster still speaks the old agreement, and the agreed internal version is raised only once every member is new. | | all at once | node by node | |---|---|---| | availability | none for the duration | full, minus one member at a time | | elapsed time | minutes | hours, or days on a large cluster | | releases live | one | two, for the whole roll | | interop burden | none | both releases must work together | | abandoning it | reinstall everything, restart everything | wave by wave, up to the abandon point | ## What must be true before the first node is touched 1. **Capacity headroom** — the cluster must carry its peak load with one member, or one upgrade wave, absent. A roll started at 95% utilisation finishes as an incident. 2. **Durability headroom** — losing a member must not drop a record's copy set below the floor a write is required to reach, or writes stall while the node is out. 3. **A stated interoperability claim** — the release must say it can run alongside the one you are on. Releases generally support a bounded skew, not an unbounded one, and a cluster left far behind may need an intermediate release on the way. 4. **A decision about the second node** — what you will do if the first one comes back unhealthy. "Continue and hope" is not a plan; "stop, reinstall the previous binaries on that one member, investigate" is. ## Where platforms differ The shape is common to the class, but what an operator *feels* during the window is not. - On platforms that **split a stream into parts** with a leader and follower copies per part, a departing member hands its leaderships to copies elsewhere and reclaims them afterwards, so the roll shows up as leadership churn and brief unavailability per part. - On a **shared-queue** broker where consumers compete for messages, there is no per-part leadership to move; the roll shows up as reconnects and the redelivery of whatever was in flight to the departing node. - On a **rented cluster**, the provider performs the roll on its own schedule, and the operator may see nothing but a series of reconnects. The mixed-version window is still there — it is simply someone else's to manage, and the tenant's job is to be sure its clients survive reconnects. ## The short version Rolling node by node converts an outage into a long interval of degraded-but-serving. The interval is not free: it is an interval in which the cluster is not one system but two releases pretending to be one, and the whole discipline of upgrading is about keeping that pretence honest until the last node is done.

  • What limits how long a cluster may be left part-way through a roll?
    Releases support a bounded version skew, not an unbounded one, so an abandoned roll can drift outside what was ever tested. Beyond that, the window doubles your debugging surface — two releases, two behaviours — and the ability to reinstall the previous release decays as new-format data accumulates. Most operators set a policy limit on how long a mixed-version window may stand and finish or revert inside it.
  • Does rolling node by node make the change invisible to clients?
    No. Each member's departure closes its connections, and clients reconnect elsewhere; in-flight requests to that member fail or are retried. A well-behaved client library handles that, but a client that treats a disconnect as fatal, or an application with no retry, will notice every wave. The roll removes the outage, not the churn.
  • Why can a cluster that is several releases behind not simply jump to the newest one?
    Interoperability between releases is usually guaranteed only across a bounded gap, because that is what was tested. Jumping further means nodes in the mixed-version window may not understand each other at all. The usual remedy is to roll to an intermediate release first, let the cluster settle at one version, and then roll again.

Repainting a bridge lane by lane keeps traffic moving, but for the whole job the bridge is half one colour and half another, and every sign has to make sense to drivers in both halves.

saying these in an interview costs you the question

  • Says the cluster can just be stopped and restarted because upgrades are quick
  • Believes rolling node by node means only one release is ever live
  • Thinks a roll is invisible to clients and needs no reconnect handling
  • Starts a roll with no spare capacity for the absent member
  • Assumes any release can run beside any other release indefinitely
  • Treats a rented cluster as having no mixed-version window at all
open as a page

During a two-pass roll, why is the agreed internal version raised only after the last broker node carries the new binaries?

level: middleimportance: must knowfreq 60%

basics

~20 s

Because the agreed internal version fixes how members speak to each other. Held at the old level, a new binary keeps talking the way an old one expects, so any node can still be reverted. Raising it early would leave old members unable to follow their peers.

open as a page

Half way through a roll the stored format version was raised; why does reinstalling the previous release now recover nothing?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Because the previous binaries cannot read records written in the newer layout. Reinstalling them recovers nothing written since the stored format version moved; the cluster has passed its abandon point, and the only way back is a restore or a second cluster, not a reinstall.

open as a page

Every broker node now runs the new release, but the capability the upgrade was for does nothing — why?

level: middleimportance: should knowfreq 42%

basics

~20 s

A capability that changes what members write or how they talk to each other arrives inert. It stays off until every broker node carries the new binaries and the agreed internal version has been raised, because a half-upgraded cluster could not tolerate it.

open as a page

What standing policy should an organisation set for how long a mixed-version window may stand and who may cross the abandon point?

level: principalimportance: should knowfreq 35%

basics

~20 s

Set a policy limit on how long a cluster may sit part-rolled, require every roll to finish or revert inside it, and make crossing the abandon point a named decision with an owner and a written recovery plan rather than a step in a runbook.

open as a page