skip to content

Upgrades & Configuration Change

Taking a live cluster from one version or setting to the next while it keeps serving, and where that change stops being reversible. Asked because a planned change is the outage whose date you chose.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

16

After a broker cluster moves to a newer release, why does an application's years-old client library keep working?

level: juniorimportance: must knowfreq 62%

answer

  1. agreed, not assumed
  2. settled per connection, at connect
  3. the older peer sets the ceiling
  4. old shapes still answered
  5. absence, not an error

basics

~20 s

Because the wire version is settled per connection, not per cluster. At the connect step each side says what it can speak, the connection uses a shape both understand, and the upgraded nodes still answer in that older shape.

solid answer

~50 s

A client and a broker node do not have to run the same release; they have to agree how to talk. On platforms that negotiate, the first exchange on a connection - the connect step - has the client advertise which revisions of each request shape it understands and the server answer with what it still supports, and the connection then speaks a shape both sides can express. Because the upgraded nodes keep serving older request shapes, the old library's requests are understood exactly as before. The agreement is bounded by whichever peer is older, so it is also why the old client gains nothing new from the upgrade: it simply keeps getting the older answers. The same agreement is re-settled on every reconnect, which is why a client upgrade takes effect only once its connections are re-established.

go deeper

for a junior

Recall that a client library and a broker cluster do not have to be on the same release, and that the two sides agree at connection time on a way of talking that both understand.

for a middle

Explain that the agreement is per connection and bounded by the older side, that upgraded nodes keep answering older request shapes, and that the agreement is re-settled only when a connection is re-established.

for a senior

Show you know the quiet failure mode: a connection on an older agreement looks perfectly healthy, so the absence of a capability produces no alert. Say how you would inventory which clients agreed what.

for a principal

Frame it as a support-span question for the estate: the platform states how far back it serves clients, and that span, not your schedule, sets when old applications must be funded to move.

## What is actually being agreed When an application's client library opens a connection to a broker node, the first useful thing that happens is not data movement. It is an agreement about **how these two will talk for the life of this connection**. On platforms that negotiate, the client sends what revisions of each request shape it understands; the server replies with what it supports; the connection then settles on shapes both can express. That outcome is the **wire version agreed per connection**, and everything afterwards on that connection is expressed in it. Two properties carry almost all the consequences: - It is **per connection**. It is not a cluster setting and not a property of the application as a whole. Two applications against the same cluster, and even two connections from the same process, can be on different agreements. - It is **bounded by the older peer**. The agreement can never exceed what the older of the two sides can express. A newer server cannot pull an old client forward, and a newer client cannot pull an old cluster forward. ## Why the old library keeps working A server release does not usually delete the older request shapes the moment it adds a new one. It keeps answering them. So when the years-old library sends the only shape it knows, the upgraded node recognises it and answers in kind. Nothing is translated or rewritten on the client's behalf - the older conversation is simply still on offer. How long that stays true is a **stated policy of each platform**, not a law: some publish how many releases back they serve, some state a time span, and some say far less. The operational consequence is the same either way: the cluster's support span, and not your appetite, is what decides when an old client finally has to move. ## What the old client gives up Working is not the same as keeping up. Anything the newer release added that the older request shape cannot express is **simply not present** on that connection. The application sees the older answers it has always seen. In the normal path nothing reports this - there is no error, because nothing failed. The visible failure mode is the opposite one: a client that *requires* a revision the server does not offer usually fails at the connect step, immediately and loudly, rather than halfway through a working day. ## The same rule in the other direction A newer client against an older cluster follows the identical rule and lands in the identical place: the agreement drops to what the cluster can express. What differs is the promise. Platforms generally make a clearer commitment to serving older clients than to accepting newer ones, so "new client, old cluster" is the direction more likely to be unsupported even when it appears to work. ## Where platforms differ This mechanism is not implemented one way across the class, and an answer that assumes one shape will be wrong somewhere: | How the contract is shaped | Where the version is settled | What an old client experiences | |---|---|---| | Per-request-type revisions, negotiated | At the connect step, per connection | Older shapes still answered; newer capabilities absent | | One protocol version per release | Advertised by the server; client must fall in the supported band | Connection refused once it falls outside the band | | A versioned remote endpoint the caller chooses | In the address the client calls | The older endpoint keeps its behaviour until it is retired | | A rented cluster on a hosted tier | Invisible to you; the provider states a supported client range | Whatever the provider's support statement allows | The agreement is also entirely separate from whatever secures the transport; those are settled independently, and one says nothing about the other. ## What this means operationally 1. **Inventory clients by connection, not by team.** The thing that decides behaviour is what each connection agreed, so a per-application list of client library versions is the artefact you actually need. 2. **Expect no signal from a downgraded agreement.** A connection quietly speaking an older shape looks healthy on every dashboard, because it is healthy. 3. **Remember the re-settle.** Upgrading a client library changes nothing until its connections are re-established; a long-lived connection keeps its old agreement until it drops. 4. **Read the support span before the upgrade, not after.** The question "how far back does this release serve?" is the one that tells you whether an application is about to stop connecting. The short version a candidate should be able to say out loud: releases do not have to match, because the two sides agree a shape at connect time; the older side bounds that agreement; and the cost of an old client is not breakage but absence.

  • An application upgrades its client library but behaves exactly as before. What is the most likely reason?
    Its connections were never re-established. The agreement is settled at the connect step and holds for that connection's life, so a long-lived connection keeps the old agreement until it drops and reconnects. Restart the application, or wait for its connections to cycle, before concluding the upgrade did nothing.
  • Does a newer client against an older broker cluster work the same way?
    The mechanism is the same - the agreement drops to what the older peer can express - but the promise is weaker. Platforms usually commit to serving older clients and say much less about accepting newer ones, so this direction can appear to work while being outside what the platform supports.
  • Why does a client that demands an unavailable revision fail at connect rather than later?
    Because the agreement is settled before application traffic flows. If the client cannot express itself within what the server offers, that is knowable immediately, so the connection is refused there. It is the friendlier failure: loud, at startup, and attributable to the version gap rather than to a request halfway through the day.

Two people who speak different amounts of a shared language work out in the first minute which words both know, and stay there. Nobody is wrong, and the richer speaker simply never uses the sentences the other could not follow.

saying these in an interview costs you the question

  • Thinks the client library must match the broker release exactly.
  • Assumes the wire version is one cluster-wide setting rather than per connection.
  • Believes an old client corrupts data instead of being served an older shape.
  • Thinks upgrading the cluster upgrades what a connected client may use.
  • Assumes the wire version is re-negotiated on every request.
  • Expects a clear error when an old client cannot use a new capability.
open as a page

A value has a cluster-wide default and a per-stream override on one stream — which value does that stream use, and which do the rest?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The override wins for that one stream; every stream or queue without an override follows the cluster-wide default. The effective value is therefore a per-object answer, and editing the default leaves overridden objects exactly where they were.

open as a page

Why is a new broker release rolled out one broker node at a time, and what runs side by side while it is?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Taking one broker node at a time keeps every other member serving, so the cluster stays available through the change. The price is a mixed-version window: until the last node is done, two broker releases are live in the same cluster.

open as a page

Why does one edited cluster-wide default change a running broker node's behaviour immediately while another waits for that node to restart?

level: middleimportance: must knowfreq 57%

basics

~20 s

Settings split into live-applying values, which a running broker node re-reads and acts on at once, and restart-only values, which the process reads while starting and holds for its lifetime. A restart-only edit is not ignored — it is pending until the next restart.

open as a page

During a two-pass roll, why is the agreed internal version raised only after the last broker node carries the new binaries?

level: middleimportance: must knowfreq 60%

basics

~20 s

Because the agreed internal version fixes how members speak to each other. Held at the old level, a new binary keeps talking the way an old one expects, so any node can still be reverted. Raising it early would leave old members unable to follow their peers.

open as a page

When moving a broker cluster and its applications to a newer release, which side goes first and why is that not a preference?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Broker nodes first, client libraries after. The wire version a connection agrees is capped by the server, so upgrading clients first buys nothing and can refuse connections, while upgrading servers first keeps every old client served and unblocks the clients that follow.

open as a page

Half way through a roll the stored format version was raised; why does reinstalling the previous release now recover nothing?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Because the previous binaries cannot read records written in the newer layout. Reinstalling them recovers nothing written since the stored format version moved; the cluster has passed its abandon point, and the only way back is a restore or a second cluster, not a reinstall.

open as a page

A new broker capability is enabled cluster-wide, yet one application never gets it and logs no error - why?

level: middleimportance: should knowfreq 55%

basics

~20 s

That application's connections agreed an older wire version, which cannot express the capability. The server answers in the older shape and simply withholds what that shape has no room for. Nothing failed, so nothing is logged.

open as a page

Expanding a live cluster is approved with a date attached: what decides whether added capacity earns work at once or only after a deliberate move?

level: middleimportance: should knowfreq 52%

basics

~20 s

Consumption shape decides. Competing consumers pulling from a shared queue absorb new capacity as soon as they reach it; a stream split into fixed parts earns a new broker node nothing until parts are assigned to it.

open as a page

Every broker node now runs the new release, but the capability the upgrade was for does nothing — why?

level: middleimportance: should knowfreq 42%

basics

~20 s

A capability that changes what members write or how they talk to each other arrives inert. It stays off until every broker node carries the new binaries and the agreed internal version has been raised, because a half-upgraded cluster could not tolerate it.

open as a page

A broker node was edited by hand months ago and behaves exactly like its peers — when does that difference surface, and how would you find it first?

level: seniorimportance: should knowfreq 46%

basics

~20 s

It surfaces at that node's next restart — usually unplanned, during a failure or a roll — because the running process holds values it read at startup. Find it first by asking every node for its effective values and comparing those against the declared estate.

open as a page

On a live broker cluster, how do you stage a value change so that a way back exists before you apply it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Record the current effective values per object, decide whether setting the old value back actually undoes the effect, apply at the narrowest scope that proves the change, and watch the behaviour it should move plus the one it could break. Staging is about reversibility, not pace.

open as a page

A team wants to expand a live cluster during a scheduled change window at 3 a.m. What does the quiet hour actually reduce, and what does it not?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A quiet hour reduces exposure, not reversibility: fewer clients affected, more spare headroom. A step that cannot be undone at noon cannot be undone at 3 a.m., so approval still turns on the half-finished state.

open as a page

As platform lead, what standing policy limit would you set on how far behind a broker client library may be?

level: principalimportance: should knowfreq 42%

basics

~20 s

Publish a floor expressed in releases or months behind, derived from the span your server release actually serves, then make it enforceable: a per-application report of who is below it, a dated cut-off, an owner for every client library, and refusal at connect as the final step.

open as a page

What standing policy should an organisation set for how long a mixed-version window may stand and who may cross the abandon point?

level: principalimportance: should knowfreq 35%

basics

~20 s

Set a policy limit on how long a cluster may sit part-rolled, require every roll to finish or revert inside it, and make crossing the abandon point a named decision with an owner and a written recovery plan rather than a step in a runbook.

open as a page

Who should be allowed to approve expanding a live cluster under traffic, and on what stated evidence should they refuse?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

The cluster's operating owner together with the owner of the affected workloads — not the requester alone. Refuse unless waiting carries a named, dated risk, someone can say when the capacity earns work, and the half-finished state has an answer.

open as a page