When moving a broker cluster and its applications to a newer release, which side goes first and why is that not a preference?
answer
- the server is the ceiling
- servers first, clients after
- clients first gains nothing
- refused at connect, not mid-stream
- the unowned library is the real constraint
basics
~20 sBroker nodes first, client libraries after. The wire version a connection agrees is capped by the server, so upgrading clients first buys nothing and can refuse connections, while upgrading servers first keeps every old client served and unblocks the clients that follow.
solid answer
~50 sThe order falls out of the agreement rather than out of taste. A connection can only settle on a wire version both peers can express, so the server's ceiling bounds what any client may use. Move the cluster first and every existing client keeps working exactly as before - the agreement simply stays where it was - while newer clients become able to ask for more as they arrive. Move the clients first and the best case is that they negotiate down and gain nothing; the worse case is a client that requires something the older cluster cannot offer and is refused at its connect step. The reason this is a real interview question is that the binding constraint is rarely technical: it is an unowned client library pinned years ago inside a live application, which turns "which side first" into "how long must the cluster keep serving that one client".
go deeper
Recall that the cluster is taken to the new release first and applications follow, and that old client libraries keep working against an upgraded cluster rather than breaking on the day.
Explain why the order is forced: the agreed wire version cannot exceed what the server can express, so the server's ceiling bounds every client and upgrading clients first changes nothing.
Show the production judgment: inventory client versions against the span the target release serves, name the application that falls outside it, and say what you do about a library with no owner.
Argue the standing policy. Per-release firefighting over old clients is a symptom; the estate needs a published floor on client age, an owner for every dependency, and funded time to hold it.
## Why the order is forced A client and a broker node settle a wire version at the connect step, and that agreement can contain only what **both** sides can express. The server is therefore the ceiling for every connection into the cluster. Two consequences follow immediately, and together they fix the order: - **Upgrading the server does not break old clients**, because the newer release keeps serving the older request shapes for a stated span. The agreement those clients reach is unchanged, so their behaviour is unchanged. - **Upgrading the client does not unlock anything** until the server can meet it. The new library negotiates down to what the cluster offers, so the work lands with no visible result - and if the library *requires* something the cluster cannot offer, it is refused at connect instead. So: **servers first, clients after.** Not a convention; a consequence of who bounds the agreement. ## What each order actually produces | Order | Immediate effect | Risk | |---|---|---| | Cluster first, then clients | Old clients unchanged; new client capability becomes reachable as each application upgrades | An application already past the cluster's supported span may stop connecting | | Clients first, then cluster | Clients negotiate down; no new capability appears | A client demanding an unavailable revision is refused at connect; teams conclude the upgrade "did nothing" | | Both at once | Nothing gained over cluster-first | Two changes in flight, and an incident cannot be attributed to either | The third row is the one candidates offer most readily and it is the weakest: simultaneity adds no benefit here, because the cluster-first order already leaves every client working. ## The constraint that actually binds In an interview, the order is the easy half. The real material is what stops you executing it: **an unowned client library**. A dependency was pinned years ago, the team that owned the application has been reorganised away, the service still carries real traffic, and nobody has authority to change its dependencies. That library, not your plan, decides how long the cluster must keep serving an old agreement. That turns the question from a per-release one into a **standing** one. Each new server release ships with a support span; each unowned client eats into it; and the day a release finally drops support for that client's shape is the day the application stops connecting - unless someone was made to act months earlier. Practical handling: 1. **Inventory before you plan.** For each application: which client library version, when its connections last re-established, and whether anyone owns it. 2. **Compare that inventory against the span** the target server release states it serves. Anything outside it is not an upgrade task, it is an outage with a date. 3. **Do not let an unowned library veto the cluster upgrade silently.** Name it, cost it, and take the decision explicitly - fund the upgrade, re-home the application, or accept the cut-off. 4. **Re-establish connections deliberately after a client upgrade.** Until a connection is re-made, the old agreement stands and the upgrade is invisible. ## Where platforms differ The rule that the older peer bounds the agreement is general, but its edges are not: - Platforms **state different spans** for how far back a server release serves clients - some several releases, some a time window, some little at all. - Support for the **reverse direction** - a newer client against an older cluster - is much less consistently promised, which is another reason not to lead with clients. - Where the surface is a **versioned remote endpoint** rather than a negotiated agreement, the caller effectively pins its own version, and the ordering question becomes when the old endpoint is retired. - On a **rented cluster**, the provider may upgrade the servers on its own schedule and simply publish a supported client range, in which case your only lever is the client side and the order is decided for you. ## The answer an interviewer is listening for Three beats. First, the rule: the agreement is capped by the server, therefore servers first. Second, the symmetry check: upgrading clients first is not dangerous so much as pointless, with a refusal at connect as the downside case. Third, and the one that separates senior from middle: the order is easy and the **sequencing constraint is organisational** - a client library nobody owns, which is why how-far-behind-may-a-client-be has to be a standing rule rather than a decision taken per release.
- Is upgrading the client libraries first actually dangerous, or just wasted effort?Usually wasted effort: the connection negotiates down and the application behaves as before, so teams conclude the upgrade did nothing. It becomes dangerous only when a newer library requires a revision the older cluster cannot offer, in which case the connection is refused at the connect step - loudly, at startup.
- An application's client library is outside the span the target server release serves. What do you do?Stop treating it as an upgrade task. It is a dated outage, so surface it before the cluster moves: either fund the library upgrade, re-home the application to an owner who can, or take an explicit decision to cut it off. Discovering it during the roll gives you none of those choices.
- Why is 'upgrade both sides in one change window' a weak answer?Because it buys nothing the cluster-first order does not already give you - old clients keep working either way - while putting two changes in flight at once, so any incident cannot be attributed to one of them. Simultaneity is a cost here with no matching benefit.
You cannot usefully teach a team new vocabulary before the person they report to understands it; they will keep having the old conversation. Widen the listener's range first, and the speakers become useful the moment they learn.
saying these in an interview costs you the question
- Says the order is a matter of team preference.
- Upgrades client libraries first so applications are 'ready'.
- Assumes a newer client can use newer capabilities against an older cluster.
- Thinks one old client caps what every other connection may use.
- Treats client version skew as a per-release problem, not a standing one.
- Plans the order without inventorying which client versions are live.