skip to content

questions

4

During a broker credential swap with traffic flowing, what must be true of every node before any client moves to the new credential?

level: middleimportance: must knowfreq 66%

answer

  1. trust first, then present
  2. every node, not most of them
  3. two rollouts, two failure modes
  4. narrow the trust set last

basics

~20 s

Every node must already accept both the outgoing and the incoming credential. Clients reconnect to whichever member answers, so a trust set widened on only some nodes produces intermittent refusals rather than a clean, visible break.

solid answer

~40 s

The swap runs in three stages and the node half comes first. Widen the trust set on every node so both the outgoing and the incoming credential authenticate — every node, because a client is given a cluster address and may reconnect to any member, so a half-widened trust set fails intermittently and looks like a network fault. Then move the clients onto the new credential at whatever pace the fleet deploys; throughout that stage both credentials work. Only when nothing presents the old one do you narrow the trust set again. The node half and the client half are separate rollouts with separate failure modes and separate evidence. Whether a node picks up new trust material without a restart differs between platforms, so establish that before scheduling anything.

go deeper

for a junior

Remember the direction of travel: the cluster learns to accept the new credential before anybody presents it. Trust is widened first and narrowed last.

for a middle

Explain why every node matters. A client reconnects to whichever member answers, so a partly widened trust set fails intermittently and disguises itself as a network fault.

for a senior

Show that you verify the halves separately — both credentials authenticating against each member, then each principal observed authenticating with the new one — rather than trusting a green pipeline.

for a principal

Own the preconditions. Whether nodes reload trust material without a restart, and whether the replacement keeps the same principal, decide whether this is a short change or a cluster-wide operation.

## The swap is three stages, and the order is fixed A credential used in front of a broker cluster is **accepted** by the nodes and **presented** by the clients. Those are two populations that change at different speeds, and keeping traffic flowing means never asking them to change at the same instant. The order is: 1. **Widen the trust set on every node** so that both the outgoing and the incoming credential authenticate. Nothing presents the new one yet — the cluster is merely prepared to accept it. 2. **Move the clients onto the new credential**, at whatever pace the fleet deploys. Throughout this stage some clients present the old credential and some the new one, and both succeed. 3. **Narrow the trust set** so only the new credential authenticates. This is the step that can break something, and it runs last, on evidence rather than on a date. Reverse any pair and you get the outage the exercise exists to avoid. Narrow first and every holder is refused at once. Move clients first and they present something no node will accept. ## Why 'every node' is the part that bites A client is normally handed an address for the cluster rather than for one member, and it reconnects to whichever member answers — after its own restart, after a blip, after a member is patched. The trust set is therefore not one thing you change; it is the same change applied to every member, and until the last member has it, the new credential works or fails depending on where a connection happens to land. The symptom is the awkward one. Not a clean break that points straight at the change, but a small share of connection failures that moves around as clients reconnect, and that reads like a network problem to everyone looking at it. ## Two rollouts, two verifications The node half and the client half fail in different places, so they are tracked and proved separately: | Stage | Where it lands | What proves it done | What its failure looks like | |---|---|---|---| | Widening the trust set | every node | both credentials authenticate against each member | intermittent refusals that follow reconnections | | Moving the clients | every client process | each principal seen authenticating with the new credential | one service left behind, quiet until it reconnects | | Narrowing the trust set | every node | the old credential is refused everywhere | the outage you were avoiding, arriving later | Tracking them as one change is how the second half gets assumed rather than checked. A pipeline reporting a successful deploy says the new credential is in place in that service; it does not say the service has opened a connection with it. ## The node-to-node hop is a third rollout Where members authenticate to one another, each node is both a presenter and an acceptor on the node-to-node hop, and the same three stages apply to it. The failure is different in kind: a mistake there does not refuse a client, it stops members talking, which halts replication and can stop the cluster accepting writes at all. Sequence and verify that hop on its own rather than folding it into the client rollout because the material looks the same. ## Three preconditions that vary, and must be settled first - **Whether a node picks up new trust material without a restart.** Some platforms re-read it periodically or on an explicit reload; others read it once at start. If yours is the second kind, the node half becomes a touch-every-member operation with all the care that implies, and that changes the plan rather than the order. - **Whether the replacement yields the same principal.** If the identity the broker derives from the new credential is unchanged, nothing downstream has to move. If it differs, that principal's grants must already be in place before any client switches — what those grants should say is a separate subject, but their existence is a precondition of this swap. - **Whether the two credentials can be told apart afterwards.** Stage three is decided on evidence, and that evidence is far weaker if the cluster's records cannot say which of the two a client presented. ## Why this is sharper on a broker than in front of a request-per-connection service Two properties do it. First, the acceptor is a cluster rather than a process, so the change is many applications of the same edit and is routinely left partly done. Second, clients hold long connections and are not asked to prove themselves again, so the client half completes invisibly and late. Widen everywhere, move, then narrow is the order that absorbs both.

  • Does widening a node's trust set take effect without restarting the node?
    It differs by platform. Some re-read trust material on a schedule or on an explicit reload; others read it once at start, which turns the node half into a touch-every-member operation with the precautions that implies. Establish which you have before the swap is scheduled, not during it.
  • The nodes authenticate to one another as well. What does that add to the swap?
    A third rollout. On the node-to-node hop each member is both presenter and acceptor, so a mistake there does not refuse a client — it stops members talking, which halts replication and can stop the cluster accepting writes. Sequence and verify it separately from the client hop.
  • Does the replacement have to authenticate as the same principal as the old credential?
    It is far simpler if it does, because nothing downstream changes. If the identity the broker derives from it differs, the new principal needs its grants already in place before clients move. What those grants should say is another subject; having them exist is a precondition of this sequence.

saying these in an interview costs you the question

  • Swaps the credential everywhere in one coordinated instant
  • Widens the trust set on the one node an operator happened to test
  • Treats a finished deploy as proof the clients moved
  • Withdraws the old credential before any evidence is gathered
  • Assumes the node-to-node hop is covered by the client rollout
open as a page

Before withdrawing a broker's old credential, what evidence shows nothing still presents it, and what will that evidence miss?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The cluster's own connection and authorization-decision records, read over an observation window longer than the slowest client's reconnect interval, and only if the two credentials are distinguishable in them. They miss every holder that has not connected during that window.

open as a page

Why does a broker credential reaching its expiry date take a cluster down all at once rather than degrading gradually?

level: juniorimportance: should knowfreq 54%

basics

~20 s

Because expiry is a date shared by everything issued in the same batch: every holder loses the right to connect at the same moment. Where the node-to-node hop uses material from that batch, replication stops alongside the clients.

open as a page

A broker's old credential was withdrawn on Monday with no errors, yet a service redeployed on Wednesday cannot connect — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

That service's connection predated the withdrawal and was never re-authenticated, so it kept working on a credential the cluster no longer accepts. The redeploy forced a fresh connection, which is the first moment the withdrawal was actually tested for it.

open as a page