skip to content

Why does one edited cluster-wide default change a running broker node's behaviour immediately while another waits for that node to restart?

level: middleimportance: must knowfreq 57%

answer

  1. not every value takes effect at once
  2. some are read only during startup
  3. an edit with no effect may be pending
  4. an unplanned restart applies it for you

basics

~20 s

Settings split into live-applying values, which a running broker node re-reads and acts on at once, and restart-only values, which the process reads while starting and holds for its lifetime. A restart-only edit is not ignored — it is pending until the next restart.

solid answer

~50 s

A broker node consults some values on every operation and others exactly once, while it starts. A **live-applying value** — a retention bound, a quota, a threshold — is re-read and takes effect on the running process, usually within seconds of being distributed. A **restart-only value** decides something the process builds at startup: where it listens, how its storage area is laid out, how large a fixed pool it allocates. Changing that cannot rewire a running process, so the edit sits **pending** and takes effect when that node next comes back. Two consequences matter operationally. First, an edit that shows no effect has not failed; it is armed. Second, the two kinds have different blast radius: a live value moves every node at roughly the same moment, while a restart-only value moves nodes one at a time as they restart, so the cluster spends a period running two different configurations. Which values fall into which class differs by platform and by release.

go deeper

for a junior

Know that some settings take effect on a running broker node and others wait for it to restart, and that an edit showing no immediate effect has usually not failed.

for a middle

Explain why the split exists — values read on each operation against values baked into structures built at startup — and say how you would check which class a value is in.

for a senior

Bring the operational consequence: a pending restart-only change fires at an unplanned restart, and a value set only on the running process reverts silently when that node comes back.

for a principal

Frame it as policy: whether a pending change may be left unapplied at all, and who is accountable when a node restarts into a configuration nobody currently on call chose.

## Two kinds of value, one settings surface An operator edits settings through one interface and reasonably expects one behaviour, but a broker node treats its settings as two different things. - A **live-applying value** is consulted repeatedly while the node serves traffic. The node picks up a new value without stopping — either because the change is distributed to it through whatever holds cluster configuration, or because it re-reads its own local source. Retention bounds, rate ceilings, thresholds that gate a behaviour and most numbers an operator tunes during an incident live here. - A **restart-only value** is read once, while the process starts, and is then baked into structures the process cannot rebuild without stopping: the addresses it listens on, the identity it presents to the rest of the cluster, how its storage area is laid out, the size of a pool allocated up front. Nothing about the edit itself tells you which kind you just made. The write is accepted either way. ## The trap: "nothing happened" is not "nothing was changed" When a restart-only value is edited, the running cluster keeps behaving exactly as before. The natural reading — *the change did not take* — is wrong and is the reason this material is asked. The change is **pending**: it is recorded, and it will be applied by each broker node at the moment that node next starts. That moment is rarely chosen. Nodes restart during a planned roll, but they also restart because a machine failed, because a platform moved them, or because someone was recovering from an unrelated incident. So a pending restart-only change is an instruction left for the worst possible moment: the node comes back **and** comes back different, and the operator handling the failure is now debugging two changes at once, only one of which they know about. The inverse is just as common. A value set directly on a running node, without also changing the source that node reads at startup, disappears the moment the node restarts — the fix that "worked" reverts, silently, at whatever hour the restart happens. ## Blast radius follows the class | Property | Live-applying value | Restart-only value | |---|---|---| | When it takes effect | On the running process, usually within seconds | When that node next starts | | How many nodes move together | Effectively all of them at once | One at a time, in restart order | | What the intermediate state looks like | Brief or none | The cluster runs two configurations until the last node is done | | What backing it out costs | Set the old value back | Another restart of every node already changed | This is why the class matters more than the value. A live-applying change is a cluster-wide event with a fast path back; a restart-only change is a multi-step procedure whose middle is a cluster deliberately running two configurations, and whose reversal is a second pass of the same length. ## Telling them apart before you apply 1. **Ask the software, not your memory.** Platforms publish which values may be changed on a running node, and the list changes between releases — a value that required a restart two releases ago may not today. 2. **Reason from the mechanism when the documentation is thin.** If the value decides something allocated or bound at startup, assume restart-only. If it is a number compared against on each operation, it is probably live. 3. **Prove it in a non-production cluster** by changing it and observing the behaviour it should move, rather than by reading the value back — reading it back usually shows the stored value, which is not the same as the value in force. 4. **Treat an accepted write as no evidence.** The interface acknowledging the change says only that it is stored. ## Where platforms differ - The **split itself is not standard**. The same conceptual setting may be live on one platform and restart-only on another, and several platforms have steadily moved values from the second class into the first across releases. - **Distribution differs.** Some platforms hold settings centrally and push a live change to every node; others expect each node to be told separately, which means a live change can reach nodes at meaningfully different times, and a node missed entirely is a difference nobody sees until it matters. - On a **rented cluster** the class is often invisible: the provider applies the change in its own way, and whether that involves restarting nodes for you is part of what you bought rather than something you control. - Some platforms let you set a value **on the running process only**, explicitly not persisted, which is useful during an incident and dangerous afterwards for exactly the revert-on-restart reason above.

  • You set a value directly on a running broker node during an incident and it fixed the problem. What must you do before you close the incident?
    Write the same value into whatever that node reads at startup, or the fix reverts the next time the node restarts — at an hour you did not choose, with nobody connecting the two events. If the value was meant to be temporary, remove it deliberately instead, and note which nodes carried it, because a value set on one node is a difference the cluster cannot show you later.
  • How does the class of a value change how you would back the change out?
    Backing out a live-applying value is a single edit that reaches the cluster as quickly as the original did. Backing out a restart-only value costs a second pass over every node you already restarted, which doubles the work and extends the period in which the cluster runs two configurations. That asymmetry is why restart-only changes are worth staging narrowly.
  • Does reading the setting back confirm that it is in force?
    Usually not. Most interfaces report the stored value, which for a pending restart-only change is the new one even though every running node is still behaving the old way. Confirmation comes from the behaviour the value is supposed to move — a bound that now bites, a ceiling that now rejects — or from asking the node what it is actually running, where the platform offers that.

saying these in an interview costs you the question

  • Assumes every setting takes effect the moment it is saved
  • Reads a change with no visible effect as a change that did nothing
  • Thinks reading the value back proves it is in force
  • Forgets that a value set on a running node can revert at restart
  • Believes the live and restart-only split is the same on every platform