Every planned node-by-node cluster change pushes the copies-behind signal past its threshold and pages the on-call, so the team keeps raising the threshold; what should they change instead?
answer
- quiet now, blind later
- two buckets, opposite treatment
- expected to move or never
- window sized to one node's catch-up
- suppression scoped and expiring
basics
~20 sRaising the threshold trades away the detection the alert existed for at every other hour. Split the signals a planned change is expected to move from those it never should, size the sustained window past the normal catch-up time, and bound any suppression to the announced change window.
solid answer
~50 sRaising the threshold is a repair aimed at the wrong half of the problem: it buys quiet during a planned change by deleting detection during the other 700 hours of the month, and nobody ever lowers it again. Start by splitting the signals into two buckets. Some legitimately move while a cluster is changed node by node - the copies-behind count, per-node request latency, the distribution of leadership - and some never should, such as a partition with no node serving it at all. For the first bucket, set the sustained window longer than the time a node normally takes to rejoin and catch up, so a shortfall that persists past that is genuinely abnormal. Keep the second bucket tight, because those conditions should still page mid-change. Only then, if it still fires, suppress narrowly for the announced change window - scoped to that signal and cluster, with an expiry.
go deeper
The point to hold on to is that an alert firing during planned work is not proof the alert is wrong, and turning its number up is not a fix - it changes the rule for every other day too.
Explain the split: some signals are expected to move during planned work, some never should, and the two need opposite treatment. Be able to say why the sustained window is sized against one node's catch-up rather than the whole change.
Show that you measure catch-up rather than guess it, and that you keep a small set of conditions deliberately paging during a change because that is when they matter most. Name the cost of the threshold edit explicitly.
The estate-level question is how the permanent and the temporary tools are kept apart. Threshold edits are permanent and invisible; make them reviewable, and make the bounded, expiring suppression the path of least resistance.
## What raising the threshold actually buys The change is announced, the operator restarts nodes one at a time, the copies-behind signal rises, the pager goes off, and somebody edits the rule. It works, in the sense that the next change is quiet. What it costs is invisible and permanent: the condition that used to notice a replication shortfall at two in the afternoon now needs a much larger one, and the number is never revisited because it is no longer causing pain. Several rounds of this produce a rule whose threshold no real incident can reach. The same logic applies to stretching the sustained window until the change fits inside it. A window long enough to cover a whole rolling change is a window long enough to hide a genuine shortfall for the same duration. ## Two buckets of signal The useful first move is to sort the conditions by whether a planned change is *expected* to move them. | Signal class | Moves during a planned node-by-node change? | Condition treatment | |---|---|---| | Replication shortfall - copies behind | Yes, by design, once per node | Sustained window sized past normal catch-up | | Per-node request latency and saturation | Yes, while load shifts to the remaining nodes | Sustained window, and a threshold set from the reduced-capacity case | | Leadership distribution across nodes | Yes, and it may stay skewed afterwards | Not a page at all; a check after the change completes | | A partition with no node serving it | No - this should not happen even mid-change | Keep it tight, keep it paging | | Writes being refused, or a reader making no progress at all | No - the point of a rolling change is that service continues | Keep it tight, keep it paging | That table is the actual answer to the question. The team has been applying one repair to every row, when the rows need opposite treatment: the top three want a condition that tolerates a bounded, self-healing excursion, and the bottom two want a condition that fires *especially* during a change, because a planned change is exactly when they are most likely to be violated. ## Sizing the window against catch-up For the tolerant bucket, the window is not guesswork: 1. Time a single node's rejoin-and-catch-up on a low-risk change, and record it. 2. Set the window comfortably above the observed worst case for one node - not for the whole change, which is a different and much longer number. 3. Re-measure as the data per node grows. Catch-up time scales with what a node has to fetch, so a window set when the cluster held a tenth of today's data is already wrong. 4. If a single node's catch-up is so long that the required window is unacceptable, that is a capacity finding, not an alerting finding, and it belongs in a different conversation. A shortfall that survives that window during a change is real: it means the node that restarted is not catching up, which is the thing you wanted to know and which the raised threshold would have concealed. ## The bounded suppression, and its limits When the window alone is not enough, suppressing the tolerant bucket for the announced change window is legitimate. Three properties make it safe rather than dangerous: it is scoped to the specific signal and cluster rather than the whole estate, it expires on its own, and it never covers the intolerant bucket. A suppression that silences everything and outlives the change is an outage you have arranged not to hear about. The discipline worth stating out loud is that the temporary and the permanent must not be confused. A threshold change is permanent by default. A suppression is temporary by default. When a team reaches for the permanent tool to solve a temporary problem, the estate slowly loses its alerting without a single decision being recorded. ## What varies across platforms How much a planned change moves these signals depends entirely on the design underneath. On platforms where each partition has a leader and a set of followers that must catch up, restarting a node guarantees a visible shortfall and a visible leadership shift. On platforms where writes are accepted by a majority, the shortfall may never surface as an operator-facing number at all. Where storage is shared rather than replicated per node, there is no catch-up phase to wait out, and the signal that moves is request latency instead. On a rented cluster, the provider may be performing its own changes without telling you, which is its own reason to prefer conditions that tolerate a bounded excursion over conditions that fire on any deviation.
- Nobody knows how long a node takes to catch up. How do you get the number?Measure it during the next low-risk change: restart one node and record how long the shortfall takes to clear. That single observation is worth more than any published figure, because catch-up time depends on how much data that node holds and how fast it can refetch. Re-measure as the cluster grows, since the number moves with the data.
- Which conditions should still page while a change is in progress?The ones whose breach means the change is going wrong: a partition with no node serving it, writes being refused, a reader making no progress at all. A rolling change is supposed to keep the cluster serving, so those conditions are most valuable exactly during it - they are the ones that tell the operator to stop.
- The team suppressed the whole cluster's alerts for a two-hour change that overran to nine hours. What went wrong in the condition design?Two things: the suppression was scoped to a cluster rather than to the signals a change is expected to move, and its lifetime was tied to a plan instead of an expiry. Scope it to the tolerant bucket, expire it on its own, and let the intolerant bucket keep paging so an overrun is still visible.
saying these in an interview costs you the question
- Raises the threshold permanently to silence change-window pages
- Stretches the sustained window to cover the whole rolling change
- Suppresses every alert on the cluster while a change runs
- Treats an unserved partition as expected noise during a change
- Sizes the window without measuring one node's catch-up time
- Never lowers the threshold again once the change is over