An alert on a reader group's record lag fires on a single sample and pages nightly on write bursts; what does a usable condition need instead?
answer
- signal, threshold, window
- one sample catches a staircase
- batches make the number step
- window outlasts write-and-drain cycle
- the window is detection delay
basics
~20 sA usable condition names three things: which signal it reads, the threshold that counts as a breach, and a sustained window the breach must hold across. Size that window longer than one full write-and-drain cycle of the flow it watches.
solid answer
~50 sAn alert condition is three choices, not one number: the signal, the threshold, and the sustained window the breach must hold continuously before the rule fires. A reader group's record lag moves in steps by construction - writers publish in batches and readers fetch in batches, so between two fetches the unread backlog climbs by everything written in the gap and then falls in one step. A rule that compares one sample against a threshold is sampling that staircase at a random phase, which is why it pages on a burst that drained by itself. Choose the window from the flow's own rhythm: longer than the writer's burst interval plus the reader's fetch-and-process cycle, so only a breach that survives a whole cycle counts. The window's main cost is detection delay - a reader that stops dead is reported only after it elapses.
go deeper
Remember the three parts by name: which measurement, what value counts as bad, and how long it has to stay bad. Being able to say that a one-off spike is not an incident is most of the credit at this level.
Explain why these numbers move in steps - batched writes and batched reads - and derive the window from that cycle rather than guessing. Say out loud that the window is detection delay, so the choice is a trade and not a free improvement.
Show that you back-test a candidate condition against history before it reaches the pager, and that you have a stance on missing samples and on a breach that flaps either side of the threshold.
The interesting call is which unit the estate standardises on. Age-based thresholds survive traffic growth and express the business question directly; record-count thresholds need re-choosing every time throughput moves, and nobody re-chooses them.
## An alert condition is three choices Teams talk about an alert as though it were a number. It is three separate decisions, and each one can be wrong on its own: - **The signal** - which measurement the rule reads, named precisely enough that two people would pick the same series: a reader group's record lag, the age of its oldest unread record, the copies-behind count, the wait time a request spends in a node's queue. - **The threshold** - the value that counts as a breach, together with its unit. `2,000,000 records` and `90 seconds of age` are different conditions on the same reader, and they fail at different moments. - **The sustained window** - how long the breach has to hold continuously before the condition fires at all. A rule with a signal and a threshold but no window is not a weaker alert; it is a different and mostly useless one. ## Why one sample proves nothing here Messaging signals move in steps, and how sharply depends on how the flow batches. Writers accumulate records and publish them together; readers fetch a batch, process it, and only then record their progress. Between two fetches, a reader's unread backlog grows by everything written in the gap, then collapses when the next batch lands. Plotted per second, a perfectly healthy reader draws a staircase whose peaks are the batch size. There is a second, more mechanical reason. Where the reader owns a stored position, lag is not measured directly - it is derived from the distance between the stream's newest record and that stored position, and the two are refreshed on different schedules. A sample taken between those refreshes can show a step that no reader ever experienced. So a single-sample rule fires on the phase of a staircase. The on-call learns within a week that the page means nothing, and the rule is then worse than no rule: it occupies the attention budget that a real condition would need. ## Choosing the window Derive it, do not round it to five minutes because five minutes looks tidy: 1. Measure the writer's burst interval - how long between the batches that arrive on this stream. 2. Measure the reader's cycle - fetch, process, record progress, fetch again. 3. Set the window longer than one full cycle of the two together, and preferably around two, so a breach must survive a complete drain to count. 4. Replay a week of history against the candidate condition and count how many times it would have fired. A condition nobody has back-tested is a guess. The same arithmetic sets the threshold's unit. A record-count threshold has to be re-chosen every time throughput changes; an age-based threshold survives a traffic change because it already asks the question the business asks - how stale is the oldest thing we have not handled. ## The window is detection delay | Window | What it does | What it costs | |---|---|---| | Shorter than one cycle | Fires on every normal batch | Trust; the page is ignored inside a week | | Around one to two cycles | Fires only on a breach that survives a drain | A detection delay equal to the window | | Much longer than the cycle | Very quiet | A reader that stopped dead goes unreported for the whole window | There is no setting that is both instant and quiet, and pretending otherwise is how a team ends up with a rule that fires at three in the morning and a rule that never fires, both watching the same signal. Pick the delay you can live with and write it down as a deliberate choice. Two practical details belong to the condition, not to the dashboard. First, decide what a gap in the data means: if the collector misses samples, does the breach count as broken, or does the condition hold? A condition that treats missing data as not-breaching goes silent exactly when the cluster is in trouble. Second, decide whether the window resets when a single sample dips below the threshold. Most evaluators require the breach to be continuous, which means a reader flapping either side of the line can stay under the alert forever; where that pattern matters, express the condition as a fraction of breaching samples inside the window instead. ## What varies across platforms The shape carries; the signal does not. Where readers own a durable position, lag comes in records and in age. Where the broker deletes a record once it is acknowledged, there is no position to subtract from, and the equivalent signal is the depth and the age of the unread backlog - but it is still a step-shaped number that needs a sustained window. Some platforms publish the number per reader group and per partition, others only a single coarse figure per queue; with a per-partition signal, the condition must also say whether it fires on any one part or on the total, and those two rules behave completely differently when one part is stuck.
- The breach dips below the threshold for one sample inside the window - should the condition restart its timer?By default most evaluators require the breach to be continuous, so one dipping sample resets it. That is usually what you want, but it means a reader oscillating either side of the line never fires at all. If that pattern matters on your stream, express the condition as a fraction of breaching samples within the window rather than an unbroken run.
- The signal is only collected once a minute. What does that do to a two-minute sustained window?It leaves the condition resting on two samples, so one missed collection either fires it or silences it. Make the window hold several samples, and state explicitly what missing data means - a condition that treats a gap as not-breaching goes quiet precisely when the cluster is unhealthy.
- Which signals deserve a much shorter window than the flow's rhythm suggests?Those whose value is never legitimate rather than merely high. A partition with no node serving it is not a large number, it is a state that should not exist, so waiting several minutes to confirm it buys nothing. Reserve long windows for signals whose normal range genuinely includes the threshold.
saying these in an interview costs you the question
- Treats the threshold number as the entire alert condition
- Fires the page on the first sample that crosses the line
- Assumes a longer sustained window costs nothing
- Picks a round window with no reference to the flow's cycle
- Never replays the condition against history before enabling it
- Leaves missing data undefined, so gaps silence the rule