skip to content

A paging rule that fires when a service's p99 latency exceeds 500 ms re-fires and resolves five times during every traffic peak. Which rule-level knobs would you reach for to damp that, and what does each one cost?

level: middleimportance: should knowfreq 52%

answer

  1. it keeps crossing the same line
  2. one crossing is not an event
  3. require continuous truth first
  4. hysteresis on the way back down
  5. quiet bought with detection latency

basics

~20 s

Require the condition to hold continuously before firing — a pending or 'for' duration — and keep the alert firing for a defined period after it clears, so brief recoveries do not close and reopen it. The first adds exactly that much detection delay; the second delays the resolved notification.

solid answer

~60 s

Flapping means the signal is crossing a single threshold repeatedly, so the fix is to stop treating an instantaneous crossing as an event. The primary knob is the pending duration — Prometheus-style rules spell it `for`, Grafana calls it the pending period — which requires the expression to evaluate true on every evaluation across that window before the alert becomes firing. Set it longer than the natural oscillation and the peak stops paging. The cost is precise and worth stating: you have added up to that duration to your detection time, so a `for: 10m` on a rule protecting a five-minute response target is wrong. The second knob is hysteresis on the recovery side — keeping the alert firing for a defined period after the condition clears — so a momentary dip does not resolve and immediately re-fire. The third is the threshold itself: if 500 ms is inside normal peak behaviour, the number is wrong rather than the timing, and it should sit where a human would actually act. Longer averaging in the underlying query does similar work, but that belongs to whoever owns the metric pipeline.

code

yaml · 11 lines
yaml
groups:
  - name: latency
    rules:
      - alert: HighRequestLatency
        expr: job:request_latency_p99_seconds:5m{job="checkout"} > 0.5
        for: 10m
        keep_firing_for: 15m
        labels:
          severity: page
        annotations:
          summary: "checkout p99 latency above 500ms for 10m"

go deeper

for a junior

Know that alert rules can require a condition to stay true for a period before firing, and be able to name that a rule which fires and resolves repeatedly is called flapping.

for a middle

Explain the mechanics: pending state, the reset on any false evaluation, and hysteresis on the recovery edge. State plainly that a pending duration converts directly into detection delay, second for second.

for a senior

Show that you tune from measurement — the observed oscillation length and whether a human ever acted at that threshold — and that you check the resulting detection time against the response target the severity promises before shipping the change.

for a principal

Own the policy: what maximum pending duration a paging rule may carry given the response commitment, whether thresholds must be derived from objectives rather than round numbers, and how the platform surfaces chronically flapping rules for retuning.

## What flapping actually is An alert flaps when the underlying signal sits close to the threshold and crosses it repeatedly. Each crossing is a state transition, each transition is a notification, and the on-call ends up with five pages and five resolutions for one afternoon of ordinary load. Nothing is wrong with the monitoring — the metric is faithfully reporting reality. What is wrong is the assumption baked into the rule: that an instantaneous crossing is an event worth waking someone for. Flapping is also corrosive out of proportion to its volume, because it trains responders to ignore the rule. A rule that has resolved itself the last nine times gets acknowledged and dismissed the tenth time too — and the tenth time is the outage. ## Knob 1: the pending duration The standard instrument is a required duration of continuous truth before the alert fires. In Prometheus-style alerting rules the field is `for`; Grafana's unified alerting calls the same idea a pending period. The rule sits in a pending state while the expression is true and only becomes firing — and therefore notifying — if it is still true when the duration elapses. Any evaluation that comes back false resets it to inactive. ```yaml groups: - name: latency rules: - alert: HighRequestLatency expr: job:request_latency_p99_seconds:5m{job="checkout"} > 0.5 for: 10m labels: severity: page annotations: summary: "checkout p99 latency above 500ms for 10m" ``` Choose the duration against the oscillation you measured, not by habit. If the signal spends two to three minutes above the line at a time, a ten-minute `for` eliminates the flapping completely; a two-minute `for` will not. **The cost is arithmetic, and you should say it out loud.** A `for` of ten minutes means a genuine, sustained breach is announced up to ten minutes late, plus the evaluation interval, plus any grouping wait downstream. If the incident response target for this severity is five minutes, a ten-minute `for` has already blown it. That is the trade the interviewer is listening for: **noise down, detection time up, and the exchange rate is one-for-one.** ## Knob 2: hysteresis on the way back down The pending duration only guards the firing edge. A signal that dips below the threshold for one evaluation will resolve the alert, and the next evaluation re-fires it — a fresh notification, and in some pipelines a fresh page. The instruments here are: - **keeping the alert firing for a defined period after the condition clears.** Prometheus added a `keep_firing_for` field for exactly this in version 2.42; the same idea appears elsewhere under other names. - **two thresholds instead of one** — fire above 500 ms, clear only below 400 ms. Nothing oscillating in the band between them can produce a transition. This is classic hysteresis and it is why thermostats do not chatter. The cost is smaller and mostly cosmetic: the resolved notification arrives later than the recovery did, and dashboards show the alert as active slightly past the real end of the problem. ## Knob 3: the threshold is simply wrong Before reaching for timers, ask an uncomfortable question: **would a human do anything at 501 ms?** If the service routinely runs at 480 ms at peak and nobody has ever acted on a brush past 500, the threshold is set at normal behaviour, and no amount of damping will make a rule about normal behaviour actionable. Move it to where action begins. Deriving the threshold from a service level objective rather than from a round number is the disciplined version of this, and the objective-driven form of the same rule is a subject of its own. ## What not to do A few common wrong answers: - **Silencing the rule.** A silence is a bounded suppression for known work, not a fix for a badly tuned rule; an open-ended silence is how a service goes unmonitored for months. - **Deleting the rule.** Sometimes right, but only after evidence — and a flapping rule is often a correct rule with the wrong timing. - **Suppressing the notification but leaving the rule as is.** You then have a rule that is firing most of the day, which destroys its value as a signal on a dashboard or in an incident review. - **Treating flap damping as a substitute for fixing the system.** If p99 latency really is brushing the line every peak, you have a capacity or performance problem the alert is honestly reporting. ## Summary of the trade Every damping knob buys quiet with latency. Pick the duration from the measured oscillation, keep it comfortably inside the response target the severity implies, add recovery hysteresis so the resolve edge does not chatter, and only then argue about the threshold number.

  • Your response target for this severity is five minutes, but the oscillation lasts eight. What now?
    You cannot damp your way out of it — an eight-minute pending duration breaks the target. Either the threshold is wrong and should move to where action genuinely begins, or the signal is too noisy at this resolution and needs a longer averaging window upstream, or this condition does not deserve a page at that response target and belongs on a slower path.
  • Someone proposes silencing the rule during every traffic peak on a recurring schedule instead of tuning it. What's your objection?
    A recurring suppression window makes the service unmonitored precisely when it is under most stress — the peak is when a real latency incident is most likely. Scheduled suppression is legitimate for known maintenance with a bounded window, not for hiding a rule that fires during normal operation. Fix the timing or the threshold instead.
  • Why is a flapping rule more damaging than its raw notification count suggests?
    It trains responders to dismiss it. After a rule has self-resolved nine times, the tenth firing gets acknowledged reflexively without investigation — and habituation applies to the whole pager, not just that rule. Persistent flapping quietly lowers the response quality of every alert that shares a channel with it.

saying these in an interview costs you the question

  • Just silence it until the peak passes
  • A longer pending duration has no downside
  • Flapping means the metric is broken
  • Delete any rule that has ever self-resolved
  • Set the same for-duration on every rule in the estate

context