Explain the states a Grafana-managed alert instance moves through — including its pending period — and what happens when the underlying query returns no data or fails outright.
answer
- Normal -> Pending -> Alerting -> Resolved
- pending period must be a multiple of the interval
- one false evaluation resets pending to zero
- NoData: NoData / Alerting / Normal / KeepLast
- Error: Alerting / Normal / KeepLast
basics
~20 sAn instance is Normal until the condition breaches, then Pending until it has breached continuously for the pending period, then Alerting, which notifies. Missing data and query failures are separate states whose handling is configurable per rule: treat as alerting, as normal, as a dedicated no-data alert, or keep the last state.
solid answer
~60 sEvaluation happens on the group's interval. If the condition is not breaching, the instance is **Normal**. On the first breaching evaluation it becomes **Pending**, and it must keep breaching for the whole pending period before it becomes **Alerting** and a notification is produced; any non-breaching evaluation in between sends it straight back to Normal, which is the flap filter. When the breach clears, the instance resolves and a resolved notification follows. Two non-boolean outcomes are handled separately. **No data** — the query returned an empty frame — is configurable: raise a dedicated no-data alert, treat it as alerting, treat it as normal, or keep the previous state. **Error** — the data source timed out or failed — is likewise configurable: alerting, normal, or keep the last state. Both special states also respect the pending period in current versions. The right choice is a judgement about the failure mode: for a rule that watches a thing which should always report, silence-by-absence is the dangerous outcome, so no-data should be visible.
code
text · 6 linest0 breach -> Pending (0/5)
t1 breach -> Pending (1/5)
t2 ok -> Normal (timer reset)
t3 breach -> Pending (0/5)
t4..t8 breach-> Alerting at t8, notification sent
t9 no data -> whichever NoData handling the rule configuresgo deeper
Name the path Normal to Pending to Alerting and say the pending period stops brief spikes from paging.
Explain the reset-on-any-false rule, the interval/pending-period arithmetic, and that no-data and error are separate configurable outcomes.
Choose the no-data and error handling per rule from the failure mode, and use hysteresis or keep-firing rather than inflating the pending period to fight flapping.
Set defaults for the estate — what no-data means by rule class, how monitoring-system failure is surfaced without paging on its own overload, and how state history is used in review.
## The normal path Each alert instance — remember, one per label set, not one per rule — carries its own state, and the rule's evaluation group drives it forward on a fixed interval. 1. **Normal.** The condition evaluated false at the last evaluation. 2. **Pending.** The condition evaluated true, but it has not yet been true for long enough. Nothing is notified. The instance is visible in the UI, which is useful when debugging a rule that never seems to fire. 3. **Alerting (Firing).** The condition has been continuously true for at least the pending period. This is the point at which the instance is handed to the notification pipeline. 4. **Resolved / back to Normal.** The condition evaluated false; the instance leaves the firing set and, if the contact point is configured to send them, a resolved notification is delivered. The pending period (historically written as a "for" duration) exists to convert a momentary spike into a non-event. Its interaction with the evaluation interval is the part people get wrong: the condition must be true at *every* evaluation during the window, and a single false evaluation resets the timer to zero. A pending period shorter than the evaluation interval therefore has no effect at all, and one that is not a multiple of the interval effectively rounds up to the next evaluation. If your data arrives every 60 seconds and you evaluate every 60 seconds, a 90-second pending period behaves like 120. ## No data An empty result is not a false condition; it is an absence of evidence, and treating the two as identical is how monitoring goes quiet exactly when the monitored thing dies. Grafana therefore makes it an explicit per-rule choice: - **No Data** (dedicated) — produces a distinct alert for the rule so a human sees "this rule stopped receiving data" rather than "everything is fine". - **Alerting** — treat absence as a breach. Appropriate when the series is guaranteed to exist while the system is healthy. - **Normal** — treat absence as fine. Appropriate for sparse, event-driven series where gaps are expected and would otherwise page constantly. - **Keep Last State** — hold whatever the instance was before. Useful across a known-flaky exporter, dangerous as a default because it can pin an instance firing or silent indefinitely. A subtlety: a no-data result usually has no labels to attach to, so the resulting alert is at rule level rather than per instance. And an instance that simply *disappears* from the query result (a host that no longer exists) is a different situation from the whole query returning nothing — Grafana stops tracking series that vanish, which is why "the alert stopped without resolving" sometimes has no notification behind it. ## Error A data-source timeout, an authentication failure or a malformed query produces an error rather than a value. The configurable outcomes mirror the no-data case: alerting (loud, correct when a silent monitoring system is worse than a false page), normal (quiet, acceptable only when someone else watches the data source), or keep last state. Errors are also where evaluation cost shows up: a query too heavy for the interval will time out under load precisely when the system is stressed, so a rule that is loud on error will page for the monitoring system's own overload. That is an argument for making the query cheap, not for making the rule quiet. ## Recovering and hysteresis A breach that oscillates around the threshold produces alternating firing and resolved notifications. Two mechanisms address it. Threshold expressions can carry **hysteresis** — a separate, lower recovery bound so an instance must clearly improve before resolving. And newer Grafana versions add a **Recovering** state with a keep-firing duration, so an instance that stops breaching remains firing for a configured period before resolving, absorbing brief dips. Both are preferable to simply raising the pending period, which delays the first real alert as well. ## Where the state ends up Instance state is persisted so it survives a restart rather than everything resetting to Normal and re-firing, and the state history is queryable, which is the fastest way to answer "was this flapping or genuinely down for an hour?". A paused rule is different from all of the above: it stops evaluating entirely and holds no state, which makes pausing a blunt tool compared with a silence.
- A rule has an evaluation interval of one minute and a pending period of thirty seconds. What is the practical effect?None beyond a single evaluation. The pending period is measured in evaluations, so a value shorter than the interval is satisfied by the very next evaluation and the instance fires immediately on the second breach — effectively no flap protection. Pending periods should be whole multiples of the evaluation interval, and any non-multiple rounds up to the next evaluation boundary.
- When should a rule treat No Data as Normal rather than as Alerting?When gaps are an expected property of the data rather than a symptom — sparse, event-driven series such as errors on a low-traffic endpoint, or batch metrics that only appear when a job runs. Treating those as Alerting produces constant noise that trains people to ignore the rule. Conversely, for a series that must exist whenever the system is healthy, absence is the most important signal there is and must never be silently normal. The dangerous default is Keep Last State, which can pin an instance firing or silent indefinitely with no one noticing.
- An instance alternates between firing and resolved every few minutes. What are your options besides raising the pending period?Add hysteresis to the threshold so the recovery bound is meaningfully lower than the firing bound, which stops oscillation around a single number, or use the keep-firing/Recovering behaviour so an instance stays firing through brief dips before resolving. You can also smooth the input by choosing a mean reduction over a longer window instead of the last sample. Raising the pending period works but pays for it by delaying every genuine alert, including the ones you care about.
saying these in an interview costs you the question
- Treating an empty query result as equivalent to a healthy condition without deciding so deliberately.
- Setting a pending period shorter than, or not a multiple of, the evaluation interval and expecting flap protection.
- Believing the pending timer accumulates across non-breaching evaluations instead of resetting.
- Using Keep Last State as a general default to reduce noise.
- Confusing pausing a rule with silencing it — pausing stops evaluation entirely and loses state.