skip to content

In Grafana's unified alerting, a Grafana-managed alert rule is defined as one or more data-source queries plus expressions rather than as a single threshold on a graph. Walk through how such a rule is evaluated and how it turns into individual firing alerts.

level: middleimportance: must knowfreq 50%

answer

  1. queries A/B -> expressions -> one condition node
  2. Reduce collapses a range to one number
  3. Math for cross-source ratios and comparisons
  4. one label set = one alert instance
  5. rule labels merged on top; annotations templated

basics

~20 s

Each query runs and returns labelled series. Server-side expressions then chain on those results — reduce a series to one number, do math, apply a threshold — and one expression is marked the rule's condition. Every distinct label set that satisfies it becomes its own alert instance.

solid answer

~60 s

A Grafana-managed rule is a small pipeline of referenced nodes. Query nodes (A, B, …) each run against a data source over a relative time range and return one or more time series, each identified by its label set. Expression nodes then operate on those references server-side: **Reduce** collapses each series to a single value with a function such as last, mean or max; **Math** does arithmetic across references, for example `$A / $B > 0.05`; **Threshold** compares a single-value result against a bound; **Resample** aligns series before combining them. One node is designated the condition. At each evaluation the pipeline runs, and the condition's result is read per label set: a non-zero value means that instance is breaching. Because the labels come from the data, one rule produces N alert instances — one per host, per route, per queue — each with its own state, and those labels are what notification policies later route on. Instance labels are then merged with the rule's own labels, and annotations are templated per instance.

code

text · 8 lines
text
A: query  errors  by route     -> series per route
B: query  requests by route    -> series per route
C: reduce(A, last)             -> one number per route
D: reduce(B, last)             -> one number per route
E: math   $C / $D > 0.05       -> 1 or 0 per route   [CONDITION]

routes breaching E each become a separate alert instance,
labelled with whatever labels the data carried

go deeper

for a junior

Describe the shape: a query, a reduction to a single number, a threshold, and one condition that decides whether the rule is breaching.

for a middle

Explain expression chaining by reference, cross-data-source Math, and that one label set equals one alert instance with its own state.

for a senior

Bring in the failure shapes — reduction function choice, empty denominators, instance churn, label cardinality — and how rule labels feed routing.

for a principal

Frame the pipeline as a contract between rule authors and routing: labels are the interface, so their vocabulary must be standardised across teams before the alert estate grows.

## Why a rule is a pipeline, not a threshold A panel threshold is a display decision. An alert rule must produce a definite yes/no per entity, at a fixed cadence, independently of anyone looking at a dashboard. Unified alerting expresses that as a directed graph of nodes: data-source queries at the bottom, server-side expressions above them, and exactly one node nominated as the condition whose output decides state. This structure buys two things a single query cannot. First, it works across data sources: query A can come from one backend and query B from another, and a Math expression can combine them because the combination happens in Grafana rather than in a query language. Second, it forces the reduction to a single value to be explicit, which is where most alerting mistakes live. ## The node types - **Query.** Runs against a data source over a *relative* time range ("now minus 5 minutes to now"), producing labelled series. The range matters: it is the window the rule sees at every evaluation, and it should comfortably exceed the data's arrival interval so a single late scrape does not empty the frame. - **Reduce.** Turns each series into one number using a function (last, min, max, mean, sum, count), with an explicit policy for non-numeric values — drop them, or treat them as zero. Many data sources return a range, and a threshold cannot be applied to a range, so Reduce is usually mandatory rather than optional. - **Math.** Free-form arithmetic and comparison over other nodes' results by reference. This is how ratios are built (`$errors / $total`) and how unit conversions or combined conditions are expressed. A comparison in Math yields 1 or 0. - **Threshold.** A simpler form of the same idea: is above, is below, is within range, is outside range — and in newer versions optionally with hysteresis, so recovery uses a different bound than firing to stop flapping around a single number. - **Resample.** Aligns series onto a common time grid so two sources with different intervals can be combined. ## From condition to alert instances The crucial idea is **multi-dimensional alerting**. The condition is not evaluated once; it is evaluated once per label set present in the data. If query A returns latency for forty routes, the rule produces up to forty alert instances, each with its own state machine, its own pending timer and its own notification. This is what makes one rule maintainable where forty copies would not be. The consequence is that labels are load-bearing in three ways: they *identify* an instance (the same label set across evaluations is the same alert, and a changed label set is a new alert plus a resolved old one), they are what notification policies match on for routing, and they determine cardinality. A label that includes something unique per evaluation — a request id, a timestamp, a raw error message — produces a stream of one-shot alerts that fire once and resolve, which is the classic self-inflicted alert storm. Rule-level labels are merged on top of the data's labels, which is how you attach routing dimensions the data does not carry, such as team or severity. Annotations are the human-facing text (summary, description, runbook link) and are templated per instance, so they can interpolate the instance's own labels and values. ## Practical failure shapes - **Forgetting the reduction.** Applying a threshold directly to a range either errors or silently evaluates on something you did not intend. Reduce first, deliberately, and choose the function on purpose: `last` alerts on the newest sample and is sensitive to a single spike; `mean` smooths and delays; `max` is for "did it ever breach". - **A ratio with an empty denominator.** When traffic stops, an errors-over-total expression can produce no data or a non-numeric result instead of a comfortable zero, which drives the rule into its no-data path rather than staying normal. - **Instance churn.** If the underlying query stops returning a series entirely (a host disappears), that instance's data vanishes and the rule handles it through its no-data behaviour rather than by firing on the old value. - **Alerting on the visualisation.** A dashboard panel's transformations and its display-time smoothing are not part of the rule; the rule re-runs its own queries. A rule that "matches the graph" only by coincidence will diverge. ## Where this fits The pipeline decides *whether* something is breaching right now. It says nothing about how long a breach must persist before it counts, how repeated notifications are grouped, or who is told — those are the pending period, the notification policy and the contact point respectively.

  • Why does a single rule sometimes produce dozens of separate alerts, and how do you control that?
    Because evaluation is per label set: every distinct combination of labels returned by the query is its own alert instance with its own state and its own notification. You control it at the query, by aggregating away labels you do not want to page on, and at routing, by grouping notifications on the labels that matter so many instances arrive as one message. What you must avoid is unbounded labels such as identifiers or free-text messages, which create a new instance on almost every evaluation.
  • You want to alert when errors exceed five percent of requests. Why is a Reduce step usually required before the comparison?
    Both queries return a time range, and a threshold or ratio comparison needs a single number per series to produce a definite yes/no. Reduce collapses each series with an explicit function so the choice is visible: `last` reacts to the newest sample, `mean` smooths over the window, `max` catches any breach within it. Skipping the reduction either errors outright or leaves the evaluation depending on implicit behaviour you did not choose.
  • How do labels and annotations differ in an alert rule?
    Labels are identity and routing: they define which alert instance this is, they are matched by notification policies and silences, and changing them creates a different alert. Annotations are human-facing text such as a summary, description or runbook link, are templated per instance, and are never matched on. The practical rule is to keep labels few and low-cardinality, and put anything descriptive or high-cardinality into annotations.

saying these in an interview costs you the question

  • Describing an alert rule as a threshold drawn on a panel, ignoring that the rule runs its own queries independently of the dashboard.
  • Applying a threshold to a time range without an explicit reduction and not knowing which sample it used.
  • Putting high-cardinality values such as request ids or error text into labels, creating one-shot alert storms.
  • Thinking one rule can only produce one alert.
  • Confusing annotations with labels and trying to route notifications on an annotation.

context