Once a Grafana alert is firing, how does it reach a particular destination? Explain the role of notification policies and contact points, and how grouping and timing settings shape what actually arrives.
answer
- labels route, rules do not
- tree: first matching sibling wins, then recurse
- continue = deliberate fan-out to a second destination
- children inherit unset settings from the parent
- group by / group wait / group interval / repeat interval
basics
~20 sFiring alerts enter a tree of notification policies. Each alert descends the tree matching on its labels; the first matching child wins unless that child is set to continue. The matched policy names a contact point and controls grouping labels, initial wait, batching interval and repeat interval.
solid answer
~60 sRouting is label-driven, not rule-driven. Every firing alert carries the labels from its data plus the rule's own labels, and it is evaluated against a tree whose root is the default policy. Child policies have label matchers; an alert takes the first child whose matchers all match, and recursion continues into that child's children. A policy marked to continue matching also lets the alert be considered by later siblings, which is how a duplicate copy goes to an audit or archive destination. The matched policy determines four things: the **contact point** (one or more integrations — chat, email, webhook, paging tool); the **group-by labels**, which decide what counts as one notification (alerts sharing those label values are batched together); **group wait**, the delay before the first message so simultaneous alerts arrive as one; **group interval**, how long before a batch that has changed is sent again; and **repeat interval**, how long before an unchanged still-firing batch is re-sent. Policies also carry mute timings. Unmatched alerts fall through to the default policy, which must therefore always be usable.
code
text · 9 linesdefault (contact: catch-all, group_by [alertname, cluster])
|- match severity=critical, continue=true -> contact: incident-bridge
|- match team=payments -> contact: payments-chat
|- match team=search -> contact: search-chat
alert {alertname=HighLatency, severity=critical, team=payments}
-> matches child 1 (critical), continue -> also tested against siblings
-> matches child 2 (payments)
-> delivered to incident-bridge AND payments-chatgo deeper
Say that firing alerts are routed by their labels through a policy tree to a contact point, which is the actual destination.
Explain first-match-wins among siblings, inheritance of unset settings, the continue option, and what group-by does to the number of messages.
Reason about the four timers as a noise-versus-latency trade-off per route, and debug misrouting by walking sibling order against the alert's real labels.
Treat the label vocabulary as the platform contract between rule authors and routing, decide where simplified per-rule routing is allowed, and make sure no path ends at an unread destination.
## Labels are the interface The single most important idea is that notification routing never mentions rules. An alert arrives at the notification pipeline as a label set plus annotations, and everything downstream — routing, grouping, silencing, muting — matches on those labels. That indirection is what lets a platform team change destinations without editing hundreds of rules, and it is why label vocabulary (team, severity, service, environment, cluster) must be agreed before the alert estate grows. ## The policy tree Notification policies form a tree rooted at a default policy that matches everything. Evaluation is depth-first and first-match-wins among siblings: the alert is tested against each child policy's matchers in order, and the first child whose matchers *all* match takes it; the search then continues among that child's own children, and so on, until no child matches and the current policy handles the alert. Two consequences follow. Ordering matters — a broad matcher placed above a narrow one shadows it, which is the most common routing bug. And an alert that matches nothing lands on the default policy, so the default must be a real destination someone reads, not a placeholder. A policy can be marked to **continue matching** subsequent sibling policies. That is the deliberate fan-out mechanism: send everything to a long-term archive webhook while also routing by team, or copy every critical alert to an incident channel in addition to the owning team's channel. Without it, routing is strictly one destination per alert. Child policies **inherit** unset settings from their parent — grouping, timers and contact point — so a tree usually sets sensible defaults at the root and overrides only what differs. Inheritance is per-field, not all-or-nothing. ## Contact points A contact point is a named bundle of one or more integrations: an email address, a chat webhook, a paging service, a generic webhook, and so on. One contact point with three integrations means one routing decision delivering to three places. Each integration has its own message template, so the same alert can be terse in a page and verbose in a chat message. Contact points also carry the choice of whether resolved notifications are sent — worth turning off for high-volume informational routes and on for anything where humans track closure. ## Grouping and the three timers Grouping is what stops a fleet-wide failure from sending five hundred messages. - **Group by** names the labels that define a batch. Alerts sharing those values are one notification. Grouping by alertname and cluster turns "forty nodes disk-full in cluster A" into one message listing forty instances. Grouping by nothing puts every alert in one giant batch; grouping by an instance-unique label defeats grouping entirely. - **Group wait** is how long to hold a *new* batch before the first send, so alerts that erupt together arrive together. Tens of seconds is typical: long enough to collect the storm, short enough not to delay a page unacceptably. - **Group interval** is the minimum time before sending an updated message for a batch whose membership has changed — more instances joined, or some resolved. - **Repeat interval** is how long an unchanged, still-firing batch waits before being re-sent as a reminder. Set it short and it becomes noise people mute; set it long and a forgotten alert stays forgotten. Hours is the usual range, and it should be longer than the group interval. These are per-policy, so a paging route can be aggressive and an informational route lazy. ## Simplified routing Recent Grafana versions let a rule name a contact point directly, bypassing the tree for simple cases. It is convenient for a small setup and it undermines the indirection the tree provides — the destination now lives on the rule again. Use it for genuinely one-off rules; keep team- and severity-based routing in the tree. ## How to debug a route When an alert goes to the wrong place, do not read the tree top-down and guess. Take the alert's actual label set from the alert list, then walk the sibling order at each level asking which matcher matches first; the shadowing sibling is almost always the answer. If it goes nowhere, check whether a silence or a mute timing swallowed it before delivery, and whether the contact point's integration is actually healthy — a webhook returning errors produces no message and no obvious alert of its own.
- An alert is being delivered to a broad catch-all destination even though a specific policy exists for its team. What is the likely cause?Sibling ordering. Policy matching takes the first sibling whose matchers all match, so a broader policy placed above the team-specific one absorbs the alert and the narrower one is never reached. Confirm by reading the alert's real label set and walking the siblings in order rather than assuming the most specific policy wins. The fix is to reorder so specific policies precede general ones, or to tighten the broad policy's matchers.
- How do you avoid five hundred separate messages when a whole cluster's worth of instances start firing at once?Group on labels that describe the shared cause rather than the individual instance — typically alert name plus cluster or service — so all those instances form one batch, and set a group wait long enough to collect the burst before the first send. The message then lists the members, and the group interval controls how often an updated list is re-sent as instances join or resolve. Grouping on a per-instance label such as host or pod would defeat this entirely and reproduce the storm.
- What is the difference between group interval and repeat interval?Group interval is the minimum delay before sending an updated notification for a batch whose contents changed — new instances joined or some resolved. Repeat interval is the delay before re-sending a batch that has not changed at all, as a reminder that the condition is still firing. Repeat interval should be substantially longer, typically hours, since its job is to prevent an unattended alert from being forgotten rather than to report progress.
saying these in an interview costs you the question
- Thinking the most specific policy wins regardless of order, rather than the first matching sibling.
- Assuming an alert matched by one policy is automatically also considered by later siblings without the continue setting.
- Grouping notifications on a per-instance label, which recreates the storm grouping was meant to prevent.
- Setting the repeat interval short to 'make sure people see it', producing noise that gets muted.
- Leaving the default policy pointed at a destination nobody reads, so unmatched alerts silently vanish.