How does Alertmanager's route tree pick a receiver, and what do group_wait, group_interval and repeat_interval each control?
answer
- One tree, matched from the top
- The first match usually ends the search
- Children borrow what they do not set
- Three timers, three different questions
- New group, changed group, unchanged group
basics
~20 sAlertmanager walks its route tree depth-first and the first matching child wins unless it sets continue. Alerts sharing the group_by labels become one notification: group_wait delays the first send, group_interval the next update, repeat_interval the reminder.
solid answer
~50 sEvery alert enters at the root `route`, which names a default receiver and the grouping defaults. Alertmanager descends the tree, testing each child's `matchers` in order; the **first** matching child wins, and evaluation stops there unless that child sets `continue: true`, which lets the alert also fall through to later siblings. A child inherits `receiver`, `group_by` and all three timers from its parent unless it overrides them, so a leaf route can be two lines. Inside a route, alerts whose `group_by` label values are identical form one **aggregation group** that produces one notification. `group_wait` (default 30s) is how long Alertmanager holds a brand-new group so siblings can join the first message; `group_interval` (default 5m) is the minimum gap before sending an update when the group's membership changes; `repeat_interval` (default 4h) is how often an unchanged group is re-sent as a reminder.
code
yaml · 14 linesroute:
receiver: catch-all
group_by: ['alertname', 'service']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers:
- team = "membership"
receiver: membership-slack
routes:
- matchers:
- severity = "page"
receiver: membership-pagerdutygo deeper
Know that alerts reach a receiver by matching labels against a tree of routes, and that grouping is what turns many similar alerts into one notification. Be able to point at where a receiver such as a chat integration is named.
Explain first-match-wins with continue, inheritance of receiver, grouping and timers from parent routes, and the distinct jobs of group_wait, group_interval and repeat_interval with their defaults.
Demonstrate diagnosis: an alert landing on the catch-all because a label was missing, a batching complaint traced to an ancestor route, or a chat channel flooded because grouping was disabled. Know that Prometheus fans out to all Alertmanagers and the cluster deduplicates.
Own the tree as shared infrastructure. Decide which labels every team must emit for routing to follow ownership, whether teams edit their own subtree or request changes, and how the configuration is reviewed and rolled out without a bad merge silencing an estate.
Prometheus decides *that* something is wrong. Alertmanager decides *who hears about it, batched how, and how often*. Both of those are configured in one YAML tree, and the tree's semantics are the part candidates most often get wrong. ## The route tree `alertmanager.yml` has exactly one root `route`. It must name a `receiver`, and it typically also sets `group_by` and the timers. Below it, `routes` is an ordered list of child routes, each with its own `matchers` against the alert's labels. Matching works like this: 1. An alert enters at the root. 2. Alertmanager tests the root's children **in the order written** and descends into the **first** one whose matchers all match. 3. That step repeats down the tree. 4. When the current node has no matching child, the alert is handled by **that node's** receiver — which is why the root's receiver acts as the catch-all. The one twist is `continue`. By default a matching child ends the search among its siblings. Setting `continue: true` on a child means that after that subtree is used, Alertmanager keeps testing the remaining siblings, so one alert can land in several routes and produce several notifications. That is how you send everything to an archival webhook while still paging the owning team. The other thing to internalise is **inheritance**: a child route inherits `receiver`, `group_by`, `group_wait`, `group_interval` and `repeat_interval` from its parent unless it sets its own. A leaf route is often just a matcher and a receiver, and its batching behaviour silently comes from three levels up. When a team complains that their notifications arrive in unexpected clumps, the setting responsible is frequently on an ancestor they never read. ## Grouping Within the route an alert lands in, `group_by` names a list of labels. All alerts whose values for **those** labels are identical belong to one aggregation group and are delivered as **one notification containing many alerts**. - `group_by: ['alertname', 'cluster']` — every instance of the same alert in the same cluster arrives as a single message. A rule that fans out to two hundred instances becomes one notification. - `group_by: ['...']` — the literal three-dot value disables aggregation entirely: every alert is its own group and gets its own notification. Useful when a downstream system wants raw events, and a reliable way to flood a chat channel otherwise. Grouping is what stops one root cause from sending fifty messages, and it is applied per route, so the routing decision and the batching decision are made together. ## The three timers This is the table worth memorising, because the names are close and the jobs are not: | Timer | Default | Starts when | What it buys you | |---|---|---|---| | `group_wait` | 30s | The first alert of a **new** group arrives | Late-arriving siblings join the first message instead of triggering a second one | | `group_interval` | 5m | The previous notification for that group was sent | A floor on how often a **changed** group re-notifies, so a growing incident does not spam | | `repeat_interval` | 4h | The previous notification for that group was sent | A reminder cadence for a group that is still firing and **unchanged** | Read them as three different questions. *How long do I hold a new group before speaking?* — `group_wait`. *A new alert just joined an existing group; how soon may I speak again?* — `group_interval`. *Nothing has changed for hours; when do I nag?* — `repeat_interval`. A worked case from a climbing-gym membership platform, where one team owns roughly seventy per cent of the alert volume: a deploy breaks their check-in service and 37 pods fire the same alert within 20 seconds. With `group_by: ['alertname', 'service']` and a 30-second `group_wait`, all 37 arrive in a single notification. Nine more pods fail two minutes later; the group changed, but `group_interval` holds the update until five minutes after the first message. If nobody fixes it, `repeat_interval` re-sends the same content four hours later. ## Receivers, and what actually leaves the process A `receiver` is a named bundle of integration configs — `pagerduty_configs`, `slack_configs`, `webhook_configs`, `email_configs` and others — and one receiver may hold several, in which case every one of them gets the notification. Receivers are referenced by name from routes, so the tree stays about *matching* and the receiver list stays about *delivery*. Two operational points that follow from all of this: - Prometheus sends its firing alerts to **every** Alertmanager it is configured with, deliberately. The Alertmanager instances gossip their notification history to each other so the duplicate does not become a duplicate page. Deduplication is Alertmanager's job, not yours. - Because routing matches on labels, routing correctness depends entirely on the labels your rules attach. A missing `team` label does not produce an error; it produces an alert that quietly lands on the root catch-all receiver.
- An alert must both page the owning team and reach a central audit webhook. How do you configure that?Put the audit route above the team routes and set `continue: true` on it. Matching a child normally ends the search among siblings, but `continue` lets Alertmanager keep testing later siblings after using that subtree, so the alert lands in both routes and produces both notifications. Without it, the first matching route is the only one that fires.
- A leaf route sets only matchers and a receiver, yet its notifications arrive in large batches. Why?Child routes inherit `group_by` and all three timers from their parent, and inheritance walks the whole way to the root. A leaf that sets none of them is batching exactly as some ancestor specified. Read the tree upward from the leaf before changing anything; the setting responsible is often several levels above the route the team edited.
- What does `group_by: ['...']` do, and when is it the right choice?The literal three-dot value disables aggregation: every alert becomes its own group and gets its own notification. It is right when a downstream system wants one event per alert and will do its own correlation — a ticketing integration, for example. Pointed at a human channel it removes the single mechanism that keeps one incident from producing fifty messages.
saying these in an interview costs you the question
- Thinks every matching route in the tree fires by default
- Confuses group_wait with the rule's for duration
- Says repeat_interval controls how often a changed group notifies
- Believes a child route needs its own timers to work
- Thinks grouping happens in Prometheus rather than Alertmanager
- Cannot say what happens when no child route matches