What does a Prometheus alerting rule contain, and what does its `for` duration change about when it fires?
answer
- A rule is more than a threshold
- It can be true and still not notify
- One field delays the transition
- inactive, pending, firing
basics
~20 sA Prometheus alerting rule pairs a PromQL expression with an optional for duration, plus labels and annotations. The for duration requires the expression to hold continuously across that window, so the alert sits in pending before it becomes firing.
solid answer
~40 sA rule inside a Prometheus rule file has five parts: `alert` (the name, which becomes the `alertname` label), `expr` (the PromQL expression), an optional `for` duration, `labels` and `annotations`. Every sample the expression returns becomes one alert instance carrying that sample's labels. Without `for`, the instance goes straight to **firing** on the first true evaluation. With `for: 10m`, it enters **pending** and must still be returned at *every* evaluation for ten minutes; one evaluation where the series disappears sends it back to **inactive** and restarts the clock. `labels` are part of the alert's identity and are what Alertmanager groups and routes on; `annotations` are free text for the human and change nothing about routing. Prometheus itself never notifies anyone — it posts firing alerts to the Alertmanagers listed under its `alerting` block.
code
yaml · 12 linesgroups:
- name: membership-api
rules:
- alert: CheckinQueueBacklog
expr: gym_checkin_queue_depth{service="checkin"} > 240
for: 7m
labels:
severity: page
team: membership
annotations:
summary: "Check-in queue depth is {{ $value }} on {{ $labels.instance }}"
runbook_url: "https://runbooks.internal/checkin-queue"go deeper
Be ready to name the fields of an alerting rule from memory — alert, expr, for, labels, annotations — and to say in one sentence that for requires the condition to persist before the alert fires.
Explain the inactive to pending to firing transition precisely, including that a single failed evaluation resets the clock, and that labels form the alert's identity while annotations are free text for the responder.
Show you have debugged this: an expression fanning out to hundreds of instances, a templated value in a label destroying deduplication, or pending state surviving a restart. Know why for filters intermittent truth but not a steadily-breached threshold.
Own the convention rather than the rule. Decide which labels every team's rules must carry so routing and ownership work estate-wide, and where the boundary sits between damping noise in the rule and damping it in the notification layer.
A Prometheus alerting rule is the object that turns a query result into something a human eventually hears about. It is worth reading its five fields carefully, because two of them look interchangeable and are not, and one of them is the single most misunderstood setting in the whole product. ## What the rule file contains Alerting rules live in YAML files that `prometheus.yml` names in its `rule_files` list. Rules are always nested inside a named group, and each alerting rule may carry: - **`alert`** — the alert's name. It becomes the `alertname` label on every alert instance the rule produces, and `alertname` is what Alertmanager groups on by default. - **`expr`** — a PromQL expression, evaluated on the group's schedule. **Every sample in the result becomes one separate alert instance**, carrying that sample's labels. A single rule written against a 2.3-million-series membership-platform metric store can therefore produce hundreds of alert instances at once, one per pod, per region or per gym site. - **`for`** — an optional duration. Covered below. - **`labels`** — extra key/value pairs stapled onto every instance, such as `severity: page` or `team: membership`. - **`annotations`** — free text such as `summary`, `description` and a runbook link. Both `labels` and `annotations` are Go templates evaluated per instance, so `{{ $labels.instance }}` and `{{ $value }}` expand to that instance's own label values and its sample value. ## Labels versus annotations: identity versus payload Alertmanager identifies an alert by its **complete label set**. Two evaluations producing the same labels are the same alert being refreshed; change one label value and you have created a different alert, while the old one resolves. | Field | Part of the alert's identity | Used for routing, grouping, silencing | Typical content | |---|---|---|---| | `labels` | yes | yes | `severity`, `team`, `cluster` | | `annotations` | no | no | `summary`, `description`, `runbook_url` | The practical consequence: never template a measured value or a timestamp into a **label**. `value: "{{ $value }}"` as a label rewrites the alert's identity on every evaluation, so it never deduplicates, never groups, and no silence written against it keeps matching. Values go in annotations, where they can change freely. ## The three states, and what `for` really does Each alert instance is in one of three states: 1. **inactive** — the expression returns no sample for that label set. 2. **pending** — the label set has appeared, but the `for` duration has not yet elapsed. 3. **firing** — the label set has been present at every evaluation for the whole `for` window; only now is it handed to Alertmanager. `for` is *not* "wait, then check once". Prometheus re-evaluates the rule on the group's interval throughout the window, and the condition must hold at **each** of those evaluations. A single evaluation where the series is absent drops the instance back to inactive and resets the clock to zero. That is precisely why `for` exists: an expression that is intermittently true — a latency threshold crossed for one scrape during a traffic peak — never accumulates enough consecutive evaluations to fire. A rule with no `for` (equivalently `for: 0s`) goes from inactive to firing on the first true evaluation. That is legal and sometimes correct, for conditions that are already durable by construction. Two details worth knowing: - The states are queryable. Prometheus exposes a synthetic `ALERTS` series with an `alertstate` label of `pending` or `firing` and a value of `1`, so you can graph how much time your alerts spend pending and whether a `for` window is doing any work at all. - A restart does not silently throw pending progress away: Prometheus persists pending state as an `ALERTS_FOR_STATE` series and restores it on startup, within a bounded tolerance, so a routine restart mid-window does not reset every clock. There is also `keep_firing_for`, the mirror image of `for`: it holds an alert in the firing state for an extra period after its expression stops matching, which damps a condition that resolves and immediately re-fires. ## What happens after it fires Prometheus sends nothing to a pager. It POSTs firing alerts over HTTP to **every** Alertmanager listed under its `alerting` block, and keeps re-sending them on a regular cadence for as long as they remain firing — the receiving side is expected to be idempotent about that. When the expression stops returning the label set, Prometheus stops sending the alert and marks it resolved. Everything a human experiences after that point — batching, suppression, delivery to a chat channel or a pager — belongs to Alertmanager, not to the rule. ## Common mistakes - Setting a `for` shorter than the group's evaluation interval, which means no meaningful delay at all. - Assuming `for` smooths a *noisy* signal. It filters *intermittent* truth; a signal that is continuously slightly over the threshold fires exactly on schedule. - Forgetting that an expression returning many samples produces many alerts, then being surprised by the notification volume. - Writing routing decisions into annotations, where Alertmanager will never see them.
- Your rule's expression returns 60 samples every evaluation. How many alerts does Alertmanager receive?Sixty separate alert instances, one per sample, each carrying that sample's labels plus the rule's `labels` block. The rule is one object; the alerts are per-series. This is why a broad expression over a large store can produce a notification storm, and why Alertmanager's grouping exists at all — the fan-out happens at rule-evaluation time and cannot be undone later.
- Someone templates `{{ $value }}` into a rule's `labels` block. What breaks?The measured value becomes part of the alert's identity, so every evaluation with a slightly different number looks like a brand-new alert and the previous one resolves. Deduplication stops working, grouping fragments, and any silence written against that label set stops matching almost immediately. Put the value in `annotations`, which are payload rather than identity.
- What does `keep_firing_for` do that `for` does not?`for` delays the transition into firing; `keep_firing_for` delays the transition out of it, holding the alert firing for an extra period after the expression stops matching. It is the tool for a condition that resolves and immediately re-fires, where the resolve/re-fire churn itself is the noise you want to suppress rather than the initial trigger.
saying these in an interview costs you the question
- Says an alert fires the instant its expression is true
- Thinks `for` checks the expression only once at the end
- Cannot distinguish labels from annotations
- Believes annotations affect routing or deduplication
- Assumes one rule always produces exactly one alert
- Thinks Prometheus itself sends the Slack message or page