Grafana's unified alerting can be delivered from files in the provisioning directory alongside dashboards and data sources. What rules does that delivery mechanism impose — object identity, what is replaced versus merged, and what happens to those objects in the UI?
answer
- provisioning/alerting: groups, contactPoints, policies, muteTimings, templates
- Rule identity = uid; changed uid = delete + recreate, state and silences lost
- Policy tree has one root → provisioning replaces the whole tree
- Provisioned = provenance marker = read-only UI unless provenance disabled
- Export API/UI generates valid YAML from a UI-built rule
basics
~20 sAlerting provisioning files declare rule groups, contact points, notification policies, mute timings and templates. Every rule needs a stable uid or reloads recreate it and lose its state. Provisioned objects are read-only in the UI. The notification policy tree is a single root object, so provisioning it replaces the whole tree rather than merging.
solid answer
~1 minUnder `provisioning/alerting/*.yaml` you can declare `groups` (folder + evaluation group + rules), `contactPoints`, `policies` (the notification policy tree), `muteTimes` and `templates`, each with matching `delete…` sections for removal. Three mechanism rules dominate: 1. **Identity is the `uid`.** Alert rules require a stable uid; if it changes, the reload deletes and recreates the rule, which resets its evaluation state and orphans existing silences and annotations. Contact points likewise. 2. **The notification policy tree is one object.** It has a single root, so provisioning `policies` **replaces the entire tree** — two teams cannot each own a file that contributes a branch. This is the answer that most surprises people and the reason routing is usually owned centrally. 3. **Provisioned objects carry a provenance marker and are read-only in the UI** — the same principle as provisioned dashboards. Editing is possible only by changing the files, or by explicitly disabling provenance (a file flag / an API header) which hands ownership back to the UI. The practical authoring loop is to build a rule in the UI, then use the **provisioning export** (API or the UI's export action) to emit the YAML, review it and commit it. Terraform exposes the same objects as resources.
code
text · 15 linesapiVersion: 1
groups:
- orgId: 1
folder: Platform
name: api-latency
interval: 1m
rules:
- uid: api-latency-p99 # STABLE. changing it = delete + recreate
title: API p99 latency high
...
policies:
- orgId: 1
receiver: default # this is the WHOLE tree; applying replaces it
routes: [ ... ]go deeper
Know that alert rules, contact points and routing can come from provisioning files and that such objects are read-only in the UI.
Explain uid-based identity and the read-only provenance behaviour, and know the export path from a UI-built rule to YAML.
Own the edges: whole-tree replacement for policies, recreate-on-uid-change consequences for state and silences, explicit delete sections, and secret handling in contact points.
Set ownership policy — who owns routing versus rules, how provenance is granted or revoked, and what review and blast-radius controls guard the high-impact policy file.
## What can be delivered as files Grafana's provisioning directory has an `alerting/` subdirectory whose files declare, at `apiVersion: 1`: - **`groups`** — evaluation groups, each bound to a folder and an interval, containing alert rules. Each rule carries a `uid`, a title, a condition, its query/expression stages, `for` duration, labels and annotations, and its no-data/error handling. - **`contactPoints`** — named receivers with their integration settings; secrets inside them follow the same secure-field and interpolation rules as data source secrets. - **`policies`** — the notification policy tree. - **`muteTimings`** — named time intervals. - **`templates`** — notification message templates. Each has a corresponding deletion section (for example `deleteRules`, listing uid + orgId) so that removal is expressible as code rather than a manual UI action. Note the asymmetry with dashboards: for dashboards, removing the file removes the object; for alerting, you generally state the deletion explicitly. ## Identity: uid discipline An alert rule's `uid` is its identity. Provisioning matches on it to decide update-versus-create. If a rule's uid is regenerated — most often because someone rebuilt the rule in the UI and re-exported it, or because a generator derives uids from a name that changed — the result is not an update: the old rule is removed and a new rule appears. Everything attached to the old identity goes with it, including its current evaluation state (so a firing alert resolves and re-fires, paging people again), existing silences that matched by rule uid, and the historical annotations tied to it. Generating uids deterministically from a stable key, and treating a uid change as a breaking change in review, is the discipline that avoids this. ## Replace versus merge This is the sharpest edge in alerting-as-code. Rules and contact points are addressable individually, so multiple files can each own their own set. **The notification policy tree cannot be.** It is a single hierarchical object with one root, so a provisioning file that declares `policies` supplies the *whole* tree; applying it replaces what was there. Practical consequences: - You cannot let each team own a file that adds a branch for its own labels. Either one central file owns routing and teams contribute changes to it through review, or routing is expressed at the rule/label level and the tree stays deliberately generic. - A partial file applied by accident can wipe routing for everyone at once. This makes the policy file a high-blast-radius artefact that deserves stricter review and, ideally, a test that renders the tree and asserts key routes still match. ## Provenance and the read-only UI Provisioned alerting objects are tagged with a provenance value. As with dashboards, this makes them **read-only in the UI**: editing is refused so the files stay authoritative and a UI change cannot be silently overwritten at the next apply. When you genuinely want the UI to own an object — handing a rule over to a team, or migrating away from files — you disable provenance, either through a flag in the file or by calling the provisioning API with the header that suppresses the provenance marker. Do that deliberately; a mixed estate where nobody knows which rules are code-owned is worse than either pure model. ## The authoring loop Hand-writing alert rule YAML is unpleasant: the query and expression stages are verbose and easy to get subtly wrong. The workflow that actually works is to build and validate the rule in the UI, then use the **provisioning export** — an API endpoint that returns YAML (or JSON), also surfaced as an export action in the UI — to emit exactly what Grafana would accept back. Review that output, set the uid deliberately, and commit it. The same export path exists for contact points and policies, which is the fastest way to bootstrap a routing file that is already valid. ## Alternative delivery The Terraform provider exposes the same objects as resources, which some teams prefer because deletions and drift are handled by state rather than by explicit delete sections, and because it works against a hosted Grafana where you cannot mount files. The trade is state ownership and a destroy blast radius. The Grafana Operator covers similar ground in Kubernetes with reconciled custom resources. The choice does not change the semantics above: uid identity, tree replacement and provenance behave the same whichever delivery path writes them. ## Scope note All of the above is about *delivery*: how alerting objects get into Grafana and what that mechanism guarantees. How to design a rule condition, choose the `for` duration, handle no-data, or structure routing for humans is alerting design, a separate topic.
- Two teams each want to provision their own notification routing. How do you organise that?You cannot split the tree across files, because the policy object has a single root and provisioning it replaces the whole thing. The workable patterns are: keep one centrally owned routing file that teams change by pull request, with review and a rendering test asserting key routes still match; or keep the tree generic — route by a team label to a per-team contact point — so teams only own their contact points and their rules' labels, which are individually addressable. Splitting the tree by file is the design that quietly deletes other teams' routing.
- What actually goes wrong if a generator produces a new uid for an existing alert rule?Provisioning treats it as a different object: the old rule is deleted and a new one is created. The new rule starts with no evaluation state, so a condition that was already firing transitions from nothing to firing and pages again; existing silences that matched the old rule uid no longer apply, so suppressed alerts start notifying; and the rule's history and annotations no longer line up. Deterministic uids derived from a stable key, and treating uid changes as breaking in review, prevent it.
saying these in an interview costs you the question
- Assuming each team can provision a fragment of the notification policy tree and have them merge
- Letting uids be regenerated, then being surprised by re-firing alerts and dead silences
- Hand-writing rule query/expression YAML instead of using the export path from a validated rule
- Expecting provisioned alert rules to be editable in the UI, or disabling provenance ad hoc so ownership becomes unclear
- Assuming deleting the file removes the alert rule the way it removes a dashboard, instead of using the explicit delete sections