Grafana lets you write alert rules that Grafana itself evaluates, or rules that are pushed down to be evaluated by the metrics backend. What does each choice change operationally, and how would you decide for a large multi-team estate?
answer
- Grafana-managed: expressions + cross-source, Grafana evaluates
- data-source-managed: pushed to the backend's ruler, scales with data
- Grafana-managed makes Grafana tier-one, needs HA + gossip
- expressions and cross-source are the only real must-haves
- one label vocabulary, rules as code, monitor the monitor
basics
~20 sGrafana-managed rules are evaluated by Grafana, can chain expressions and combine data sources, and depend on Grafana's own availability and state store. Data-source-managed rules live in the backend's ruler, evaluate where the data is, scale with it, and are limited to that backend's query language.
solid answer
~50 sWith **Grafana-managed** rules, Grafana schedules evaluation, runs the query/expression pipeline, keeps instance state in its database, and hands firing alerts to its embedded Alertmanager. That gives cross-data-source rules, server-side expressions, per-rule no-data and error handling, and one UI for everything — at the cost that Grafana is now a critical alerting component that must be highly available, and that every evaluation is a query issued from Grafana over the network. With **data-source-managed** rules, Grafana is an editor: the rule is written into the backend's own rule store and evaluated there, close to the data, scaling with that backend and surviving Grafana being down. You give up expressions, cross-source rules and Grafana's richer state handling, and routing then belongs to that ecosystem's own notification component rather than to Grafana's policies. For an estate, the practical split is to let each metrics platform own the rules over its own data, and reserve Grafana-managed rules for what genuinely needs them: cross-source correlation and sources with no ruler of their own.
code
text · 4 linesrule reads one metrics backend, simple threshold -> data-source-managed (ruler)
rule combines metrics + logs, or two backends -> Grafana-managed (expressions)
rule reads a source with no ruler (SQL, cloud API) -> Grafana-managed (only option)
rule must survive Grafana being down -> data-source-managedgo deeper
Know that Grafana can either evaluate a rule itself or hand it to the metrics backend to evaluate, and that only the first supports combining data sources.
Contrast the capability sets — expressions and cross-source versus scaling and independence — and note where state and notifications live in each case.
Reason about capacity and availability: evaluation duration against the interval, high-availability requirements, and the operational cost of two routing systems.
Set the estate policy — placement rules, one label vocabulary, rules provisioned from version control, unified silencing, and external monitoring of the alerting path itself.
## Two evaluation homes Unified alerting can produce a rule in two shapes, and the difference is *where evaluation happens*. A **Grafana-managed** rule is stored in Grafana, scheduled by Grafana, and evaluated by Grafana issuing queries out to whatever data sources it references. Its instance state lives in Grafana's database and its firing alerts go to the embedded Alertmanager, which owns policies, contact points, silences and mute timings. A **data-source-managed** rule is authored through Grafana but written into the metrics backend's own rule store. From then on the backend evaluates it on its own schedule; Grafana can display it but is not in the path. Notifications leave through that ecosystem's own routing component. ## What Grafana-managed buys - **Expressions.** Reduce, Math, Threshold, Resample run server-side in Grafana, so a rule can express things the data source's query language cannot, and can compare results across sources. - **Cross-data-source rules.** Query A from one backend, query B from another, combine them. There is no other place this can happen. - **Any data source.** Backends with no rule-evaluation capability of their own — many logging, tracing, SQL and cloud-API sources — can only be alerted on this way. - **Uniform semantics.** Per-rule no-data and error handling, one state model, one place to see every rule regardless of the underlying system, and one permission model tied to folders. ## What it costs - **Grafana becomes tier-one.** If Grafana is down or its database is unavailable, those rules are not evaluated. That forces a real high-availability deployment: several instances with shared state, an alerting cluster so the embedded Alertmanager instances gossip and deduplicate notifications rather than each sending its own, and monitoring of the alerting subsystem itself. - **Evaluation load is Grafana's problem.** Every rule is a query issued over the network at its group's interval. Rule count multiplied by frequency multiplied by query cost is a capacity number someone has to own, and it degrades under exactly the conditions that make alerts matter. Evaluation duration approaching the group interval is the leading indicator. - **State is a database.** Instance state and history are persisted; at large rule counts that store and its retention become an operational concern of their own. ## What data-source-managed buys - **Evaluation next to the data**, so no cross-network query per evaluation and far better scaling for high rule counts over one backend. - **Independence from Grafana's availability** for the rules that matter most about that backend. - **The backend's own ecosystem** — its rule language, its recording rules, its existing routing and its existing operational tooling, which a metrics platform team may already run well. And what it costs: no expressions, no cross-source rules, the source's query language only, its state and no-data semantics rather than Grafana's, and split ownership of routing between two systems, which is the part that hurts the humans. ## Deciding for an estate The decision is rarely all-or-nothing; it is a policy about which rules go where. 1. **Default to the data's home.** If a rule reads one metrics backend that has a capable ruler and the platform team runs it, put the rule there. It scales, it survives Grafana, and it lives with the recording rules it depends on. 2. **Use Grafana-managed for what only it can do.** Cross-source correlation, sources without a ruler, and rules that genuinely need server-side expressions. 3. **Do not split routing casually.** Two notification systems mean two places to silence, two sets of contact points and two mental models during an incident. Either centralise routing in one Alertmanager that both paths feed, or accept the split deliberately and document which alerts come from where. 4. **Everything as code.** Whichever home, rules should be provisioned from version control rather than clicked into a UI, so they are reviewable, diffable and restorable. Ad-hoc UI editing at estate scale produces rules nobody can account for. 5. **Enforce a shared label vocabulary across both paths.** Routing, silencing and grouping all key on labels; if the two paths use different names for team or severity, no policy can treat them uniformly. 6. **Monitor the alerting system itself,** from outside it. Evaluation duration, evaluation failures, notification delivery errors and a periodic synthetic alert that must arrive end to end — a monitoring system that fails silently is worse than none, because it is trusted.
- What does running Grafana-managed alerting in high availability actually require?More than one Grafana instance is not enough on its own: the instances must share alert state and their embedded Alertmanagers must form a cluster so they gossip notification state, otherwise every instance evaluates independently and each sends its own copy of every notification. Scheduling must also be coordinated so a rule is not evaluated redundantly. And silences must be visible cluster-wide, or an operator silencing on one node will still be paged by another.
- You have ten thousand rules over a single metrics backend. Why is Grafana-managed evaluation a poor fit?Every rule becomes a query issued from Grafana across the network at its evaluation interval, so Grafana's capacity — not the backend's — becomes the binding constraint, and instance state for ten thousand rules is a substantial database workload. The backend's own ruler evaluates locally, scales with the storage it already runs, and keeps rules alongside the recording rules they depend on. Grafana-managed evaluation earns its cost only where expressions or cross-source queries are genuinely needed.
- What is the main organisational risk of mixing both kinds of rule?Routing and suppression fragment. Data-source-managed rules notify through their own ecosystem's routing while Grafana-managed rules go through Grafana's policy tree, so during an incident there are two places to silence, two sets of contact points and two label conventions. Either funnel both into a single Alertmanager, or make the split explicit and documented, including where each class of alert is silenced and who owns each path.
saying these in an interview costs you the question
- Assuming Grafana-managed rules keep working while Grafana is down.
- Choosing Grafana-managed rules for everything because the UI is nicer, ignoring the evaluation load and the availability requirement it creates.
- Believing multiple Grafana instances give alerting high availability without clustered alert state and Alertmanager gossip.
- Splitting rules across both models without unifying the label vocabulary or the silencing story.
- Treating expressions as an optional nicety rather than the specific capability that justifies Grafana-managed evaluation.