Grafana-managed alert rules are organised into evaluation groups inside a folder. What does a group actually control at evaluation time, and how should you use grouping when a folder holds many rules?
answer
- group = interval + sequential order, inside a folder
- groups run in parallel, rules within a group in series
- cheap/fast vs expensive/slow in different groups
- group != notification grouping
- pending period counts evaluations, so interval changes rescale it
basics
~20 sA group sets the evaluation interval shared by its rules and evaluates them in order, one after another, at each tick. It controls scheduling and ordering only — it has nothing to do with how notifications are grouped or routed.
solid answer
~50 sAn evaluation group is a named bucket inside a folder with a single interval. Every rule in it is evaluated at that interval, sequentially in the group's defined order, so a slow rule delays its successors within the same group; independent groups run concurrently. That gives you two levers. **Interval** is chosen per group, so cheap fast-moving checks can sit in a ten- or thirty-second group while expensive aggregate queries sit in a five-minute group, instead of everything sharing one compromise. **Concurrency and blast radius** are shaped by splitting rules across groups: a heavy query in its own group cannot starve unrelated rules, and spreading rules over several groups avoids a thundering herd of simultaneous queries against one data source. What a group does *not* do is bundle notifications — that is the notification policy's grouping. Also note the pending period is measured in evaluations of the group's interval, so changing a group's interval silently changes every member rule's effective flap protection.
code
text · 6 linesfolder: payments (team permission boundary)
group "30s-critical" interval 30s -> rule A, rule B (cheap instant queries)
group "5m-aggregate" interval 5m -> rule C (heavy range query)
A and B evaluate in series every 30s; C runs independently every 5m.
Rule C slowing down cannot delay A or B.go deeper
Know that a rule lives in a folder and a group, and that the group sets how often it is evaluated.
Add that rules in a group run in order at that interval while separate groups run in parallel, and that grouping is unrelated to notification bundling.
Design the grouping scheme around cadence and query cost, watch evaluation duration against the interval, and treat interval changes as behaviour changes for pending periods.
Own the estate-level layout: folders as the permission and ownership boundary, groups as the cadence and isolation boundary, plus a load model for the shared data sources every rule queries.
## What a group is In unified alerting, a Grafana-managed rule lives in a folder, and within that folder it belongs to a named evaluation group. The group is the scheduling unit. It carries one property that matters above all others — the evaluation interval — and it defines an ordering over its member rules. At each tick of the interval, the scheduler walks the group's rules in order and evaluates them one after another. Different groups are independent and run in parallel. So the group is simultaneously a *rate* (how often) and a *serialisation boundary* (what waits for what). ## The two levers grouping gives you **Rate matching.** Evaluation frequency should track the data and the urgency, not a global default. A rule over a metric scraped every fifteen seconds, protecting a user-facing path, justifies a short interval. A rule that runs an expensive aggregation over a day of logs does not — evaluating it every thirty seconds burns query capacity for an answer that cannot change that fast. Putting them in the same group forces one interval on both, so the cheap fast one gets slow or the expensive slow one gets hammered. Separate groups let each be right. **Isolation and load shaping.** Because a group is sequential, a rule whose query takes twenty seconds pushes back everything behind it in that group; if the group's interval is thirty seconds and the total work exceeds it, evaluations start to overrun and effectively skip. Keeping expensive rules in their own group contains that. In the other direction, having *all* rules in one group with the same interval means every query hits the data source at the same instant — a self-inflicted spike. Distributing rules across groups, and where needed offsetting them, smooths the query load. ## What a group is not This is the question's real trap. A group does not aggregate notifications. Two rules in the same evaluation group produce entirely separate alerts and separate notifications unless the *notification policy* groups them by matching labels. Nor does it create an ordering guarantee you can rely on for correctness — the sequential evaluation is a scheduling property, not a transaction; do not design a rule that assumes another rule in the group already ran and did something. ## Interaction with the pending period The pending period on a rule is enforced in evaluations. If a group's interval is one minute and a rule has a five-minute pending period, the condition must hold across five consecutive evaluations. Move that rule into a five-minute group and the same pending period now means five evaluations spanning twenty-five minutes of wall clock — the rule became dramatically less sensitive without anyone editing it. Any change to a group's interval is therefore a change to every member rule's behaviour, and pending periods should be re-checked as multiples of the new interval. ## Choosing a grouping scheme Common workable schemes: - **By cadence.** Groups literally named for their interval, so the interval is visible in the rule's location. Simple and easy to reason about, and it makes the cost of a rule explicit at authoring time. - **By cost.** A group for cheap instant queries and a separate one for heavy range queries, so the expensive ones cannot delay the cheap ones. - **By ownership.** Folders per team (which is also the permission boundary), with groups inside for cadence. This usually reads best in practice because folder permissions and rule ownership already align. Avoid the scheme that feels most natural and is wrong: grouping by *service or incident*, in the hope that the group will bundle the resulting notifications. It will not, and the rules will end up sharing an interval that suits none of them. ## Operational signals Watch for evaluation duration approaching the interval — that is the leading indicator of overrun and skipped evaluations, and it degrades exactly when the system is under stress and the queries get slower. Watch for a single group that has accumulated dozens of rules over time; it started as one interval choice and has become a serial queue. And when many rules query the same data source, remember that the data source's own concurrency limits, not Grafana's, may be the binding constraint.
- You move a rule from a one-minute group into a five-minute group without editing the rule. What changes?Two things. It now evaluates five times less often, so detection latency rises accordingly. More subtly, its pending period is counted in evaluations of the new interval, so a five-minute pending period that used to mean five minutes of continuous breach now spans twenty-five minutes of wall clock. Any group interval change should be followed by re-checking every member rule's pending period as a multiple of the new interval.
- Why can putting every rule in one evaluation group hurt at scale?Rules within a group evaluate sequentially, so total evaluation time grows with the number of rules and a single slow query delays everything behind it; once the total exceeds the interval, evaluations overrun and are effectively skipped. It also concentrates every query into the same instant, producing a periodic spike against the data source rather than a smooth load. Splitting into several groups restores parallelism, contains slow rules, and spreads the query load.
- A colleague wants related alerts to arrive as one message and proposes putting the rules in the same evaluation group. Is that right?No — evaluation groups control scheduling, not notification bundling. Aggregating related alerts into one message is done by the notification policy, which groups firing alerts by chosen labels and holds them for a group-wait period before sending. The correct approach is to give the related rules a shared label such as service or cluster and group notifications on it, leaving evaluation grouping free to be chosen on cadence and cost.
saying these in an interview costs you the question
- Believing an evaluation group bundles notifications for its rules.
- Not realising rules inside a group evaluate sequentially, so one slow query delays the rest.
- Changing a group's interval without re-checking member rules' pending periods.
- Treating the evaluation interval as a global setting rather than a per-group choice matched to data freshness and query cost.
- Designing a rule that depends on another rule in the same group having already evaluated.