skip to content

In Alertmanager, when do you write an inhibition rule instead of creating a silence?

level: seniorimportance: should knowfreq 46%

answer

  1. Two ways to make an alert stop notifying
  2. One is conditional, one is timed
  3. Configuration versus a runtime object
  4. Matching labels must agree on both sides
  5. Neither stops the rule from evaluating

basics

~20 s

Write an Alertmanager inhibition rule for a standing cause-and-effect relationship: it mutes matching alerts only while a source alert is firing, and lifts itself. Create a silence for a known, time-boxed window such as maintenance.

solid answer

~50 s

An **inhibition rule** lives in `alertmanager.yml` under `inhibit_rules`, with `source_matchers`, `target_matchers` and an `equal` list of labels that must hold the same value on both sides. While at least one source alert is firing, matching target alerts are muted — automatically, and only within the scope the `equal` labels tie together. A **silence** is created at runtime through the Alertmanager UI or API with matchers, an expiry and a comment; it mutes anything that matches for that window and then dies. A relationship you can state in advance and want enforced forever is an inhibition rule; a window you know about now is a silence. When one root cause fans out into fifty alerts, inhibition is the right tool, because it is conditional on the cause still firing and lifts by itself when the cause resolves. Neither one stops rule evaluation: the alerts are still firing, just not notified.

code

yaml · 6 lines
yaml
inhibit_rules:
  - source_matchers:
      - alertname = "CacheTierDown"
    target_matchers:
      - severity = "warning"
    equal: ['cluster', 'service']

go deeper

for a junior

Know that Alertmanager can stop notifications in two ways: a silence you create for a time window, and an inhibition rule that mutes alerts while another alert is firing. Neither deletes the alert.

for a middle

Explain the three parts of an inhibition rule, especially why the equal labels are needed to keep suppression scoped, and contrast that with a silence's matchers and expiry. Say where each one is created.

for a senior

Show judgement on a real fan-out: pick inhibition because it is conditional, scoped and reviewed, and diagnose the rule that silently never matches. Know that suppressed alerts are still firing and still visible.

for a principal

Own the policy. Decide which cause-and-effect relationships are encoded centrally, who may create long silences and how they expire, and where suppression stops being legitimate and becomes a substitute for fixing the rules.

Alertmanager has two suppression mechanisms and they are not interchangeable. One is a standing rule in configuration; the other is an ad-hoc, expiring object created at runtime. Choosing between them comes down to whether the thing you want suppressed is *conditional on another alert* or *bounded by a clock*. ## Inhibition Inhibition is declarative and lives in `alertmanager.yml`: ```yaml inhibit_rules: - source_matchers: - severity = "critical" target_matchers: - severity = "warning" equal: ['alertname', 'cluster', 'service'] ``` Three parts do the work: - **`source_matchers`** select the alerts that *cause* suppression. - **`target_matchers`** select the alerts that get *muted*. - **`equal`** names labels that must have **identical values** on the source and the target for the muting to apply. That `equal` list is the whole point and the field people forget. Without it, a single critical alert anywhere in the estate would mute every warning everywhere. With `equal: ['cluster', 'service']`, a critical alert on the check-in service in one cluster mutes only warnings about that same service in that same cluster. Get the list too wide and inhibition does nothing; get it too narrow — include a label like `instance` that differs between cause and effect — and it also does nothing, silently. There is no error either way, which is why a rule that has never inhibited anything can sit in a config for a year unnoticed. Two behaviours worth knowing. Inhibition applies for exactly as long as a matching source alert is firing, so it lifts itself the moment the cause resolves — nobody has to remember to remove it. And an alert that matches both the source and the target matchers is not inhibited by itself, so a rule that overlaps on both sides does not self-suppress. ## Silences A silence is imperative and created at runtime, from the Alertmanager UI, its API, or `amtool silence add`. It carries a set of matchers, a start and end time, a creator and a comment. Anything matching is muted until it expires, and then it is simply gone — no configuration change, no deploy, no review. That is the strength and the weakness. A silence is instant and requires no access to the config repository, which is what you want at 02:40 during a planned migration. It is also unconditional: it mutes whatever matches for the whole window, including a genuinely new and unrelated problem that arrives during it. A broad silence with a long expiry is the most common way a team goes quietly blind. ## Side by side | | Inhibition rule | Silence | |---|---|---| | Where it lives | `alertmanager.yml`, reviewed and deployed | Runtime object, UI / API / `amtool` | | Trigger | Another alert firing, scoped by `equal` labels | A time window you chose | | Lifetime | As long as the source alert fires | Until its expiry | | Made by | The team that owns the config | Anyone with access, during an incident | | Suppresses | Only alerts matching `target_matchers` in that scope | Anything matching, for the duration | | Risk | Silently never matches, or mutes too broadly | Outlives its reason and hides new problems | ## Choosing, and the fan-out case Take a climbing-gym membership platform whose metrics store holds about 2.3 million series, and where one team's services produce roughly seventy per cent of the alert volume. A shared cache tier fails, and 46 alerts fire: one for the cache itself, the rest for the dependent services timing out. An inhibition rule is the right answer here, for three reasons: 1. **It is conditional.** Suppression exists only while the cause alert fires. When the cache recovers, the downstream alerts become visible again with no human action. 2. **It is scoped.** The `equal` labels confine the suppression to the affected cluster or environment; an unrelated failure elsewhere still pages. 3. **It is written once.** The cause-and-effect relationship is a property of the architecture, not of tonight. It should be reviewed like code. A silence would work tonight and be wrong as a policy: someone must create it under pressure, guess a duration, and remember to expire it, and while it lives it hides anything else that matches. The converse is just as clear. For a maintenance window on Thursday morning, there is no source alert to hang an inhibition on; the trigger is a clock, so it is a silence — narrowly matched, with an owner and an expiry that ends with the work. ## What neither one does Both act on **notification only**. Prometheus still evaluates the rule, the alert is still firing, the `ALERTS` series still exists, and the alert is still visible in Alertmanager marked as suppressed. Nothing about a dashboard changes. That is a feature: post-incident you can still see everything that fired, including what was suppressed and why. It is also the trap for anyone who reasons that a quiet pager means a healthy system. The related failure mode is treating suppression as a fix for a noisy alert set. Both mechanisms hide symptoms of a rule that fans out too widely or fires on a cause nobody acts on; neither improves the rule.

  • An inhibition rule has been in the config for a year and has never suppressed anything. What do you check first?
    The `equal` list. Every label in it must carry the same value on the source and the target alert, so including one that differs between cause and effect — `instance` is the classic — means the rule can never match. Check next that the source alert actually fires with the labels the matchers expect. Neither mistake produces an error.
  • A team says an alert is suppressed, so the underlying problem must be gone. What is wrong with that?
    Suppression acts on notification only. The rule is still evaluated, the alert is still firing, and it is visible in Alertmanager marked suppressed and on any dashboard querying the alert state. A quiet pager means nobody is being told, not that the system is healthy — which is exactly why silences need owners and expiries.
  • Why is a broad silence with a long expiry considered a hazard rather than a convenience?
    It is unconditional: for its whole window it mutes everything matching, including a genuinely new problem that appears an hour later. Nothing in the mechanism reasons about causes, so it cannot distinguish the failure you expected from the one you did not. Narrow matchers, a short expiry and a named owner are the discipline that keeps it safe.

saying these in an interview costs you the question

  • Thinks suppression stops the alerting rule from evaluating
  • Writes an inhibition rule without an equal label list
  • Believes a silence lifts itself when the cause resolves
  • Says an inhibited alert disappears from Alertmanager entirely
  • Treats suppression as the fix for a badly scoped rule
  • Confuses inhibition with grouping several alerts into one message