A shared database primary fails. Its own alert fires, and so do the alerts of thirty dependent services, paging thirty teams at once. How would you suppress the downstream pages, and what goes wrong with dependency-based suppression?
answer
- one cause, thirty owners
- source suppresses target
- scope with labels that must match
- the model drifts from the architecture
- who fires first matters
basics
~20 sUse inhibition: a firing source alert suppresses matching downstream alerts, scoped by labels that must be equal on both — same cluster or region — so suppression cannot leak across environments. The risk is that the encoded dependency graph drifts from reality and starts hiding genuinely independent failures.
solid answer
~60 sInhibition is the mechanism built for exactly this: you declare that while a source alert matching one set of labels is firing, target alerts matching another set are not notified, provided both agree on a list of labels that must be equal — typically cluster, region and environment — so a database failure in staging can never mute production. In practice you inhibit the noisy downstream symptom alerts with the higher-level cause alert, and route the one surviving page to the team that owns the database. Three things go wrong. First, the rule encodes a dependency graph that drifts as architecture changes, so a service that stopped depending on that database keeps being suppressed. Second, a race: if the downstream alerts have shorter pending durations than the source, they fire first and page before the source exists to inhibit them. Third, over-broad matchers hide an independent failure that happens to look like collateral damage. Inhibition suppresses notification, not alert state, so the suppressed alerts must remain visible to responders and in the incident timeline.
go deeper
Know that one failure low in the stack sets off alerts in every service above it, and that alerting systems can suppress the downstream ones while a known cause alert is firing.
Explain the shape of an inhibition rule — source matcher, target matcher, and the labels that must be equal on both — and why that equality scoping is what keeps a staging failure from muting production.
Demonstrate the operational judgment: the firing-order race between cause and symptom pending durations, the drift of a hand-written dependency model, and the requirement to give suppressed teams a non-paging path to the same information.
Decide whether the org encodes dependencies by hand at all, or derives suppression from a service catalogue; set who owns the rules, how staleness is audited, and how much risk of a missed independent alert you are willing to buy cross-team quiet with.
## The problem: alerts in a dependency relationship Grouping collapses alerts that are *peers* — the same symptom on many instances. It does nothing for alerts in a *dependency* relationship, where a single failure at the bottom of the stack causes correct, distinct, differently-owned alerts to fire above it. Thirty teams paged for one database is not thirty incidents; it is one incident and twenty-nine notifications that will produce no useful action. The cost of getting this wrong runs both ways. Leave it alone and every shared-infrastructure failure becomes an org-wide pager storm in which the one team that can fix it is drowning in cross-talk. Suppress too eagerly and a team sits unaware through an outage that was genuinely theirs. ## The mechanism: inhibition Dependency suppression is usually expressed as an inhibition rule with three parts: - a **source matcher** — which firing alert does the suppressing (for example, the database-down alert at critical severity); - a **target matcher** — which alerts get suppressed while it fires (the dependent services' error-rate alerts, or everything at warning severity from services labelled as depending on that datastore); - a set of **labels that must be equal** on both — cluster, region, environment. That third part is the one candidates forget, and it is the one that keeps the rule safe. Without it, a source alert anywhere suppresses targets everywhere: a database failure in staging silences production symptoms. With it, suppression is scoped to the blast radius the dependency actually has. Crucially, inhibition suppresses **notification**, not the alert. The downstream alerts still evaluate, still fire, still appear in the alert console and still land in the incident timeline. That matters during the postmortem, when you want to know which services were affected and when. ## Failure modes worth naming in an interview **1. The dependency graph drifts.** An inhibition rule is a hand-written model of who depends on whom, and it is correct on the day it is written. Services move off the shared database, new ones arrive, a cache is inserted in front. Nobody updates the inhibition rules, because nothing fails when they are stale — the failure is silent and shows up only as a page that never arrived. Any inhibition rule set needs an owner and a review cadence, and ideally a source of truth (service catalogue, dependency labels emitted by the services themselves) rather than a hand-maintained list. **2. The race on the firing edge.** Inhibition only works if the source alert is *already firing* when the target fires. If the database alert carries a two-minute pending duration and the downstream error-rate alerts carry thirty seconds, the downstream pages go out ninety seconds before there is anything to inhibit them. The fix is to make the source detect at least as fast as its targets — which sometimes means giving the cause alert a shorter pending duration than you would otherwise choose. A related trap is that inhibition works only where the alerts meet: alerts flowing through separate notification pipelines, or a per-service tool that notifies directly, never reach a common point where suppression can apply. **3. Over-broad matchers hide independent failures.** If the target matcher is 'any warning in production', then during any database incident *every* warning in production is muted, including one that has nothing to do with the database. This is the noise-reduction equivalent of a stuck-open valve, and it is worst precisely when it matters — during an incident, when a second, unrelated problem is more likely, not less. **4. The team that needed to know.** Suppressing a downstream team's page removes their notification but not their responsibility: their users are seeing errors. The suppression only works socially if there is a broadcast path — an incident channel, a status page, an automatic notification of affected owners — that reaches them without paging thirty people. Inhibition changes *how* they find out, and you have to design that replacement path deliberately. ## Alternatives and complements - **Alert on the symptom higher up.** If one objective-based alert on user-facing failure would have fired anyway, thirty component alerts add nothing at page severity and could be tickets. This is often a cleaner answer than a large inhibition rule set. - **Dependency-aware correlation** from a service graph, where the pipeline derives suppression from live dependency data rather than static rules. Better hygiene, more moving parts, and it fails in more interesting ways. - **Routing rather than suppression.** Route all downstream alerts of a known cause to one aggregated channel instead of thirty pagers — the information survives, the pages do not. ## What good looks like A small, owned set of inhibition rules for genuine shared-infrastructure dependencies — datastore, cluster control plane, network, a critical upstream — each scoped by equality labels, each with a source alert that detects at least as fast as the targets it suppresses, reviewed when the architecture changes, and paired with a broadcast path that tells affected owners what is happening without paging them. And an explicit acceptance that inhibition trades a small chance of a missed independent alert for a large reduction in cross-team noise — a trade worth naming rather than glossing over.
- Your inhibition rule is configured correctly, yet the downstream teams were paged anyway. What is the most likely cause?A timing race. Inhibition only suppresses targets while the source is already firing, so if the cause alert carries a longer pending duration than the downstream symptom alerts, the downstream pages go out before the source exists. Align the durations so the cause detects at least as fast. The other candidate is alerts that never meet — a team notifying from a separate pipeline that the suppression never touches.
- Thirty teams no longer get paged. How do they find out their service is degraded?You have to build that path deliberately: an incident channel with the affected-service list, a status entry, or an automatic non-paging notification to the owners of every suppressed alert. Inhibition should change how people learn, not whether they learn. If suppressed alerts vanish without a trace, the first review after an incident will demand the pages back.
- Would you ever prefer routing over inhibition for this scenario?Often, yes. Routing the thirty downstream alerts to a single aggregated channel keeps all the information — including which services were hit and when — while producing zero pages. Inhibition is stronger medicine and carries the risk of hiding an unrelated alert; routing degrades more gracefully when the dependency model is wrong.
- How do you keep an inhibition rule set from going stale?Give it an owner and derive it from a source of truth where possible — service-catalogue dependencies or labels the services themselves emit — rather than a hand-maintained list. Then audit it against incidents: after each shared-infrastructure outage, check which suppressions actually applied and which should have. Staleness here fails silently, so it needs an active check.
saying these in an interview costs you the question
- Inhibit all warnings in prod while any critical is firing
- Inhibition stops the downstream alerts from firing at all
- Write the dependency rules once, they don't need review
- Downstream teams don't need to know if we're fixing it
- Grouping and inhibition solve the same problem