skip to content

Your team will fail over a database at 02:00 and expects a burst of alerts. What is a maintenance silence, what does it actually stop, and why should every silence carry an expiry, a narrow scope and an owner?

level: juniorimportance: should knowfreq 48%

answer

  1. planned work, expected alerts
  2. mutes notification, not evaluation
  3. always bounded in time
  4. narrow matchers keep the rest live
  5. who made it and why

basics

~20 s

A maintenance silence suppresses notifications for alerts matching a set of criteria during a bounded window. It stops the notification only — rules keep evaluating and alerts stay visible. It needs an expiry so it cannot outlive the work, a narrow scope so unrelated failures still page, and an owner so someone can explain it.

solid answer

~50 s

A silence is a temporary, matcher-scoped suppression you create *before* planned work so that expected alerts do not page anyone. The key thing to be clear about is what it suppresses: notification. The alert rules keep evaluating, the alerts still fire, and they remain visible in the alert console — you have muted the pager, not blinded the monitoring. Three properties make a silence safe. An **expiry**, because an open-ended silence is the classic way a service ends up unmonitored for months; set it to the planned window plus a small margin and extend deliberately if the work overruns. A **narrow scope**, matching the specific service and cluster under maintenance rather than everything at paging severity, so a genuinely unrelated outage at 02:15 still reaches somebody. And an **owner and reason** recorded on it, so the next person who finds it knows whether it is still needed. Afterwards, check what fired under the silence before you let it expire quietly.

go deeper

for a junior

Be able to say that a silence suppresses notifications for matching alerts during planned work, that it must have an end time, and that the alerts themselves still fire and stay visible.

for a middle

Contrast a silence with pausing or deleting a rule and with rerouting, explain what each one stops, and describe how matcher scope decides which unrelated failures can still page during the window.

for a senior

Show the operating discipline: size the window to the plan, keep the user-facing symptom alerts outside the matchers, read what fired under the silence afterwards, and confirm alerting is live again before calling the work done.

for a principal

Set the policy — a maximum silence duration, mandatory owner and reason, automatic expiry reporting, and an audit of long-lived silences — so that no service can drift into being unmonitored without someone being told.

## What a silence is for Some alerts are correct and expected. If you are about to fail a database over, the connection errors, the elevated latency and the replica-lag alerts are all going to fire, and they are all going to be true. Paging a human about them is pure noise — the human already knows, because they are the one doing it. A maintenance silence is the instrument for that: a temporary suppression, created in advance, scoped by matchers to the alerts you expect, and bounded by a start and end time. Alertmanager-style silences and the equivalents in other alerting stacks all share that shape, and scheduled recurring versions exist for regularly-timed work. ## What a silence actually stops This is the part candidates get wrong, and it is worth being precise: - Alert **rules keep evaluating.** Nothing about the metric pipeline changes. - Alerts **still fire and still exist.** They appear in the alert console, usually flagged as silenced, and they are available afterwards for review. - Only the **notification** is suppressed. No page, no ticket, no chat message. That distinction matters twice. During the work, an engineer can look at the console and see exactly which of the expected alerts fired — useful confirmation that the failover behaved as predicted. Afterwards, the incident or change review can reconstruct what happened, because the alert history is intact. It also distinguishes a silence from two neighbouring actions: **pausing or deleting the rule**, which stops evaluation and leaves no record, and **routing the alert elsewhere**, which keeps a notification but sends it somewhere quieter. ## Expiry: the single most important property An open-ended silence is how services go dark. The pattern is always the same: someone silences an alert during an incident or a migration, the work finishes, nobody remembers, and three months later a real failure produces no page at all. Because the failure mode is *silence*, nothing complains. So: every silence gets an end time, sized to the planned window plus a modest margin. If the work overruns, extending the silence is a deliberate act someone has to take — which is exactly the checkpoint you want. A good platform makes the maximum silence duration a policy (hours, not weeks), lists silences that are about to expire, and reports long-lived silences to their owners for renewal or removal. ## Scope: match narrowly The lazy silence matches everything at paging severity, or everything in the production cluster. It works — nothing pages — and it also means that if an unrelated service falls over at 02:15, nobody finds out until morning. Maintenance windows are not a safe time to be blind: you are actively changing production, so the probability of a problem is *higher* than usual, not lower. Match the service, the cluster and, where you can, the specific alert names you expect. Leave the broad user-facing symptom alerts outside the silence, so that if the failover genuinely breaks the product, someone still hears about it. A useful test question: *if this maintenance goes catastrophically wrong, which alert tells us — and is it inside my matchers?* If the answer is yes, narrow the scope. ## Owner and reason Silences are found later by people who did not create them, usually while wondering why an alert did not fire. Record who created it, why, and what work it covers. Most tools capture the creator automatically and offer a comment field; use it, and reference the change or ticket. A silence with no explanation will either be removed at the worst moment or left in place forever, and neither is good. ## After the window Two habits separate teams that do this well: 1. **Read what fired under the silence** before it expires. If something fired that you did not expect, the maintenance did something you did not expect. 2. **Confirm the silence has actually ended** and alerting is live again — a quick check that the alerts have cleared. If they have not, you have not finished the work; you have finished the window. ## The trade A silence buys quiet at the cost of a blind spot. You keep the price low by bounding it in time, narrowing it in scope, and recording who owns it — and by treating a silence as something you use for *known, planned* work, never as a fix for a rule that fires too often in normal operation. That rule needs retuning or retiring, not muting.

  • What is the difference between silencing an alert and pausing or deleting its rule?
    A silence stops notification only — the rule still evaluates, the alert still fires, and it is visible in the console and in the history afterwards. Pausing or deleting the rule stops evaluation, so there is no record that the condition ever occurred. For bounded maintenance you want the silence, because you keep the evidence of what the change did.
  • The maintenance overruns and the silence is about to expire. What do you do?
    Extend it deliberately, for a specific additional window, and say so in the change channel. The expiry doing its job is a feature: it forces a conscious decision at a moment when someone should be asking whether the work is still on track. What you do not do is remove the end time to make the prompt go away.
  • A teammate silences a chronically flapping alert for two weeks 'until we get to it'. What's wrong with that?
    It converts a tuning problem into a blind spot with a two-week fuse, and the fix predictably never happens. A rule that fires too often in normal operation should have its threshold or pending duration corrected, be demoted off the paging path, or be retired outright — all of which leave the team with a monitored service. A long silence leaves them with an unmonitored one.

saying these in an interview costs you the question

  • Silence everything in the cluster, it's just for an hour
  • Silences don't need an end time if you remember to remove them
  • A silence stops the alert rule from evaluating
  • Use a long silence instead of fixing a noisy rule
  • Nothing fired under the silence, so nothing to review

context