How do you decide which Datadog monitors notify a human, and keep one failure from becoming a notification storm?
answer
- Not every monitor should reach a person
- The notification message decides who is told
- Grouping decides how many alerts one failure produces
- Composite monitors stop dependent symptoms notifying
- Scheduled downtime mutes while evaluation continues
basics
~10 sA Datadog monitor only reaches a person if its notification message addresses a recipient handle; without one it records state silently. Storms are controlled by deliberate query grouping, composite monitors, and scheduled downtimes.
solid answer
~50 sOn a platform where anything can be monitored, the design question is not what to watch but what to **tell someone about**. In Datadog a monitor evaluates a query against thresholds and changes state, but only reaches a human if its notification message addresses a recipient handle. A monitor with no handle is a recorded signal you read during an investigation - what most monitors should be. Storm control then comes from three places. First, **grouping**: a monitor query grouped by a dimension raises one alert per group, so grouping by host turns one shared-dependency failure into an alert per host. Group by the smallest unit a responder would treat differently. Second, **composite monitors**, so a symptom only notifies while its known cause is healthy. Third, **scheduled downtimes**, which mute notification during deploys and maintenance while the monitor keeps evaluating and recovering normally.
go deeper
Know that a Datadog monitor watches a query against a threshold and changes state, and that it messages people only through the recipients named in its notification message. Recognise that not every monitor is meant to interrupt anyone.
Explain the mechanics: alert and warning thresholds, how a query grouped by a dimension produces one alert per group, and what a scheduled downtime mutes while leaving evaluation running.
Show judgement on a real estate. Say why you notify on symptoms and record causes, how you pick the grouping dimension to match who responds, and how composite monitors and tag-scoped downtimes keep one root cause from generating dozens of notifications.
Own notification volume as a governed number: who may create a notifying monitor, the ownership and response required before one is allowed to, how deploy windows are covered by default, and the review that deletes monitors nobody has ever acted on.
## Two different questions A platform that can monitor anything invites a specific failure: teams create monitors because they can, point them all at the on-call handle, and within a quarter nobody reads the notifications. Good monitor design separates two questions that look like one. **Should this condition be detected?** Almost always yes. Detection is cheap and a recorded state history is what you read during an investigation. **Should this condition interrupt a person?** Much more rarely. Interruption is expensive, and its cost is paid by everyone who is later trained to ignore the channel. Datadog makes the distinction concrete. A monitor evaluates a query on a schedule, compares it with its thresholds, and transitions between states. Notification is a separate act, driven by **the recipients addressed in the monitor's notification message**. A monitor with no recipient in its message is a fully working monitor that records state and appears on the monitor list without waking anyone. This is a feature, not an incomplete configuration, and using it deliberately is most of what separates a curated alerting estate from a noisy one. ## What to check before a monitor is allowed to notify - **Is it a symptom or a cause?** Notify on what users experience. A cause-level condition - a queue depth, a saturated cache - is better recorded, then read during the investigation the symptom triggered. - **Is there an action?** If the responder's honest first move is to look at a dashboard and wait, the monitor is telling them something they cannot act on. - **Is it already covered?** Five monitors that all fire on the same underlying failure are one monitor and four notifications. - **Does it distinguish severity?** A warning threshold alongside the alert threshold lets a condition be visible before it is urgent, without a second monitor. - **What does it do when the data stops?** A monitor that treats missing data as a breach fires constantly on a sparse or batch-driven metric; one that ignores missing data stays silent through the outage that stopped the data arriving. Choose explicitly, per monitor. ## The three storm controls **1. Grouping.** A Datadog metric monitor whose query groups by a dimension becomes a multi-alert monitor: each group is evaluated and alerts independently. This is powerful and it is the most common self-inflicted storm. Group by host on a condition caused by a shared dependency and one failure produces one notification per host. The rule is to group by the smallest unit a responder would treat **differently**: grouping by service when each service has a different owner is useful; grouping by container instance almost never is, because nobody responds per container. **2. Composite monitors.** A composite evaluates the states of other monitors together, so a symptom monitor can be made to notify only while its known dependency monitor is healthy. This is how you stop a database outage generating a notification from every service that talks to it, without giving up per-service detection. **3. Scheduled downtimes.** A downtime mutes notification for monitors matching a scope - typically a tag such as an environment or a service - for a window, optionally recurring. The essential detail is that a muted monitor **keeps evaluating**: it still transitions state and still recovers, so you get an accurate history of what happened during a maintenance window rather than a hole in it. Scope the downtime by tag rather than by listing monitors, so monitors created next month are covered without anyone remembering to add them. ## Putting it together on a real estate | Condition | Notify? | Why | |---|---|---| | Booking confirmations fall below expected rate | Yes, grouped by service | User-visible symptom with a clear owner | | One node's disk filling | No, recorded only | Handled by automation; read during investigation | | Payment provider returning errors | Yes, single alert, ungrouped | One cause, one notification, however many callers | | Latency above target on a canary build | Warning threshold only | Visible without interrupting; rollback is a decision | | Nightly reconciliation job did not report | Yes, with missing-data handling on | Silence is the failure mode for a batch job | For a 41-service estate, the discipline that holds up over time is that every notifying monitor has a named owner and a stated response, everything else records, deploy windows are covered by a recurring downtime scoped by tag, and a monthly review deletes monitors that have notified without ever changing what anyone did. Notification volume is a design output, and it is the number worth watching.
- When is a Datadog composite monitor the right answer, and when is it over-engineering?It earns its place when one known cause reliably produces many dependent symptoms - a shared datastore, an authentication service - and you want detection on every dependent while notification comes from one place. It is over-engineering when the dependency is not actually reliable, because the composite then suppresses a real independent failure. Composites also add a second thing to maintain, so use them where the fan-out is large enough to justify it.
- What do the missing-data settings on a Datadog monitor do, and when do they hurt?They decide whether the absence of data is treated as a condition worth alerting on, and how long the gap must last first. Enabled, they catch the case where an outage stops the telemetry that would have alerted. They hurt on naturally sparse metrics - batch jobs, low-traffic endpoints, anything that legitimately reports nothing at night - where an over-eager window makes the monitor fire every quiet period until people mute it.
- Why scope a scheduled downtime by tag rather than by selecting individual monitors?A tag scope covers monitors that did not exist when the downtime was written, which matters most for recurring deploy or maintenance windows. A hand-picked list decays: someone adds a monitor, the next maintenance window pages the on-call, and the fix is remembered only after the fact. The tag scope also documents intent - this window covers a particular environment or service - rather than an arbitrary set of names.
saying these in an interview costs you the question
- Sends every monitor to the on-call handle by default
- Thinks a monitor with no recipient in its message still notifies someone
- Groups a monitor by host on a condition caused by a shared dependency
- Believes a muted monitor stops evaluating during the window
- Treats missing data and a threshold breach as the same condition
- Creates a new monitor for a symptom already covered by another