CloudWatch offers static-threshold alarms, anomaly-detection alarms and composite alarms. What does each one actually evaluate, and what would make you pick one over the others?
answer
- fixed number versus learned band
- band width in standard deviations
- composite watches alarms, not metrics
- ALARM(), AND, OR, NOT rule
- suppressor collapses correlated storms
basics
~20 sA static alarm compares a metric statistic to a fixed number. An anomaly-detection alarm compares it to a band that CloudWatch learns from the metric's own history. A composite alarm evaluates no metric at all — it is a boolean rule over the states of other alarms.
solid answer
~50 sA static alarm is the default: pick a statistic, a threshold and an M-of-N window. It is predictable, explainable and the only sensible choice when the number is meaningful in itself — a queue depth, a disk at 90%, a quota you must not cross. An anomaly-detection alarm replaces the number with a model: CloudWatch trains on the metric's recent history, learns its daily and weekly shape, and the alarm thresholds against `ANOMALY_DETECTION_BAND(m1, 2)` using operators such as `GreaterThanUpperThreshold` or `LessThanLowerOrGreaterThanUpperThreshold`. That earns its place on strongly seasonal metrics where no fixed number works — request volume, login rate — but it needs history, it will happily learn a degraded state as normal, and you cannot explain to anyone exactly why it fired. A composite alarm has no metric: its `AlarmRule` combines other alarms with `ALARM()`, `OK()`, `AND`, `OR`, `NOT`, and its `ActionsSuppressor` can hold notifications while a broader alarm is active. It is how you turn a storm of correlated alarms into one signal.
go deeper
Be able to say what each type watches: a fixed threshold, a learned expected range, and the states of other alarms — and that only the first two look at a metric at all.
Explain the mechanics — ANOMALY_DETECTION_BAND with a standard-deviation width and the upper/lower comparison operators, and an AlarmRule built from ALARM(), AND, OR and NOT.
Show judgment about when each earns its place: static where the number means something, anomaly detection for seasonal metrics with real history, composites to correlate conditions and suppress a storm behind one shared-dependency alarm.
Own the shape of the alarm estate — how many alarms a team may own, where correlation and suppression structure lives, and the standing cost of anomaly detectors that nobody can explain or tune.
## Static alarms: a number you chose The standard metric alarm compares one statistic to one threshold over an M-of-N window. Everything about it is inspectable: you can point at the graph, point at the line, and say why it fired. It is the right choice whenever the threshold has intrinsic meaning rather than being an empirical guess. A filesystem at 90% full, an SQS `ApproximateAgeOfOldestMessage` above the time your business tolerates, a Lambda concurrency figure approaching the account limit, a certificate expiring — these are all facts about the system, not observations about typical behaviour. Static thresholds are also the only kind that behaves predictably on a brand-new resource with no history. Where they fail is on metrics whose normal value moves: throughput, request volume, active sessions. A number that is right at midday is wrong at midnight, which is the failure mode that pushes people toward the other two types. ## Anomaly-detection alarms: a band you trained CloudWatch anomaly detection fits a model to a metric's history — up to about two weeks of data — capturing its trend and its daily and weekly seasonality, and produces an expected range for each future point. In the alarm you do not supply a threshold value; you supply the band as a metric math expression and point the alarm at it: ``` ANOMALY_DETECTION_BAND(m1, 2) ``` The second argument is the band width in standard deviations — larger means fewer, more extreme alerts. The alarm's `ThresholdMetricId` names the band expression, and the comparison operator is one of `GreaterThanUpperThreshold`, `LessThanLowerThreshold` or `LessThanLowerOrGreaterThanUpperThreshold`. That last one matters: a *drop* in order volume is often the more important signal, and it is the one a static upper-bound alarm can never give you. The honest limitations: - **It needs history.** A new service or a metric with a few days of data gives you a band that means little. - **It normalises whatever it sees.** A slow degradation, or an incident left running for a week, gets absorbed into "normal". Anomaly detection catches sharp deviations, not drift. - **Deployments and campaigns look like anomalies**, because they are. Expect noise around any legitimate step change until the model catches up. - **It is not explainable.** "The model says this is unusual" is a poor opening line at 3am, and nobody can tune it in the way they can move a number. - **You still choose an M-of-N window**, and it still costs — the detector is billed in addition to the alarm. Use it as a wide safety net over seasonal business metrics, alongside static alarms on the things you actually understand — not as a replacement for them. ## Composite alarms: a rule over other alarms A composite alarm evaluates no metric. Its input is an `AlarmRule` — a boolean expression over the states of other alarms: ``` ALARM("api-5xx-rate") AND (ALARM("db-cpu-high") OR ALARM("db-connections-high")) ``` The functions available are `ALARM()`, `OK()`, `INSUFFICIENT_DATA()`, combined with `AND`, `OR`, `NOT` and parentheses. Children may themselves be composite, so you can build a small hierarchy. Two distinct jobs follow from that. **Correlation.** Requiring two independent conditions before paging removes a whole class of false alarms — high CPU that nobody is suffering from, an error rate on a service nobody is calling. Conversely, `NOT ALARM("maintenance-in-progress")` can gate an alarm on a condition you know about. **Suppression.** A composite alarm can name an `ActionsSuppressor` — another alarm whose ALARM state stops this one from taking its actions, with `ActionsSuppressorWaitPeriod` and `ActionsSuppressorExtensionPeriod` controlling the grace either side. The canonical use is a shared-dependency alarm: when the database is down, forty service alarms turn red, and you would like one page rather than forty. The suppressor lets the children keep evaluating (so the dashboard is still truthful) while only the parent notifies. The cost of composites is that state now lives in two places. A composite that never fires because a child alarm has been sitting in INSUFFICIENT_DATA for a month is a common and quiet failure — `INSUFFICIENT_DATA()` is available in the rule precisely so you can handle it explicitly rather than by accident. ## How to choose Start static, because it is explainable and cheap to reason about. Reach for anomaly detection only when the metric's normal value genuinely moves with time of day and no fixed number works, and accept that it is a net rather than a precision instrument. Add composites once you have enough alarms that they fire in correlated groups — they are a structuring tool on top of alarms you already have, never a substitute for having the right underlying alarm.
- Why is a drop in a metric often the signal anomaly detection gives you that a static alarm does not?A static alarm on a busy metric almost always guards the upper side, because the lower bound that would be correct at midday would fire every night. The band moves with the daily shape, so `LessThanLowerOrGreaterThanUpperThreshold` can catch traffic falling out from under you — a broken client, a DNS change, a stuck producer — at any hour, without a separate threshold per time of day.
- A composite alarm has been in OK for weeks while its children clearly fired. What do you check first?Whether a child spent that time in INSUFFICIENT_DATA rather than ALARM. `ALARM(child)` is false when the child is unevaluated, so an AND rule quietly stops firing. Check each child's state history, then either fix the underlying metric gap or make the rule explicit with `INSUFFICIENT_DATA()` so an unevaluated child is treated as a condition rather than as silence.
- What does an ActionsSuppressor change about the alarms it suppresses?Nothing about their evaluation — children keep computing and keep showing their true state, so dashboards and history stay honest. It only stops the composite alarm from executing its actions while the suppressor alarm is in ALARM, with a wait period before suppression takes effect and an extension period after it clears. It is a notification control, not a mute on the underlying signal.
saying these in an interview costs you the question
- Treating anomaly detection as a replacement for static thresholds
- Expecting a useful band on a metric with days of history
- Assuming a composite alarm evaluates metrics of its own
- Forgetting that a child in INSUFFICIENT_DATA breaks an AND rule
- Believing suppression stops the child alarms from evaluating