A CloudWatch alarm on your service's error-count metric stayed in INSUFFICIENT_DATA throughout a real outage instead of firing. Explain how a CloudWatch alarm evaluates a metric — period, evaluation periods, datapoints to alarm — and how the missing-data treatment decides what happens when the metric stops arriving.
answer
- M breaching datapoints out of N
- gaps are not breaches by default
- TreatMissingData has four settings
- silence looks like INSUFFICIENT_DATA
- actions fire on transitions only
basics
~20 sA CloudWatch alarm counts how many of the last N periods breached the threshold and fires when M of them did. During the outage the metric stopped being published, so there was nothing to compare, and the default missing-data treatment leaves the alarm unevaluated rather than breaching.
solid answer
~60 sAn alarm is defined by a period, a statistic, a threshold, `EvaluationPeriods` (N) and optionally `DatapointsToAlarm` (M): it looks at the last N periods and goes to ALARM when at least M of those datapoints breach. The trap is what counts as a datapoint. Many metrics are only published when something happens — a `Sum` of errors emits nothing when the service is emitting nothing at all — so an outage that stops the process produces a gap, not a breach. `TreatMissingData` decides how a gap is scored: `missing` (the default) treats it as neither good nor bad and leaves the alarm in its current state, dropping to INSUFFICIENT_DATA when the whole evaluation range is empty; `notBreaching` scores it as OK; `breaching` scores it as a breach; `ignore` freezes the state entirely. The fix is either `breaching` on a metric that must always be present, or alarming on a metric that is emitted continuously — a heartbeat, or a request-count metric — rather than one that only appears on failure.
code
bash · 12 linesaws cloudwatch put-metric-alarm \
--alarm-name checkout-errors \
--namespace MyApp \
--metric-name Errors \
--statistic Sum \
--period 60 \
--evaluation-periods 5 \
--datapoints-to-alarm 3 \
--threshold 5 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data breaching \
--alarm-actions arn:aws:sns:eu-west-1:111122223333:ops-pagergo deeper
Be able to name the alarm's inputs — period, statistic, threshold, evaluation periods — and the three states, and say that a gap in the data is not the same thing as a breach.
Explain M-out-of-N evaluation and all four TreatMissingData values, and work out why an alarm on an error count goes to INSUFFICIENT_DATA when the process publishing it dies.
Show the operational reasoning: pick a continuously-published metric or a breaching treatment so silence is detected, compute the real detection latency from period times evaluation periods, and know that actions fire only on transitions.
Own the convention across a fleet — which signals every service must publish so that absence is detectable, what the default period and evaluation shape should be, and how alarm state reaches a notification pipeline that can escalate.
## The five settings that define an alarm A metric alarm is not "a line on a graph". It is a small state machine over a fixed set of inputs: - **Period** — how wide each datapoint is (60 seconds, 300 seconds, and so on). This must match how the metric is actually published; asking for a 60-second period on a metric published every five minutes gives you four empty periods out of every five. - **Statistic** — how the raw observations inside a period are collapsed: `Sum`, `Average`, `Maximum`, `Minimum`, `SampleCount`, or a percentile such as `p99`. - **Threshold and comparison operator** — for example `GreaterThanOrEqualToThreshold` with a threshold of 5. - **EvaluationPeriods (N)** — how many recent periods the alarm considers. - **DatapointsToAlarm (M)** — how many of those N must breach. When you omit it, M equals N. The M-out-of-N form exists to trade responsiveness against flapping. `3 out of 3` needs a sustained condition and reacts slowly; `2 out of 5` tolerates a single spiky period while still catching a persistent problem within roughly five periods. ## The three states An alarm is in `OK`, `ALARM`, or `INSUFFICIENT_DATA`. Actions — an SNS notification, a scaling action — are attached to **transitions** into a state, not to being in it. An alarm that has been in ALARM for six hours has sent exactly one notification. That surprises people who expect a repeat page, and it is why re-notification is a job for whatever consumes the SNS topic. ## What actually happens when data is missing CloudWatch first tries hard to find datapoints. If some of the last N periods are empty, it extends the window backwards — the *evaluation range* — collecting up to N real datapoints from a longer stretch of time before it gives up. Only genuinely absent data reaches the `TreatMissingData` setting, which has four values: | Setting | A missing period is scored as | Effect | |---|---|---| | `missing` (default) | not evaluated | Alarm keeps its current state; goes INSUFFICIENT_DATA when the whole range is empty | | `notBreaching` | good | Gaps push the alarm towards OK | | `breaching` | bad | Gaps push the alarm towards ALARM | | `ignore` | — | State is frozen; the alarm never changes on missing data | The failure in the question is the default at work. The alarm watched `Sum` of an error metric. When the service died it stopped publishing anything, including errors. Every period became missing, the evaluation range emptied, and the alarm slid into INSUFFICIENT_DATA — technically correct, operationally useless. ## Three ways to fix it **Score gaps as breaching.** If the metric is genuinely published on every period the service is healthy, `--treat-missing-data breaching` makes silence itself the alarm condition. Use it only when you are sure of the publication cadence, or a deployment restart will page you. **Alarm on a metric that is always present.** AWS-vended metrics differ here: an SQS queue publishes `ApproximateAgeOfOldestMessage` every minute whether or not anything is happening, whereas Lambda publishes `Errors` only when invocations occurred. Prefer the continuously-published signal, or watch a positive signal (request count, successful heartbeat) with a `LessThanThreshold` comparison, which fires when the number drops rather than when a rare number appears. **Fill the gap in metric math.** An expression such as `FILL(m1, 0)` substitutes zero for missing periods, converting a gap into a real datapoint before the alarm sees it. That is the right tool when you want zero to mean zero — but note it also means an alarm on `FILL(errors, 0)` will never fire from silence, which may be the opposite of what you want. ```bash aws cloudwatch put-metric-alarm \ --alarm-name checkout-errors \ --namespace MyApp --metric-name Errors \ --statistic Sum --period 60 \ --evaluation-periods 5 --datapoints-to-alarm 3 \ --threshold 5 --comparison-operator GreaterThanOrEqualToThreshold \ --treat-missing-data breaching \ --alarm-actions arn:aws:sns:eu-west-1:111122223333:ops-pager ``` ## Two more sharp edges An alarm on a **percentile** statistic has an extra setting, `EvaluateLowSampleCountPercentile`. Set it to `ignore` and periods with too few samples do not move the alarm — a single slow request in a quiet minute will not page you. And an alarm evaluates on a delay: the metric has to be published, aggregated and made available before the period can be judged, so the practical detection time is roughly `period × N` plus a minute or two. Choosing a five-minute period with three evaluation periods buys you a sixteen-minute worst case before anyone is told.
- Your alarm has been in ALARM for two hours and nobody has been paged since the first message. Is that a bug?No — alarm actions fire on a state transition, not repeatedly while the state holds. CloudWatch sends one message into the SNS topic when the alarm enters ALARM and another when it returns to OK. Repeat notification, escalation and deduplication belong to whatever consumes the topic; CloudWatch itself has no re-notify interval.
- How would you choose between 3-out-of-3 and 2-out-of-5 for the same threshold?Both need a sustained problem, but they trade differently. 3-of-3 requires an unbroken run, so one recovered period resets the count — slow to fire and easy to miss an intermittent fault. 2-of-5 tolerates gaps and spiky recovery, catching flapping conditions, at the cost of firing on two unlucky periods spread across five. Match the shape of the failure you expect.
- What does the low-sample-count setting do on a percentile alarm?`EvaluateLowSampleCountPercentile` decides whether periods with too few observations are judged at all. With `evaluate` (the default) a p99 computed from three requests counts like any other datapoint; with `ignore` such periods are skipped, which stops a single slow request during quiet hours from taking the alarm into ALARM.
saying these in an interview costs you the question
- Assuming a missing datapoint counts as a breach by default
- Thinking one breaching datapoint always fires the alarm
- Expecting repeated notifications while the alarm stays in ALARM
- Alarming on an error count that only exists when traffic exists
- Setting a 60-second period on a metric published every five minutes