skip to content

A CloudWatch alarm on the raw count of 5xx responses from your load balancer pages at every traffic peak and stays silent during a quiet-hour outage. How would you use CloudWatch metric math to alarm on an error rate instead, and what must you handle for the alarm to behave when traffic drops to zero?

level: seniorimportance: should knowfreq 46%

answer

  1. a count scales with traffic
  2. ratio is scale-free
  3. gaps in, gaps out
  4. FILL and IF guard the denominator
  5. alarm needs exactly one series

basics

~20 s

Alarm on a ratio rather than a count: use CloudWatch metric math to divide the 5xx count by the request count and threshold the percentage. Then guard the expression so that zero traffic produces a defined value instead of a gap that leaves the alarm unevaluated.

solid answer

~50 s

A count scales with traffic, so any fixed threshold is simultaneously too low at peak and too high at 3am. In CloudWatch you build the alarm from a `Metrics` array instead of a single metric: `m1` is `HTTPCode_Target_5XX_Count` with `Sum`, `m2` is `RequestCount` with `Sum`, both with `ReturnData` false, and an expression `e1` with `ReturnData` true computes `100 * m1 / m2`. Two things then need care. First, when there is no traffic both series stop producing datapoints, and arithmetic over missing data yields no datapoint — so wrap the inputs in `FILL(m1, 0)` and guard the denominator with `IF` so a quiet period evaluates to zero rather than a gap, and set the missing-data treatment deliberately. Second, a rate is meaningless on tiny volumes: three requests, one failure, 33%. Fold a minimum-volume condition into the `IF`, or pair the rate alarm with a request-count alarm in a composite. The expression must return exactly one time series, and `SEARCH()` expressions cannot back an alarm.

code

json · 16 lines
json
[
  { "Id": "m1", "ReturnData": false,
    "MetricStat": {
      "Metric": { "Namespace": "AWS/ApplicationELB",
                  "MetricName": "HTTPCode_Target_5XX_Count",
                  "Dimensions": [{ "Name": "LoadBalancer", "Value": "app/prod-alb/50dc6c495c0c9188" }] },
      "Period": 60, "Stat": "Sum" } },
  { "Id": "m2", "ReturnData": false,
    "MetricStat": {
      "Metric": { "Namespace": "AWS/ApplicationELB",
                  "MetricName": "RequestCount",
                  "Dimensions": [{ "Name": "LoadBalancer", "Value": "app/prod-alb/50dc6c495c0c9188" }] },
      "Period": 60, "Stat": "Sum" } },
  { "Id": "e1", "ReturnData": true, "Label": "5xx rate %",
    "Expression": "IF(FILL(m2,0) > 100, 100*FILL(m1,0)/m2, 0)" }
]

go deeper

for a junior

Recognise that CloudWatch can compute expressions over metrics, and that an error rate is a better alarm signal than an error count because it does not move with traffic volume.

for a middle

Build the alarm: a Metrics array with two MetricStat entries and one expression marked ReturnData, and explain what FILL and IF do to missing datapoints.

for a senior

Show the operational judgment — the zero-traffic and small-sample failure modes, deliberately choosing the missing-data treatment, and adding a separate low-traffic alarm because a rate alarm can never fire from silence.

for a principal

Decide the convention for the estate: which ratio signals every service alarms on, whether volume gating lives inline or in a composite, and how those alarm definitions get generated consistently rather than hand-built per service.

## Why the count alarm is wrong in both directions An absolute error count is a proxy for two different things at once: how broken the system is, and how busy it is. At 5,000 requests per minute, 50 errors is a one-percent blip nobody needs to wake for. At 30 requests per minute, 20 errors is two-thirds of your traffic failing and sits comfortably below the same threshold. Any single number you pick is wrong at one end of the daily curve, which is why the count alarm both pages spuriously at peak and sleeps through the quiet-hour outage. The rate — errors divided by requests — is scale-free, which is exactly the property you want in a threshold. ## Building the expression A CloudWatch alarm can be defined over a `Metrics` array rather than a single namespace/metric pair. Each entry is either a `MetricStat` (a real metric with a period and statistic) or an `Expression`. Exactly one entry sets `ReturnData: true`, and that is what the alarm thresholds: ```json [ { "Id": "m1", "ReturnData": false, "MetricStat": { "Metric": { "Namespace": "AWS/ApplicationELB", "MetricName": "HTTPCode_Target_5XX_Count", "Dimensions": [{ "Name": "LoadBalancer", "Value": "app/prod-alb/50dc6c495c0c9188" }] }, "Period": 60, "Stat": "Sum" } }, { "Id": "m2", "ReturnData": false, "MetricStat": { "Metric": { "Namespace": "AWS/ApplicationELB", "MetricName": "RequestCount", "Dimensions": [{ "Name": "LoadBalancer", "Value": "app/prod-alb/50dc6c495c0c9188" }] }, "Period": 60, "Stat": "Sum" } }, { "Id": "e1", "ReturnData": true, "Label": "5xx rate %", "Expression": "IF(FILL(m2,0) > 100, 100*FILL(m1,0)/m2, 0)" } ] ``` When an alarm is defined this way you do not set `MetricName`, `Namespace`, `Statistic` or `Period` on the alarm itself — the period lives inside each `MetricStat`. ## The zero-traffic problem Metric math is not SQL: an operation whose input has no datapoint at a timestamp produces no datapoint at that timestamp. Two failure shapes follow. If the numerator is missing (no errors this minute, which is the healthy case), `m1/m2` yields nothing, and the alarm sees a gap rather than a reassuring zero. With the default missing-data treatment the alarm holds its previous state — so an alarm that fired once can stay in ALARM through a healthy period. If the denominator is missing or zero (genuinely no traffic), there is nothing sensible to divide by at all. `FILL(m1, 0)` substitutes zero for missing periods in a series, converting the first case into a real datapoint. The second needs a conditional: `IF(condition, trueValue, falseValue)` evaluates elementwise across the time series, so `IF(FILL(m2,0) > 100, 100*FILL(m1,0)/m2, 0)` yields the real rate when volume is meaningful and a flat zero otherwise. That single expression solves the zero-division and the small-sample problem together. Set `TreatMissingData` deliberately on top of this. Once the expression always produces a datapoint, `notBreaching` is usually right — but remember you have now built an alarm that will *never* fire from silence. If a dead load balancer is a condition you also need to catch, that is a separate alarm on request count with a `LessThanThreshold` comparison, not a tweak to this one. ## Volume gating, and when to split it out The minimum-volume guard can live either inside the `IF` (as above) or as a second alarm combined with a composite rule such as `ALARM(error-rate-high) AND ALARM(traffic-present)`. Inline is simpler and keeps everything in one object; the composite form is clearer to read on a page and lets you reuse the traffic alarm elsewhere. Either way, do not skip it — an unguarded rate alarm on a low-traffic service is a pager that fires on statistical noise. ## Functions worth knowing, and the limits Beyond `FILL` and `IF`, the frequently useful ones are `SUM`, `AVG`, `MAX` and `MIN` across an array of series; `RATE(m1)` for per-second change; `DIFF(m1)` for the difference between consecutive points; `METRICS()` to refer to every metric in the request; and `SEARCH()` to match many series by a query. `SEARCH()` is the important exception: it returns an unknown number of series, and an alarm needs exactly one, so **you cannot create an alarm directly on a `SEARCH()` expression**. You either wrap it in an aggregating function that collapses it to a single series, or accept that it is a dashboard tool. More generally, if your `ReturnData` expression yields multiple series, alarm creation fails. Finally, mind the periods. Mixing a 60-second metric with a 300-second one in the same expression forces alignment, and an alarm on a five-minute period reacts in five-minute steps no matter how fine the other input is.

  • Why does 100 * m1 / m2 produce no datapoint at all during a healthy minute with zero errors?
    Because the error metric publishes nothing when there are no errors, and metric math skips timestamps where an input has no datapoint — so the expression emits a gap, not a zero. `FILL(m1, 0)` materialises the missing periods as zeros before the division. Without it the alarm evaluates on gaps and, with the default missing-data treatment, simply holds whatever state it was already in.
  • You want the same rate alarm across forty target groups without writing forty alarms. What stops you?
    The natural tool, `SEARCH()`, returns many series, and an alarm must threshold exactly one — so you cannot create an alarm on a `SEARCH()` expression. Your options are to collapse it with an aggregating function (which hides which target group is failing), or to generate one alarm per target group from your account baseline tooling. Dashboards can use `SEARCH()` freely; alarms cannot.
  • Would you gate on minimum volume inside the expression or with a composite alarm?
    Both work. Inline `IF` keeps it as one object and one state to reason about, which suits a service-level alarm you copy everywhere. A composite — rate alarm AND traffic-present alarm — makes the two conditions separately visible on a page and lets the traffic alarm be reused, at the cost of three objects to keep in step. Pick one convention and apply it consistently.

saying these in an interview costs you the question

  • Thresholding an absolute error count and calling it an SLI
  • Assuming metric math treats a missing datapoint as zero
  • Dividing by a request count that can be zero
  • Alarming on a rate with no minimum-volume guard
  • Trying to create an alarm on a SEARCH() expression

context