skip to content

A cost anomaly rule on a checkout API's daily spend fires every Monday morning — why, and what baseline stops it?

level: seniorimportance: should knowfreq 44%

answer

  1. compared against what, exactly?
  2. shape, not a constant
  3. this Monday against recent Mondays
  4. median and spread, not a mean
  5. a slow ramp never trips it

basics

~20 s

A flat threshold treats every day as identical, so a recurring weekly peak always looks anomalous. A useful baseline compares like with like — this Monday against recent Mondays — using a robust centre and a minimum absolute money delta before anything is reported.

solid answer

~50 s

Anomalous means *deviating from expectation*, so the rule is only as good as the expectation it carries. A fixed daily number is a baseline that says every day should look the same, which is false for any workload with weekly shape — so the Monday peak trips it every week and the team learns to ignore it. Replace the constant with **comparable periods**: this Monday against recent Mondays, this hour against the same hour on recent days. Use a **robust centre and spread** — a median rather than a mean, which one outlier drags toward itself — and require a **minimum absolute delta**, so a tiny service doubling its trivial spend does not report. Evaluate **per dimension**, not on the account total, where one thing rising and another falling cancel out. And re-establish the baseline deliberately after any intentional change.

code

pseudocode · 13 lines
pseudocode
# compare like with like: this Monday against recent complete Mondays
comparable = [d.spend for d in history
              if d.weekday == today.weekday and d.isComplete]

centre = median(comparable)
spread = median(abs(x - centre) for x in comparable)

delta = today.spend - centre

# both tests must pass: deviation relative to normal variation,
# AND enough absolute money to be worth anyone's attention
if delta > 3 * spread and delta > minimumAbsoluteDelta:
    report(scope = today.scope, expected = centre, observed = today.spend)

go deeper

for a junior

Recall that an anomaly is a deviation from an expectation, and that a fixed daily number is an expectation that every day costs the same. Know why that fails for a workload busier on some days than others.

for a middle

Explain how to build a baseline from comparable periods — same weekday, same hour — and why a median beats a mean when the thing you are detecting is an outlier.

for a senior

Demonstrate the operating judgment: a minimum absolute delta so trivial services stay quiet, evaluation per dimension so offsetting movements do not cancel, re-anchoring after deliberate changes, and knowing that a gradual ramp needs a different signal entirely.

for a principal

The trade-off you own is which blind spot the organisation accepts. Anomaly, cumulative threshold and per-unit figure each miss something the others catch, and a programme that funds only one should be able to name what it has chosen not to see.

## What "anomalous" has to mean An anomaly is a deviation from an expectation, which means every anomaly rule contains a **baseline** whether or not its author thought about one. A rule of the form *"alert if daily spend exceeds this number"* has a baseline too: a constant. That baseline asserts that every day of the week, every day of the month, and every phase of the workload's life should cost the same. For a workload with a weekly traffic shape, that assertion is false once a week, on schedule. This is why the Monday alert is not a tuning problem. Raising the constant until Monday stops firing sets it above the busiest day, which means the rule can now only detect a change larger than the entire weekly swing — and it has become a rule that alerts on nothing while appearing to be on. ## Building a baseline that matches the shape - **Compare like with like.** The comparison for a Monday is recent Mondays, not yesterday. Where spend has an hour-of-day shape, the comparison for 09:00 is 09:00 on recent days. - **Use a robust centre.** A mean over recent comparable periods is pulled toward exactly the kind of outlier you are trying to detect; if last Monday was the runaway, it inflates the expectation and hides the next one. A median resists that. - **Measure spread, do not guess it.** A fixed percentage band is another constant in disguise. Derive the band from how much the comparable periods actually varied. - **Require a minimum absolute delta.** Relative deviation alone means a small service that went from a trivial amount to twice a trivial amount reports as loudly as one that added real money. Report only when both the relative and the absolute test pass. - **Evaluate per dimension and per scope.** On an account total, a rise in one dimension and a fall in another cancel, and the rule stays quiet through a real change. - **Exclude the in-progress period.** Today is not finished and the most recent hours are still arriving, so including them drags the baseline down and manufactures both false highs and false lows. - **Re-establish after a deliberate change.** Triple capacity on purpose and the old baseline will report the new level as anomalous every day until enough comparable periods have accumulated at the new level. That is the rule working correctly and being useless; the repair is to re-anchor when you make the change, not to argue with it afterwards. ## The two failure modes | | False positive | False negative | |---|---|---| | Typical cause | Recurring shape treated as flat; a deliberate change; an in-progress period compared to a complete one | A slow ramp the baseline follows; offsetting movements inside one total; a cost that was wrong from day one | | What it costs | The rule gets muted, which removes the detection you thought you had | Nothing reports, and the spend accrues for weeks | | Repair | Shape the baseline, re-anchor on purpose, compare complete periods | Compare against a longer window too, evaluate per dimension, watch unit cost separately | ## What anomaly detection structurally cannot catch This is the part that matters most for judgment, because it defines what else you need: 1. **A gradual ramp.** A baseline built from recent comparable periods follows a slow climb, so every individual day is unremarkable against the days before it. Ten per cent a week is invisible to an anomaly rule and obvious against the period budget. 2. **A cost that was always wrong.** Something provisioned oversized on the day it was created, or left running after its purpose ended, has no anomaly in it at all — its baseline is the waste. Finding those is a deliberate hunt through what exists, not a signal that arrives. 3. **A rise proportional to genuine growth.** If spend doubled because work doubled, the anomaly rule is right that spend moved and wrong that anything is wrong. Only a per-unit figure separates those two cases. So the three signals cover different ground: an **anomaly** catches a sudden change against the workload's own shape, a **budget threshold** catches the cumulative total regardless of how gradually it got there, and a **unit figure** catches a change that growth would otherwise excuse. A programme with only one of them has a blind spot you can name in advance. ## What the alert does and does not tell you When it fires correctly, the alert says: spend in this scope moved away from what this scope normally does, on this day. It does not say which dimension moved, what generated the charge, whether the cause is a design choice or an accident, or whose workload it was. Every one of those is a separate investigation with a separate owner. The anomaly rule's job is to make sure someone starts that investigation in days rather than at the invoice.

  • Why a median of comparable periods rather than a mean?
    A mean is pulled toward the outliers it is meant to detect. If last week contained the runaway you missed, it raises the expectation and helps hide the next one. A median of recent comparable periods barely moves for a single extreme value, so the baseline keeps describing normal rather than drifting toward the incident.
  • Your team triples capacity deliberately on a Tuesday. What should happen to the baseline?
    Re-anchor it as part of the change. Otherwise the new level reports as anomalous every day until enough comparable periods accumulate at that level — the rule behaving correctly and being useless, which is precisely how a team learns to mute it.
  • What kind of cost increase will an anomaly rule never report, however well tuned?
    A gradual one. The baseline is built from recent comparable periods, so a slow climb is absorbed into the expectation a little at a time and no single day stands out. That case belongs to the cumulative budget threshold and to a per-unit figure, not to anomaly detection.

saying these in an interview costs you the question

  • Uses one fixed daily number as the baseline for a seasonal workload
  • Raises the threshold until the weekly peak stops firing
  • Expects the old baseline to hold after a deliberate capacity change
  • Runs the rule on the account total rather than per dimension
  • Believes anomaly detection will catch a slow, steady climb
  • Includes the in-progress day when computing the baseline