How do you score anomalies in a strongly seasonal daily metric without flagging every Monday?
answer
- the calendar is signal, not noise
- compare like with like
- same weekday, one cycle back
- threshold from residual spread, not raw noise
- known peaks need an event calendar
basics
~20 sScore the residual against a seasonal baseline instead of the raw value: compare each day with the same weekday one cycle earlier, then flag unusually large residuals. A Monday is only anomalous relative to other Mondays.
solid answer
~50 sThe problem is that the raw level is dominated by seasonality, so any fixed band flags whichever weekday is systematically highest or lowest. The fix is to score residuals, not levels. The cheap, strong baseline is seasonal-naive: predict `y_t` with `y_{t-m}`, where `m` is the seasonal period — 7 for day-of-week — and score `r_t = y_t - y_{t-m}`. Then estimate the spread of the residual series itself and flag points far out in that residual distribution. Two things you must handle. Known calendar events need their own treatment: a recurring peak such as Black Friday is not an anomaly, so give it an event flag or exclude it from scoring, otherwise you guarantee an annual false alarm. And the threshold has to come from the residual spread, not the raw noise: differencing two noisy observations roughly doubles the variance, so `Var(y_t - y_{t-m}) = 2*sigma^2` when the seasonal pattern is exact and the errors are independent.
go deeper
Be ready to say why a fixed band on raw daily data flags whole weekdays, and that the fix is comparing each day with the same day of the previous cycle rather than with the overall average.
Explain the mechanics: the seasonal-naive baseline y_{t-m}, why differencing two observations roughly doubles the residual variance, and why the alarm threshold must be estimated from the residual series itself.
Show operational judgment: an event calendar for known peaks, per-period spread estimates, log-scale scoring for multiplicative seasonality, and persistence rules so a single breached residual does not page anyone.
Own the tradeoff between a simple, explainable baseline the on-call team trusts and a richer model that scores better but is harder to debug at 3am. Decide who maintains the event calendar and how detector changes are reviewed.
## Why raw thresholds fail on seasonal data Put a fixed band around a daily orders series and it will be busy every week. Suppose weekend volume runs a third below weekday volume. Any band tight enough to catch a genuine Tuesday collapse is breached by every Saturday, and any band loose enough to accept Saturday cannot see a Tuesday problem until it is catastrophic. The metric's biggest source of variation is not noise, it is the calendar — and the calendar is entirely predictable, so it belongs in the baseline rather than in the alarm. The reframing: **an observation is anomalous relative to what the season predicts for it, not relative to the series average.** Score residuals. ## Seasonal-naive residual scoring The simplest baseline that respects the calendar is the seasonal-naive forecast: the prediction for time `t` is the observed value one full cycle earlier, `y_{t-m}`. For day-of-week effects on a daily series, `m = 7`; for an annual pattern on daily data, `m = 365`. The anomaly score is the residual ``` r_t = y_t - y_{t-m} ``` usually standardised by a spread estimate computed over the residual series itself, so the score is in comparable units across metrics. What this buys you: the day-of-week shape is removed without fitting anything, the level is removed, and a linear trend is reduced to a constant offset (`m` times the per-step slope), which you can subtract as the residual median. It needs one full cycle of history to start, and nothing else. When the seasonal shape is stable but noisy, residuals against a smoothed seasonal profile — the residual component of a decomposition — score better than seasonal-naive, because they do not carry a second observation's noise. That is the trade in the next section. ## Set the threshold from the residual spread A classic mistake is to take the per-observation noise level `sigma` and place the alarm at `3*sigma`. Seasonal-naive residuals are a difference of two observations. If the seasonal component is exact and the errors are independent with variance `sigma^2`, then ``` Var(y_t - y_{t-m}) = sigma^2 + sigma^2 = 2*sigma^2 ``` so the residual standard deviation is about `1.41*sigma`, and a threshold set from raw noise over-flags by a wide margin. Always estimate the spread from the residuals you will actually score, and estimate it robustly, over a window long enough to cover the pattern but recent enough to reflect the current regime. ## Calendar events and multiple seasonalities Day-of-week is rarely the only cycle. Retail and consumer metrics carry a weekly cycle, an annual cycle and a set of moving holidays that do not repeat on the same date. A `m = 7` baseline handles Monday and leaves every holiday exposed; a `m = 365` baseline handles fixed-date holidays and misaligns the weekday pattern, since 365 is not a multiple of 7. Practical handling: - Maintain an **event calendar** — recurring peaks and troughs, promotions, releases, known outages — and either suppress scoring on those dates or give the baseline an event term. A predictable annual peak that alarms every year is not detection, it is a scheduled interruption. - **Score within season** when the spread differs by period: overnight hours are quiet and tight, midday is loud and wide, so a single global threshold is simultaneously too tight and too loose. Either fit a per-period spread or work on the log scale when seasonality is multiplicative (amplitude proportional to level). - Beware the **echo**: with seasonal-naive, an anomalous value becomes the baseline exactly one cycle later. A one-day outage flags on its own day and flags again a cycle later with the opposite sign, when a perfectly normal day is compared against the outage. Cleaning confirmed anomalies out of history before they are reused as a baseline removes the echo; a smoothed seasonal profile is less vulnerable to it. ## Direction, persistence and what you do with a flag Most operational metrics only care about one side — you page on orders collapsing, not on orders being unusually good — so use a one-sided rule and keep the other side as a data-quality signal. Require persistence before acting: a single breached residual is a candidate, several consecutive same-signed breaches is an incident. And separate the two outcomes this scoring produces: an isolated large residual is a point anomaly, while a run of same-signed residuals of similar magnitude is the beginning of a level shift, which needs re-baselining rather than an alert every day forever. ## Interview-ready summary Remove the predictable part, score what is left, set the threshold from the spread of that residual series, keep an event calendar for known peaks, and require persistence before you wake anyone. If a candidate proposes a fixed band on raw seasonal data, everything else they say about detection is downstream of a broken premise.
- Why does a seasonal-naive residual have larger variance than the underlying noise?Because it differences two noisy observations rather than comparing one observation to a smooth estimate. With independent errors of variance `sigma^2` on both `y_t` and `y_{t-m}`, the residual variance is `2*sigma^2`, so its standard deviation is about 1.41 times the per-observation noise. Thresholds derived from raw noise therefore over-flag; derive them from the residual series you actually score.
- Your metric has both a weekly and an annual pattern. What breaks with a single seasonal period?A weekly baseline treats every holiday as an anomaly, and an annual daily baseline misaligns weekdays because 365 is not a multiple of 7, so last year's same date is a different day of the week. You need both cycles handled — a baseline carrying weekly and annual terms plus an explicit calendar of moving holidays — or you accept a predictable stream of false alarms.
- How would you handle a metric whose seasonal amplitude grows as the level grows?That is multiplicative seasonality: peaks scale with the level, so an additive residual is small early in the series and large later, and a single threshold drifts out of calibration. Score on the log scale, which turns the multiplicative pattern into an additive one and makes residual spread roughly constant, or score relative residuals as a percentage of the baseline rather than absolute differences.
Judging a shop by raw daily takings is like judging a city's traffic by counting cars at 3pm and at 3am with the same rule. You compare a Monday with other Mondays, not with the weekly average.
saying these in an interview costs you the question
- Puts a fixed band on the raw seasonal level
- Sets the threshold from raw noise, not residual spread
- Treats a recurring holiday peak as an anomaly every year
- Uses one global spread across quiet and busy periods
- Lets an uncleaned anomaly become next cycle's baseline