skip to content

What does a CUSUM chart detect that a per-point threshold rule misses?

level: middleimportance: nice to knowfreq 30%

answer

  1. evidence accumulates, verdicts do not
  2. small shift, every point still inside limits
  3. running sum with a floor at zero
  4. slack k tolerates, interval h alarms
  5. k is half the shift you target

basics

~20 s

CUSUM accumulates signed deviations from the target mean instead of judging each point alone, so it catches a sustained small shift that never breaches a per-point limit. Slack k targets the shift size; threshold h triggers the alarm.

solid answer

~60 s

A per-point rule asks "is this observation extreme?" and a 0.5-sigma shift in the mean never is — almost every point still lands inside the limits, so a permanent degradation can run for months unnoticed. CUSUM asks a different question: is the evidence *accumulating*? It keeps a running sum of deviations from the target mean, with a slack `k` subtracted each step, and alarms when the sum exceeds a decision interval `h`. In standardised units the upper side is `S_t = max(0, S_{t-1} + (x_t - mu0)/sigma - k)`, alarm when `S_t > h`; the lower side mirrors it. The reset at zero is what makes it work: noise cannot push the statistic up, only a persistent bias can. You tune `k` to about half the shift you care about in sigma units and `h` to buy the false-alarm rate you can afford — `k = 0.5, h = 5` is the standard starting point. The price is speed on large spikes, where a per-point rule wins, so most monitors run both.

code

python · 17 lines
python
import random, statistics

def run_length(k, h, shift, seed):
    rng = random.Random(seed)
    s, n = 0.0, 0
    while True:
        n += 1
        z = rng.gauss(shift, 1.0)   # standardised observation
        s = max(0.0, s + z - k)     # upper one-sided CUSUM
        if s > h:
            return n

k, h = 0.5, 5.0
quiet = [run_length(k, h, 0.0, i) for i in range(500)]
shifted = [run_length(k, h, 1.0, 1000 + i) for i in range(500)]
print("mean points between false alarms:", round(statistics.mean(quiet)))
print("mean points to detect a 1-sigma shift:", round(statistics.mean(shifted)))

go deeper

for a junior

Be ready to say that some changes are too small to trip a per-point limit yet persist, and that accumulating deviations over time is how such a change gets caught.

for a middle

Explain the recursion itself: deviations from a target mean, slack subtracted each step, the floor at zero, and the alarm when the running sum passes the decision interval. Know that k is set to about half the shift you target.

for a senior

Show you have tuned one in production: choosing k and h from the tolerable false-alarm rate, estimating the reference mean and spread from a clean window, deciding the post-alarm reset policy, and feeding residuals rather than raw seasonal values.

for a principal

Own the choice between cumulative and per-point monitoring across a fleet: which failure modes each covers, whether one detector family is worth standardising on, and how much detection delay the organisation is buying with each notch of quiet.

## The blind spot in per-point rules A per-point rule compares each observation to a fixed band and treats the observations as independent verdicts. It is excellent at catching a big, obvious excursion — one point, way out, alarm. It is close to useless against a *small but permanent* change in the mean. Work the numbers. Suppose the process mean shifts by 0.5 standard deviations and stays shifted. Points are now centred at 0.5 sigma instead of 0. A three-sigma rule fires when a point exceeds 3 sigma, which for the shifted process means exceeding 2.5 sigma above its own new mean — an event with probability around 0.6% per point. You will wait on the order of 150 observations to see one, and the one you do see looks like an isolated spike rather than a regime change. Meanwhile every one of those 150 points carried a little evidence that the mean had moved, and the per-point rule threw all of it away. That discarded evidence is what CUSUM collects. ## The recursion Standardise the observation, `z_t = (x_t - mu0)/sigma`, where `mu0` is the in-control target mean and `sigma` the in-control standard deviation, both estimated from a clean reference period. Then run two one-sided statistics: ``` S_hi(t) = max(0, S_hi(t-1) + z_t - k) S_lo(t) = max(0, S_lo(t-1) - z_t - k) ``` both starting at zero, alarming when either exceeds the decision interval `h`. Three things are doing the work. **The sum.** Persistent bias adds up; independent noise does not. Under a real shift of size `delta` sigma, each step adds about `delta - k` on average, so the statistic climbs linearly and hits `h` in roughly `h / (delta - k)` observations. **The slack `k`.** Subtracting `k` every step is what stops the statistic drifting up on noise alone. In control the increment averages `-k`, so the sum is pushed down toward the floor. `k` is the size of shift you are willing to tolerate, and the standard choice is half the shift you want to detect: to detect a 1-sigma shift, use `k = 0.5`. **The reset at zero.** Without it, an old quiet period of below-target values would bank credit and delay detection of a later rise. Flooring at zero means the statistic only remembers the current stretch of evidence in one direction, which is why CUSUM responds to a shift that starts today rather than to the whole history. ## Tuning, and what the tuning buys The performance currency is **average run length (ARL)**: the mean number of observations before an alarm. Two versions matter — ARL when nothing is wrong, which is the mean gap between false alarms and should be large, and ARL under the shift you care about, which is the detection delay and should be small. With `k = 0.5`, the textbook pairing is `h = 5`: a two-sided chart then averages roughly 465 observations between false alarms while detecting a 1-sigma shift in about ten. Dropping to `h = 4` alarms sooner but falsely far more often. Raising `h` always buys quieter operation at the cost of slower detection — there is no setting that improves both. Two practical notes. First, `mu0` and `sigma` are *estimates*; a biased reference period silently biases every subsequent decision, so re-estimate after a confirmed level shift and never estimate them from a window containing the incident. Second, decide what happens after an alarm: reset the statistic to zero and continue, or hold it and stop the process. Resetting without fixing the cause produces a periodic re-alarm, which on-call reads as flapping. ## Where it sits among alternatives - **Per-point limits** are fast on large spikes and useless on small drifts. - **CUSUM** is the sharpest tool for a shift of a *known* size, because `k` is tuned to that size — and correspondingly less efficient for shifts much larger or smaller. - **EWMA** — an exponentially weighted moving average compared to its own limits — behaves similarly, with a smoothing constant playing the role `k` plays, and degrades more gracefully across a range of shift sizes. - **Retrospective changepoint estimation** answers a different question: given a whole history, *where* did the mean change? CUSUM is sequential and answers *has it changed by now*, which is what monitoring needs. Running a per-point rule and a CUSUM together is common and sensible: the first catches the outage, the second catches the slow bleed. ## Assumptions worth stating out loud CUSUM's run-length properties assume roughly independent observations around a stable mean with stable spread. Real metrics are autocorrelated and seasonal, so you do not feed raw values in — you feed residuals after the predictable structure is removed. Feed a seasonal series to a raw CUSUM and it will alarm on the season, every cycle, which is a detector reporting on the calendar rather than on the process.

  • What is the role of the slack parameter k, and how do you choose it?
    `k` is subtracted from every increment, so in control the statistic drifts down to its floor at zero and noise cannot accumulate into an alarm. It encodes the shift size you are willing to ignore. The convention is `k = delta/2` in sigma units, where `delta` is the shift you want to catch: `k = 0.5` targets a one-sigma shift. Too small and you alarm on noise; too large and a genuine small shift never accumulates.
  • Why is the CUSUM statistic floored at zero?
    Without the floor, a long quiet or below-target stretch would bank negative credit that a later genuine rise must first pay off, delaying detection by however long the quiet period ran. Resetting to zero makes the statistic depend only on the current run of one-directional evidence, so a shift beginning today is detected on today's evidence rather than on the whole history.
  • What breaks if you feed CUSUM a raw seasonal metric?
    Its run-length behaviour assumes roughly independent observations around a stable mean. A seasonal series spends every cycle systematically above and then below the target, so the statistic accumulates on schedule and alarms once per cycle — reporting the calendar, not the process. Feed residuals after the predictable structure has been removed, and re-estimate the reference mean and spread after any confirmed level shift.
  • When would you prefer a per-point rule over CUSUM?
    When the failure mode you fear is a large, abrupt excursion — an outage, a dropped feed, a runaway value. A per-point rule fires on the first offending observation, while CUSUM still needs the sum to climb past `h`, costing you a few points of delay. The usual production answer is to run both: a fast rule for catastrophes and a cumulative one for slow degradation.

A per-point rule is a bouncer checking each guest for a weapon. CUSUM is an accountant noticing the till is five pounds short every night: no single evening looks criminal, the running total does.

saying these in an interview costs you the question

  • Thinks CUSUM is just a rolling average with a threshold
  • Confuses the slack k with the alarm threshold h
  • Omits the reset to zero, letting old evidence bank credit
  • Feeds raw seasonal values straight into the chart
  • Claims a threshold can improve false alarms and delay at once

context