In a time series, what is the difference between a point outlier and a level shift?
answer
- ask what happens after the surprise
- does the series come back?
- one bad point versus every point after
- Nile flow before and after 1899
- drop the point, or re-baseline
basics
~20 sA point outlier is one observation far from expectation, after which the series returns to its old level. A level shift moves the series to a new baseline that persists. The difference decides whether you drop a point or re-baseline.
solid answer
~50 sA point outlier (an additive outlier) is a one-off: a single timestamp is far from what the surrounding history predicts, and the very next observation is back to normal. A level shift is a changepoint in the mean: from some date onward the series settles around a different baseline and never returns. The classic textbook case is the Nile annual flow series, which drops to a visibly lower mean from around 1899, when construction of the Aswan dam began, and stays there. The distinction drives the treatment. An outlier you can winsorise, impute or exclude from fitting, because the underlying process did not change. A level shift you must not delete — you re-baseline: add a step term from the shift date, refit on post-shift data, or reset the detector, or every subsequent point keeps alarming against a mean the world no longer has.
go deeper
Be ready to name the two shapes and say what happens after each: the outlier's series returns to its old level, the level shift's does not. Say plainly that one is dropped or imputed and the other is re-baselined.
Explain the mechanics of the distinction: persistence rules over a window, why a lag or rolling feature carries an outlier forward for the window length, and how a step indicator lets a model absorb a changepoint instead of alarming forever.
Show you have operated this. Describe how you confirm a shift against a release or pipeline change before acting, how you keep a change log of known shifts, and why an adaptive baseline that quietly learns a broken level is worse than a noisy alert.
Own the policy: who is allowed to declare a changepoint, whether re-baselining is automatic or reviewed, and how historical series stay comparable across definition changes. Frame it as metric governance, not as a detector setting.
## The shapes an anomaly can take In ordered data "anomaly" is not one thing. What matters is what the series does *after* the surprising observation, and interviewers use this question to see whether a candidate has that reflex. **Point outlier (additive outlier).** A single observation sits far from what the surrounding history implies, and the next observation is back on the old path. A one-day traffic spike from a press mention, a sensor glitch, a duplicated batch load. The underlying process never changed; one measurement did. **Transient (decaying) outlier.** A shock hits and the series returns to its old level gradually rather than instantly — an outage day followed by a few days of recovery as the backlog clears. The signature is a run of same-signed residuals that shrink toward zero. **Level shift (changepoint in mean).** From some date onward the series oscillates around a new baseline and shows no tendency to come back. This is what "changepoint" usually means. The canonical example is the Nile annual flow record, which sits at one mean through the 1870s-1890s and at a distinctly lower mean from roughly 1899 onward, coinciding with the start of Aswan dam construction. One number changed permanently; the noise around it did not. **Drift / trend change.** No jump at all: the slope changes. The series climbs where it used to be flat. Detectors tuned for jumps often miss this entirely, because no single point and no single step is large. **Variance change.** The mean is unchanged but the series becomes noisier. Amplitude alone will not reveal it; the spread of residuals will. ## Why the distinction decides the treatment The treatments are close to opposite. For a point outlier, the right move is usually to keep it out of estimation: winsorise it, impute the seasonal expectation in its place, or fit a model that is robust to it. Leaving it in is not neutral — an extreme value distorts the mean and variance you compare future points against, and if your features include lags or rolling windows, the one bad value re-enters the model for as many periods as the longest window. That is why an outlier can hurt forecasts for weeks after the series itself has recovered. For a level shift, deleting the points is destructive: they are the new truth. The options are to add a step regressor (an indicator that is 0 before the changepoint and 1 after) so the model can absorb the jump, to refit on post-shift data only, or to reset the monitor's reference mean. Do none of these and a detector alarms every single day forever, because it keeps comparing correct data to a stale baseline. That is the most common production failure this question is really probing. ## How you tell them apart - **Persistence.** Require evidence over a window, not a point. One breach is an outlier candidate; a run of same-signed residuals of similar size is a level-shift candidate. A simple rule — flag a shift only after k of the next n points stay on the same side of the old mean — separates the two cheaply. - **Magnitude versus duration.** Point outliers are usually large and short. Level shifts can be small and forever, which is exactly why per-point thresholds are bad at them and cumulative methods are good at them. - **Shape of the recovery.** An instant return means point outlier; a decaying return means transient; no return means level shift. - **Corroboration.** A level shift almost always has a story — a release, a pricing change, a new data source, a definition change. If you cannot find one, suspect the pipeline before you believe the world changed. ## Common mistakes Treating every large residual as "an anomaly" and routing them all to the same handler. Deleting post-shift observations as bad data. Re-baselining after a single unusual day, which permanently bakes a spike into your reference. And silently letting an adaptive detector *learn* a broken level: if a logging bug halves your counts and the detector's rolling baseline adapts within a week, the alarm stops and the bug becomes invisible — the shift is real, the cause is not, and the fix is a change log of known shifts, not a quieter threshold.
- How would you distinguish a level shift from a slow drift?A level shift is a jump between two flat regimes: residuals against the old mean are all the same sign and roughly the same size from the changepoint onward. A drift has no jump — the deviation grows steadily, so the residual run is same-signed but increasing. Fit the two competing shapes (step versus slope) over the suspect window and compare fit; a drift also usually has no single candidate date attached to it.
- Why can a single point outlier damage forecasts long after the series recovers?Because it re-enters the model repeatedly. Any lag feature, rolling mean or smoothed level carries the bad value forward for the length of its window, and a seasonal baseline that looks back one full cycle will reuse it a whole cycle later. It also inflates the estimated variance, which widens intervals and desensitises the detector. Cleaning the value once, at ingestion, is cheaper than fighting its echoes.
- When is the right response to a confirmed level shift to change nothing?When the shift is a known, intended, one-time re-definition and downstream consumers already account for it — for example a metric deliberately re-scoped. Then you annotate the series and the model with the date, add a step term so history stays usable, and leave the raw data alone. What you never do is quietly delete or patch post-shift observations to make old dashboards line up.
A point outlier is a pothole: you swerve once and drive on. A level shift is the road being permanently rerouted: keep steering by the old map and every turn is wrong.
saying these in an interview costs you the question
- Calls any large residual an anomaly without asking whether it persists
- Deletes post-shift observations as bad data
- Assumes anomalies are always single points
- Re-baselines the detector after one unusual day
- Lets an adaptive baseline silently absorb a broken level