Why does a fixed 500-downloads-a-day rule miss a user whose SaaS downloads tripled?
answer
- one number is set by the loudest legitimate user
- headroom under the threshold is where the drip lives
- compare each entity with itself
- ratios over small counts are noise
- a rolling baseline can absorb a slow ramp
basics
~20 sA single global threshold has to sit above the loudest legitimate account, so it sits far above a typical one. A user going from six downloads a day to forty has tripled and is still nowhere near the number. Comparing each account to its own history is what makes the change visible.
solid answer
~50 sOne number for the whole estate is set by its noisiest legitimate member. If a reporting account genuinely pulls six hundred documents a day, the threshold has to clear that, and every ordinary user then has hundreds of downloads of headroom to hide in. Someone drip-collecting at forty a day never trips it, and never will, because they can measure the boundary as easily as you can. The fix is a per-entity comparison: derive each account's own distribution from real history — say a median and a high percentile over sixty to ninety days — and alert on a departure from that. It costs you state per entity, an explosion of keys, a cold start for every new account, and noise on low-volume users where a jump from one to four is meaningless. So pair the ratio with an absolute floor, and exclude entities with too little history rather than alerting on them.
code
text · 9 linesdownloads per user per day, 30-day summary
user median p95 yesterday
reporting-svc 420 610 585
m.ferrand 6 14 41
p.oyelaran 3 9 38
k.svensson 8 17 11
...
rule in force: alert when downloads_per_day > 500go deeper
Be ready to say why one number for everyone is set by the heaviest legitimate account, and why that leaves ordinary users a large gap to operate in unnoticed.
Explain the derivation: per-entity distribution over real history, a backtest that predicts alert volume, and a percentile chosen for the volume it produces rather than picked for looking round.
Demonstrate that you know what per-entity costs — state, cardinality, cold start, noisy small counts — and that you would combine a relative departure with an absolute floor and a qualifying condition before letting it page.
Own the argument that a threshold is an artefact with an owner, a derivation and a review date, and that a rule producing volume with no escalations is failing regardless of how carefully it was derived.
## Why one number cannot work A count-over-window rule fires when some entity exceeds a count within a period — downloads per user per day, failed authentications per account per hour, records read per service account per week. The tempting implementation is one threshold for everyone. The trouble is that the threshold must be high enough not to fire on legitimate heavy users, and in every real estate the heaviest legitimate user is one or two orders of magnitude above the median. A backup or reporting identity pulling hundreds of documents daily forces the number up, and once it is up, the entire population of ordinary users has enormous unmonitored headroom underneath it. That headroom is the whole attack surface of the rule. Slow collection through a sanctioned client is specifically shaped to live there: the account is legitimate, the client is approved, the destination is the normal one, and the only anomalous thing is the *rate relative to that account*. Against a global threshold there is nothing to detect. Worse, a global threshold is a published boundary — anyone who can observe whether alerts fire can find the edge and stay below it. ## Deriving a threshold instead of choosing one "Tuned" means the number came from data, not from a round figure someone liked. The method is: 1. Take a representative slice of history — long enough to include the estate's cycles, and free of known incident activity. 2. Compute the per-entity distribution of the count you intend to alert on: median, p95, p99, maximum. 3. Pick the candidate threshold and *backtest* it: how many alerts would this rule have produced over that period, and on which entities? 4. Sample the would-be alerts and label them. The output you care about is the daily volume an analyst would inherit and the share of it that is worth working. 5. Record the number with its derivation and a review date, because the estate drifts. A threshold nobody has backtested is a guess, and the first time you learn its volume is when it lands in the queue. ## What a per-entity baseline buys, and what it costs Per-entity means each account is compared with itself: alert when today's count exceeds that account's own p95, or exceeds its median by some multiple. It makes proportional change visible regardless of the account's absolute scale, which is exactly what a global number destroys. The costs are real and you should name them: - **State.** You must store and maintain a profile per entity, refresh it on a schedule, and decide how far back it looks. That is a data-management job, not a rule. - **Cardinality.** Tens of thousands of entities multiplied by several fields is a lot of keys, and the cost of computing the baseline grows with them. - **Cold start.** Every new account starts with no profile. The correct behaviour is to *withhold* the rule for that entity until it has a minimum history, not to alert on it. - **Small numbers.** A user whose median is one download a day trips a 4x ratio by opening four documents. Ratios over small counts are noise, so combine the ratio with an absolute floor: alert only when the count is both anomalous for the entity and above a level that could matter. - **Self-poisoning.** If the baseline is a rolling window and the adversary ramps slowly, their activity is inside the window and becomes the new normal. Guard this with a long-horizon cumulative check — total volume over ninety days against the same period last quarter — which a slow ramp cannot escape by being gradual. ## The shape that actually works In practice the deployable rule is a conjunction rather than a single test: an entity-relative departure, plus an absolute floor, plus at least one qualifying condition that raises the prior — an unfamiliar client, a session from an address the account has not used, activity outside the account's usual hours, or a departing employee flag. Each term is individually weak; together they produce a volume an analyst can actually work. And measure the rule by what it yields, not by whether it exists. If it has produced two hundred alerts and no escalations, the threshold is not tuned, whatever process produced it. ## The interview answer Say plainly that a global threshold is calibrated by the loudest legitimate user and therefore blind to proportional change in everyone else; that the fix is per-entity comparison derived from backtested history; and then, unprompted, name the costs — state, cardinality, cold start, noisy small numbers and a baseline the adversary can walk into. Candidates who only give the first half sound like they have read about baselining; candidates who give both halves sound like they have run one.
- What stops the adversary's own activity from being absorbed into the per-user baseline?Nothing, if the baseline is a short rolling window and the increase is gradual — the ramp becomes the new normal within a few refreshes. Counter it with a long-horizon cumulative check that a slow ramp cannot escape, such as ninety-day total volume against the same period a quarter earlier, and by rebuilding baselines from a period you have reason to believe is clean.
- How much history do you require before a per-user threshold is allowed to alert on that user?Enough to cover the longest legitimate cycle you care about and enough observations for the statistic to mean anything. Below that, exclude the entity rather than alert on it — an account with four days of history will breach any percentile you compute from it. Say explicitly what the minimum is, so the exclusion is a policy rather than an accident.
- How do you know the number you derived is the right one?Backtest it: replay the rule over historical data, count the alerts it would have produced, sample and label them. You are choosing a workable daily volume and an acceptable share of it being benign. A threshold that has never been replayed against real history is a guess, and the queue is where you will discover that.
saying these in an interview costs you the question
- Picks a round number and calls the rule tuned
- Claims per-entity baselines are free of operational cost
- Applies a ratio to users whose normal count is one or two
- Never backtests the threshold against historical data before shipping
- Assumes the adversary's activity cannot be inside the baseline window