skip to content

A promo-abuse detector's positive rate triples during a January sale. What happens at its fixed threshold?

level: seniorimportance: should knowfreq 42%

answer

  1. separate more positives from different positives
  2. within-class rates ignore the prior
  3. precision moves with the mix, recall does not
  4. the queue breaks before the metric does
  5. choose which invariant you hold fixed

basics

~20 s

Flagged volume rises sharply. If only the base rate moved, precision at the fixed cut rises while recall holds, so the queue floods rather than degrades. If the traffic itself changed shape, false positives rise and precision can fall.

solid answer

~50 s

Split it into two cases. Under a pure base-rate rise -- the same kinds of abusers, just more of them -- the per-class score distributions are unchanged, so TPR and FPR at the fixed cut are unchanged, recall is unchanged, and precision actually *rises* because a larger share of flagged items are genuine. What breaks is volume: three times the abuse means far more items above the cut than the review team can clear. Under covariate shift -- new promo mechanics, different traffic -- the score distributions themselves move, false positives can rise, and precision can fall. Distinguish them by watching flag rate, score distribution and audited precision together. Then decide which invariant you hold -- the score cut, the daily volume, the precision or the expected cost -- because you cannot hold all four; re-derive the cut on recent labels and alert on the flag rate.

go deeper

for a junior

Be ready to notice that a fixed cut lets the number of flagged items float with traffic, so a busy period produces far more alerts even though nothing about the model changed.

for a middle

Explain the arithmetic: recall and false-positive rate are within-class quantities and do not move with the class mix, while precision does, so a higher positive rate at a fixed cut means a cleaner flagged set and a bigger one.

for a senior

Show the diagnosis. Separate a pure base-rate rise from a change in the traffic itself using flag rate, score distribution and an audit of flagged items, then re-derive the cut on recent labels rather than trusting a seasonal number.

for a principal

Decide and document which invariant the organisation holds -- the cut, the volume, the precision or the expected cost -- since they conflict, and put the monitoring and the review cadence in place before the season that breaks the stale one.

## The scenario A promo-abuse detector was tuned in December on December traffic, and a threshold was fixed on its score. A January sale arrives and the share of sessions that are genuinely abusive triples. Nobody retrains, nobody moves the cut. What happens? The answer people give reflexively -- "precision drops because there is more noise" -- is usually backwards, and getting the direction right is what separates a candidate who has operated a detector from one who has read about one. ## Case one: pure base-rate shift Suppose the abusers behave exactly as they did in December and the honest users do too; there are simply more abusers in the mix. Formally, the conditional score distributions `P(score | abuse)` and `P(score | legitimate)` are unchanged and only the prior `pi = P(abuse)` has risen. Consequences at a fixed score cut: - **TPR and FPR are unchanged.** Both are computed *within* a class, and neither class's score distribution moved. - **Recall is unchanged**, since recall is TPR. - **Precision rises.** Precision is `pi*TPR / (pi*TPR + (1-pi)*FPR)`. Increasing `pi` increases the numerator's share, so the flagged set gets *cleaner*. - **Volume rises, a lot.** Both true and false positives grow in absolute count -- true positives because there are more positives, false positives because sale traffic is heavier overall. - **The model's probabilities are now wrong.** It learned a December prior, so its outputs understate January risk. A cut expressed as "flag if the probability exceeds 0.04" no longer means what it meant, even though a cut expressed as "flag the top 500 scores" still does. So the failure mode is operational, not statistical: the review queue floods, backlog grows, and effective response time collapses -- while every offline metric on the flagged sample looks fine or better. ## Case two: covariate shift A sale rarely brings *only* more abusers. It brings different traffic: new users, bulk purchases, coupon stacking that is legitimate this month, referral patterns that were rare in December. Now `P(score | legitimate)` shifts upward -- ordinary behaviour starts scoring like abuse -- and FPR at the fixed cut rises. Precision can fall, sometimes sharply, and recall may move in either direction. ## Telling the two apart You do not need labels to start: - **Flag rate up, score distribution shape roughly unchanged, audited precision up** -- prior shift. The cut still separates as well as it did; there is just more of everything. - **Flag rate up, score distribution visibly shifted upward for the bulk of traffic, audited precision down** -- covariate shift. Ordinary sessions are now scoring high, and the model is finding a new kind of pattern. Sampling flagged items for human review gives realised precision within days. Recall is harder because it needs labels on unflagged items; a small random audit outside the flagged set, or complaints and clawback signals, are the usual proxies. ## Which invariant do you hold? This is the senior point. Four things people implicitly want to hold constant, and they conflict: 1. **The score cut.** Simple, auditable, reproducible -- and it lets volume float, which is what floods a queue. 2. **The daily volume.** Take the top N scores each day, where N is what the team can clear. Self-correcting under any kind of drift and trivially capacity-safe -- but the precision and recall you deliver now float silently, and a genuinely quiet day means you review 400 mostly-innocent sessions because you promised to review 400. 3. **The delivered precision.** Re-derive the cut from recent labelled data so the flagged set stays at the promised quality. Needs a steady label supply and lags by however long labelling takes. 4. **The expected cost.** Recompute the break-even from current costs and current probabilities. The most principled and the most demanding, since the probabilities must be trustworthy under the new mix. There is no universally right answer. A capacity-bound human queue usually wants (2) with (3) as a monitored guard-rail: take the top N, but alarm when audited precision of that top N falls below the floor, because that means you are burning reviewer time on innocents. ## What to put in place beforehand - **Monitor the flag rate as a first-class alert**, not a dashboard nobody opens. A tripling shows up within hours and needs no labels. - **Monitor the score distribution**, at a few percentiles, for the same reason. - **Audit a sample of flagged items** continuously so realised precision is a number you have, not a number you estimate after an incident. - **Write down the threshold's provenance** -- the data it came from, the criterion, the date. A cut with no recorded origin is one nobody will dare to move. - **Expect seasonality.** A cut fitted in December on December traffic is a seasonal artefact. Fitting on a period that spans the variation, or re-deriving on a schedule, beats discovering it during the sale. ## The concise interview answer "If only the base rate moved, recall and FPR hold, precision improves and volume explodes -- so the pain is capacity, not accuracy. If the traffic changed shape as well, precision can fall and I would tell them apart with the flag rate, the score distribution and an audit of flagged items. Then I would decide explicitly which invariant we hold -- cut, volume, precision or cost -- rather than letting a stale December number decide it for us."

  • Why does precision rise under a pure base-rate increase when recall does not move?
    Recall is the true-positive rate, computed only among real positives, so it cannot see the class mix at all. Precision looks at the flagged set, which contains both classes, and its composition follows the prior: with three times as many abusers scoring the same way, a larger share of everything above the cut is genuine abuse. Same detector, cleaner flagged set, unchanged catch rate.
  • What does pegging the cut to a fixed daily review capacity give up?
    Control of quality. Taking the top 400 scores each day is capacity-safe under any drift, but the precision and recall you deliver then float with the traffic: on a quiet day you review hundreds of innocent sessions to fill the quota, and on a bad day you miss abuse that sat at rank 401. It needs a precision floor monitored alongside it, with the quota shrinking when audited precision drops.
  • Which monitoring signal warns you first, before any labels arrive?
    The flag rate, followed by the score distribution's upper percentiles. Both are computable the moment predictions are made and need no ground truth, so a tripling is visible within hours while labels may take days or weeks. Alert on a deviation from a rolling baseline rather than an absolute number, and treat a sharp move as a prompt to check whether the mix or the traffic itself changed.

It is a metal detector at an airport: doubling the number of travellers carrying keys does not change the detector's sensitivity, it just triples the queue at the search table.

saying these in an interview costs you the question

  • Says precision must fall when the positive rate rises
  • Claims recall changes when only the class prior moves
  • Ignores that flagged volume can exceed review capacity
  • Treats a threshold fitted in one season as permanent
  • Cannot distinguish prior shift from a change in traffic
  • Monitors only offline metrics and never the flag rate

context