As reviewed labels accumulate from an anomaly-alert queue, should the task be re-posed as supervised classification?
answer
- who chose which items got reviewed?
- unreviewed is unlabelled, not negative
- the model distils the incumbent detector
- buy unbiased labels with random review
- exploration budget versus finds this quarter
basics
~20 sOnly with a correction for how those labels were produced. Reviewers judged only what the incumbent detector surfaced, so a classifier trained on them learns to imitate that detector and stays blind wherever it never looked.
solid answer
~50 sA year of reviewer dispositions looks like a supervised dataset but is not a sample of the population. Everything below the incumbent's alert cut was never reviewed, so those items have no label, and the labelled region is exactly the region the old detector considered suspicious. Train a classifier on it and you distil the incumbent plus reviewer habits: it will score beautifully on the alert stream and never surface a mode the incumbent misses. The way out is to buy unbiased data with review capacity — reserve a slice of the budget for randomly sampled items, so you get labels outside the alert region and can measure the supervised model on something other than its own training distribution. Then run both channels: supervised for the modes that have become closed-set and repeat, novelty for whatever is unlike anything seen. The real decision to own is how much scarce review capacity you spend on exploration versus immediate yield.
go deeper
Know that labels collected from an alert queue only describe the items someone chose to review, so they are not a random sample of everything the system saw.
Explain why a classifier trained on reviewed alerts reproduces the incumbent detector's coverage, and why holding out part of that same reviewed set flatters the new model.
Show how you would obtain unbiased labels in practice: a reserved random-review slice, logging the incumbent score and cut for every item, and evaluating candidates on the random sample rather than the queue.
Own the budget split between exploration and immediate yield, defend it to a team measured on finds, and state the risk posture that decides whether distilling the incumbent is acceptable at all.
## The situation A weld-inspection line has run a novelty-based flagging system for a year. Operators reviewed the flagged welds and recorded a disposition — genuine defect, cosmetic, false alarm. There are now thousands of labelled records. The obvious next step is a supervised classifier. The obvious next step is a trap unless you understand where the labels came from. ## Why the labelled set is not a sample of the population Three structural problems sit inside those dispositions. **1. Selective labelling.** Only items above the incumbent's alert cut were ever reviewed. Below it, nothing is labelled — not "negative", *unlabelled*. If you treat unreviewed items as negatives you are asserting the incumbent never misses, which is precisely the assumption you would be building a new model to test. **2. The labels encode the incumbent's notion of anomaly.** The positives in your training data are the intersection of "actually a defect" and "the old system flagged it". A defect mode the incumbent scores as ordinary appears zero times. A supervised model fitted here converges toward reproducing the incumbent, faster and perhaps cheaper, but with the same blind spots. That is a real product outcome — sometimes an acceptable one — but it must be a decision, not a surprise. **3. Reviewer behaviour is part of the label-generating process.** Dispositions drift with training, workload and what the team was recently told to prioritise. During a scrap-rate push, borderline welds get called defects; a quarter later they do not. The label is a joint product of the item and the review context. ## The seductive evaluation The trap closes at evaluation time. Hold out a slice of the reviewed dispositions and the new classifier scores extremely well, because train and test share the same selection. It is being graded only on items the incumbent already surfaced — the population where the incumbent is by construction good. Nothing in that measurement can detect the failure that matters: the defects nobody ever looked at. ## Buying unbiased data The fix is to spend review capacity on information rather than only on yield. - **Reserve a random-review slice.** Send a small random sample of *unflagged* items to review each week. These give labels drawn from the real population, which lets you estimate how much the incumbent misses and gives a supervised model examples outside the alert region. - **Evaluate on the random slice, not the queue.** The random sample is the only measurement that can compare a candidate model against the incumbent on equal terms. It is small, so the estimates are noisy and you must accept wide intervals rather than pretend precision. - **Record what was shown, not just what was decided.** Log the incumbent's score and the alert cut for every item, so that later analysis can reason about the selection rather than guess at it. - **Keep a permanent novelty channel.** Even after supervision takes over the known modes, something must be watching for items unlike anything in the training data — that channel is the only source of genuinely new positives, and it is also how the supervised model's future training data stays fresh. ## When re-posing as supervised is genuinely right Migrate the known part of the problem to supervision when the evidence supports it: - The defect or attack modes have become **closed-set and repeating**: the same handful of causes account for most confirmed cases quarter after quarter. - You hold enough confirmed positives per mode that a model can learn the mode itself rather than just distance from normal. - You have some unbiased data — the random slice — with which to check that the supervised model is not merely mimicking the incumbent. - The operational win is concrete: supervision names the mode, which routes the fix to the right process step, whereas a novelty score only says "strange". And keep the framing hybrid. A supervised model for known modes plus a novelty channel for the rest is not fence-sitting; it is the design the problem's structure implies, because the two channels answer different questions. ## The organisational tradeoff to own The scarce resource is reviewer time, and it has two competing uses. Spent on top-ranked alerts it maximises confirmed finds this quarter. Spent on random samples it produces almost no finds and buys the only unbiased measurement you will ever have. A leader has to choose a split, defend it to a team judged on finds per week, and revisit it — a larger exploration slice early while the system is young, a smaller maintenance slice once the mode mix is understood. The second call to own is the risk posture. Distilling the incumbent into a faster classifier is a legitimate efficiency project if the mode mix is stable and the cost of a missed novel mode is low. On an inspection line where an unseen failure mode ships defective parts to a customer, it is not. Say which world you are in, and let that drive the framing rather than the availability of a convenient pile of labels.
- How large should the random-review slice be?Large enough to estimate the miss rate with a usable interval, and no larger, because every random review is a find the team did not make. In practice I would start from the smallest sample that can detect a miss rate the business would care about, run it for a fixed period, and report the interval honestly rather than a point estimate. If the confirmed base rate is extremely low, I would say up front that a plain random sample may be too weak and stratify it.
- The team argues that random reviews waste investigator time. How do you answer?I agree that it costs finds, and I put a number on it, because the argument is legitimate. Then I make the counter-cost visible: without unbiased data we cannot say what we are missing, we cannot compare a candidate model to the incumbent fairly, and we will not notice a new failure mode until it reaches a customer. It is an insurance premium paid in reviewer-days, and it should be sized and reviewed like one.
- Would you ever treat unreviewed items as negatives?Only as an explicit, documented approximation, and only when the incumbent's miss rate is known to be low from independent evidence such as a random audit. Silently labelling everything unreviewed as clean encodes the assumption that the incumbent never misses, which biases the new model toward the old one and hides the exact errors you are trying to find.
- What signals tell you the problem has become closed-set enough for supervision to lead?The same small set of confirmed causes dominating quarter after quarter, new modes appearing rarely and being absorbed quickly, enough confirmed cases per mode for a model to learn the mode rather than mere strangeness, and a random-sample check showing the supervised candidate beats the incumbent outside the alert region rather than only inside it.
saying these in an interview costs you the question
- Treats unreviewed items as confirmed negatives
- Evaluates the new model only on the alert queue
- Claims a year of dispositions is a representative sample
- Drops the novelty channel once supervision performs well
- Ignores that reviewer standards drift over time