How do you decide an unsupervised anomaly detector is good enough to send alerts to analysts?
answer
- no labels means no accuracy number
- judge only the top of the ranking
- expert review of the top k
- stability across seeds and subsamples
- recall needs replayed incidents or below-cut sampling
basics
~20 sJudge the top of the ranking, not the whole model. Have experts review the highest-scored items, measure precision at that k, check the ranking is stable across refits, and accept that recall stays unknown until labels accumulate.
solid answer
~50 sWith no labels there is no accuracy to report, so the decision rests on three things. First, **precision at k**: have an expert adjudicate the top 100 scored rows and record what fraction were genuine - that is the only quality number you can measure honestly on day one. Second, **stability**: refit with different seeds and subsamples and check that the top of the ranking barely moves; a ranking that reshuffles run to run is not something to act on. Third, **explainability and volume**: an analyst needs to see which features drove each score to adjudicate it at all, and the alert volume must fit what the team can work through. Recall stays unmeasured, so estimate it separately - replay known past incidents, or review a small random sample from below the cut - and treat the deployment as the start of label collection rather than the end of evaluation.
go deeper
Remember that without labels there is no accuracy to report, and that the practical measure is how many of the highest-scored items an expert confirms as genuine.
Be able to define precision at k, explain why recall cannot be computed from a flagged-only review, and name the substitutes: replaying known incidents and sampling below the threshold.
Demonstrate the operating loop - blind adjudication, reviewer agreement, stability checks across seeds, alert volume against capacity, and capturing every review outcome as a label for later evaluation.
Own the launch gate and the economics behind it: what precision floor justifies analyst time, how reviewer trust is spent by a noisy launch, and how the review workflow is designed so that months of adjudication compound into an asset.
## Why the usual evaluation vocabulary does not apply Accuracy, a confusion matrix, a hold-out score - all of them need a ground-truth column that does not exist here. You have a scored list and a domain expert whose time is expensive. Every honest evaluation of an unsupervised detector is built out of exactly those two ingredients, so the question becomes: what can you learn from a bounded amount of expert attention, and what can you never learn from it? ## Precision at k: the one number you can get on day one Take the top `k` scored rows, where `k` is what a reviewer can realistically get through - a hundred is a common unit of work - and have them adjudicated. **Precision at k** is the fraction confirmed as genuine anomalies. It matches how the system will actually be used: nobody consumes the whole ranking, they work down from the top until the queue is empty. What it is good for: comparing two candidate detectors on the same data with the same reviewer, tracking degradation over time, and setting an expectation with the team that consumes the alerts. What it hides: - **Recall is invisible.** Precision at k says nothing about the anomalies sitting below the cut, and there is no way to compute recall without labels for the rest of the data. - **The reviewer defines the truth.** Two experts may disagree about the same row, and a reviewer who sees the score before deciding will drift toward confirming it. Adjudicate blind where possible, and measure agreement between reviewers on a shared subset before trusting the number. - **Interesting is not the same as actionable.** A detector often surfaces rows that are genuinely unusual and completely useless - a data-entry artefact, a test account, a known batch job. Score the review on "is this worth acting on", not "is this statistically odd", or you will optimise the wrong thing. ## Getting a handle on what you are missing Since recall cannot be measured, approximate it: - **Replay known incidents.** Any confirmed past events give you a handful of true positives. Score the historical data and check where those rows land in the ranking. A detector that buries a known incident at rank 40,000 is disqualified regardless of its precision at 100. - **Inject synthetic anomalies.** Perturb real rows in ways a domain expert agrees are realistic, and see how high they rank. The caveat matters and should be stated aloud: you only measure sensitivity to the kinds of anomaly you knew how to construct, which are rarely the ones that hurt you. - **Sample below the cut.** Reserve a small random sample of unflagged rows for review in every cycle. It is the only mechanism that can ever tell you the threshold is far too strict, and without it a quiet system is indistinguishable from a blind one. ## Stability as an admission requirement A detector built on random subsampling, or on a neighbourhood size you chose by feel, can produce a top-100 list that changes substantially between runs. Refit with several seeds, several subsample sizes and a range for any neighbourhood parameter, and measure the overlap between the resulting top-k lists. Low overlap does not merely mean noise: it means the number you measured as precision at k belongs to one arbitrary run and does not predict the next one. Stability is a precondition for the evaluation being meaningful at all. ## The organisational half of the decision A detector that is technically good and operationally unusable fails. - **Volume against capacity.** Alerts that nobody has time to open are equivalent to alerts that were never raised, except that they create the illusion of coverage. - **Adjudicability.** An analyst cannot act on "score 0.81". They need which features were unusual and what a typical comparable row looks like. If the output cannot be explained to the person who must act on it, the review loop never produces reliable labels either. - **The trust budget.** Reviewer confidence is spent, not renewed. A launch with poor precision teaches the team to ignore the queue, and the second launch inherits that. Launching narrow and strict, then loosening as precision holds, is usually the right sequencing. - **Labels are the real deliverable.** Every adjudication is a label. After a few months of review you own a labelled set that supports far sharper evaluation than anything available at launch. Design the review workflow to capture the outcome, the reviewer and the reason from day one, and remember that labels gathered only from flagged rows are a biased sample - which is what the below-cut sampling is for. ## A defensible launch gate Something like: precision at 100 above an agreed floor from a blind expert review; the known historical incidents ranked inside the top band; top-100 overlap above an agreed level across seeds; daily volume inside review capacity; and every alert carrying a per-feature reason. Numbers vary by domain - the shape of the argument is what an interviewer is listening for.
- Why can't you report recall for an unsupervised detector at launch?Recall needs the count of anomalies you missed, which requires labels for the rows you never flagged. Expert review only covers the top of the ranking, so it produces precision and nothing else. The workable substitutes are replaying confirmed historical incidents to see where they rank, injecting realistic synthetic anomalies, and reviewing a small random sample from below the cut.
- How do you keep expert review from producing biased labels?Adjudicate blind to the score where practical, mix in rows from below the cut so reviewers are not seeing an all-suspicious queue, have two reviewers overlap on a subset and measure their agreement, and write down the decision rule so the definition of an anomaly does not drift month to month. Also record the reason, not just the verdict.
- Two detectors have similar precision at 100. What separates them?Stability of the top of the ranking across seeds and parameter settings; whether their alerts are adjudicable, since a score with no per-feature reason costs far more reviewer time per item; how they rank known historical incidents; the overlap between their alert sets, because a detector that surfaces a different failure mode may be worth running alongside rather than instead; and the cost of scoring at production volume.
saying these in an interview costs you the question
- Reports accuracy for a detector trained without labels
- Reviews only flagged rows and calls the result full evaluation
- Treats a high anomaly score as proof the row is actionable
- Ignores that the top-k list changes between random seeds
- Ships alert volume far beyond what reviewers can process
- Shows analysts a bare score with no per-feature reason