When do you pose anomaly detection as one-class modelling of normal data instead of supervised classification?
answer
- rare is not the same as unseen
- does the future look like the labels?
- model normal when the positives are open-ended
- unusual is not the same as important
- clean normal data is the hidden requirement
basics
~20 sChoose one-class modelling when the anomalies you will face are not represented by the anomalies you have labelled, so no boundary learned from past examples transfers. Choose supervised classification when labelled anomalies are plentiful and representative of what comes next.
solid answer
~50 sThe deciding factor is whether "anomalous" is a class at all. A supervised classifier learns a boundary from labelled examples, which assumes future positives look like past positives. In network intrusion that assumption fails by definition: next month's attack is a technique absent from your history, so a classifier trained on last year's attacks will be strong on known modes and blind to the new one. Posing it as one-class or density work inverts the problem — model what normal traffic looks like and flag departures from it — which requires no examples of the thing you are hunting. The costs are real: you need reasonably clean normal data, normality has to be stable enough that drift does not flood you with alerts, and statistically unusual is not the same as interesting to the business. Rarity alone is not a reason to abandon supervision; a rare but well-labelled and repeating class is still a classification problem.
go deeper
Know the basic split: supervised needs labelled examples of the thing you are looking for, while one-class modelling learns what normal looks like and flags departures without ever seeing an anomaly.
Explain why a classifier trained on past anomalies assumes future anomalies resemble them, and why rarity by itself is a class-imbalance issue rather than a reason to change framing.
Demonstrate the operating judgement: contaminated normal data, drift in normal turning into alert floods, the gap between unusual and important, and the two-channel design for known plus unknown modes.
Own the framing as a long-lived commitment. Decide what the organisation maintains, who reviews the queue, what evidence would justify moving a mode from the novelty channel to a supervised one, and what coverage you are accepting you will not have.
## The question behind the question Interviewers asking this are checking whether you can distinguish **rare** from **unseen**. They are different problems with different framings, and conflating them is the single most common error here. ### Supervised framing: anomaly as a class You have labelled examples of anomalies and of normal items, and you learn a boundary between them. This framing makes one strong assumption: **the anomalies you will meet are drawn from the same distribution as the anomalies you labelled.** It is a closed-set assumption. When it holds, supervised learning is usually the stronger choice, because it uses the structure of the anomalies themselves, not just their distance from normal, and because you can measure it honestly against labels. ### One-class framing: anomaly as "not normal" You model only the normal population — its region of support, its density, its typical structure — and score new items by how poorly they fit. No labelled anomalies are required, and no assumption is made about what an anomaly looks like beyond "unlike normal". This is sometimes called novelty detection, and it is the natural framing for **open-set** problems where the categories at deployment time were absent at training time. ### The deciding criteria **1. Are future anomalies represented in your labels?** In network intrusion, next month's attack type is unseen by definition — the whole point of an attack is to do something the defender has not enumerated. An adversary actively selects the technique you have no examples of. That argues for modelling normal traffic and flagging departures, because the model does not need to have seen the attack. Contrast a manufacturing line with three well-understood defect modes that recur every week: those are labelled, repeating and closed-set, and supervised classification is the right frame. **2. How many labelled anomalies do you actually have, and how were they obtained?** A handful of labelled positives is weak evidence about a whole class. Worse, if the labels came from what a previous alerting system surfaced, they describe that system's notion of anomaly rather than the phenomenon. **3. Is rarity the only issue?** If positives are rare but well-labelled and behave consistently, that is class imbalance inside a supervised problem, handled by the usual means, and switching to one-class modelling throws away real information for no reason. "Only 0.5% positive, so I will use anomaly detection" is a wrong answer. **4. Do you have a clean, representative sample of normal?** One-class framing depends on it. If your "normal" training data quietly contains the very anomalies you want to find, the model learns to consider them normal. If normal drifts seasonally, every seasonal shift becomes an alert storm. **5. Is unusual the same as bad?** This is the framing trap. Density-based scoring finds statistically atypical records: a legitimate but enormous corporate transfer, an unusual-but-fine maintenance window, a new product launch. Nothing in the framing knows what the business cares about. Supervised framing does, because the labels encode it. ### The open-set middle ground Consider weld inspection on a production line. You trained on three defect modes captured during commissioning, and production is now throwing modes nobody photographed — porosity patterns from a new wire supplier, say. A pure classifier will confidently assign every new defect to one of its three known classes, because that is all its output space contains; it has no way to say "this is none of the above". A pure one-class model will flag every unusual weld, including harmless cosmetic variation. The honest framing is usually **both channels**: a supervised model for the known, labelled modes where it is accurate and explainable, plus a novelty channel that catches items far from anything the training set contained and routes them to a human. The two answer different questions — *which known defect is this?* and *is this something we have never seen?* — and an item can trigger either. ### Consequences of the framing choice - **Evaluation.** A supervised framing can be measured against held-out labels. A one-class framing usually cannot, because the ground truth for "anomalous" does not exist at scale; you end up validating on whatever a human reviewer confirms, which covers only what you surfaced. - **Output shape.** One-class scoring produces an ordering of items by strangeness, consumed as a review queue rather than as decisions. - **Explanation burden.** A supervised model can point at the pattern it recognised. A novelty score says only "this is unlike your normal data", so the operational design has to include how a human works out why. - **Maintenance.** Supervised models decay when anomaly modes shift. One-class models decay when *normal* shifts, which happens for entirely innocent reasons far more often. ### Answer shape Ask whether future anomalies are represented in the labels; if not, model normal. Say explicitly that rarity alone does not force the one-class framing. Then name the cost you take on — clean normal data, drift sensitivity, and the gap between unusual and important — and propose the two-channel design where both known and unknown modes matter.
- You have 4,000 labelled intrusions from last year. Does that settle it in favour of supervised classification?No. Volume is not representativeness. Those 4,000 cover techniques that were current last year and that the old detection stack managed to catch, so they are a biased sample of a moving target. I would train a supervised model on them because known modes still recur and it will be accurate and explainable there, but I would keep a novelty channel on normal traffic as well, because the label set cannot contain next month's technique.
- Your normal-traffic training window probably contains some undetected intrusions. Does that break the one-class framing?It degrades it rather than breaking it. A small fraction of contamination mostly widens the estimated normal region slightly, but anything present in volume gets learned as normal and becomes permanently invisible. I would use the cleanest window I can defend, prefer periods with independent confirmation of health, and treat any large cluster of look-alike outliers with suspicion, since a coordinated campaign inside the training window can look like an ordinary sub-population.
- How would you tell an interviewer the difference between an outlier and something the business cares about?An outlier is a statistical statement: the record sits in a low-density region of the data. Business interest is a value statement: acting on this record saves money or prevents harm. They coincide only when the process generating harm also generates statistical strangeness. When they diverge, and they often do, the one-class framing will keep surfacing legitimate rarities, and the fix is either labelled feedback or business rules layered on the ordering.
Airport security has both a watchlist and a behaviour-detection team. The watchlist catches people you already know about; the behaviour team exists precisely because tomorrow's threat is not on it.
saying these in an interview costs you the question
- Switches to anomaly detection just because positives are rare
- Assumes a classifier generalises to attack types never labelled
- Ignores that training normal data may contain hidden anomalies
- Equates statistically unusual with operationally important
- Thinks a closed-set classifier can output none of the above