skip to content

How do you set an assumed contamination rate, or a one-class SVM's nu, with no labels?

level: seniorimportance: should knowfreq 45%

answer

  1. it is an assumption, not an estimate
  2. a threshold on a ranking, not a refit
  3. nu sits inside the training objective
  4. anchor on how many alerts get reviewed
  5. revise it from the review outcomes

basics

~20 s

Set the contamination rate from review capacity, not from a guess at the truth: it only chooses where to cut a score ranking, so pick the cut that yields an alert volume your team can actually work through.

solid answer

~50 s

First separate two different knobs. For a score-based detector, the assumed contamination rate is a post-hoc decision threshold: it cuts the ranking at the quantile corresponding to that fraction and changes nothing about the ranking itself. A one-class SVM's `nu` is different - it enters the training objective, upper-bounding the fraction of training points allowed to fall outside the learned boundary and lower-bounding the fraction of support vectors, so changing it refits the model. In both cases, with no labels, the honest anchor is capacity. If a security operations team can genuinely work through 50 flagged login sessions a day, set the boundary so it emits about 50, then measure how many of those the analysts confirm. Cross-check that number against any historical incident rate you have and against a score-distribution plot, and treat the first value as provisional.

go deeper

for a junior

Know that these parameters decide how many rows get flagged, not how accurate the detector is, and that the value is assumed rather than learned from the data.

for a middle

Be able to state the mechanics: contamination cuts a score ranking at a quantile and leaves the ranking untouched, while nu bounds the fraction of training points outside a one-class SVM's boundary and the fraction that become support vectors.

for a senior

Show the operating discipline - anchor the threshold on review capacity, keep raw scores so the cut can move without refitting, sample below the cut, and re-check after population drift.

for a principal

Frame it as an economic decision: the cost of an analyst hour against the cost of a missed event, and the failure mode where alert fatigue destroys a technically sound system. Say who owns the number and how often it is renegotiated.

## Two knobs that look alike and are not Candidates routinely conflate these, and separating them is half the answer. **Assumed contamination rate.** Score-based detectors - an isolation forest, a density ratio - produce an ordering. To emit a yes/no decision you must cut that ordering somewhere. Declaring a contamination rate of 1% means "threshold at the 99th percentile of the scores". Crucially it is a **decision** parameter: the ranking is identical whether you assume 0.1% or 10%, and only the label boundary moves. If you keep the scores rather than the labels, you can change your mind at any time without refitting. **A one-class SVM's `nu`.** This one sits inside the optimisation. `nu` in `(0, 1]` simultaneously upper-bounds the fraction of training points permitted to fall on the wrong side of the learned boundary and lower-bounds the fraction of training points that become support vectors. Raising `nu` therefore buys a tighter boundary that rejects more of the training data, at the cost of a more complex model held up by more support vectors. Changing it means refitting, and the resulting decision boundary is a different shape, not the same shape cut in a different place. ## What the number cannot be derived from With no labels there is no quantity in the data that tells you the true anomaly rate. Any claim to have estimated it from the unlabelled data alone is really an assumption about the shape of the score distribution. So say plainly that the parameter is an assumption, and choose the assumption from something outside the data. ## Anchors that actually work **Review capacity first.** This is the most defensible anchor and the one senior candidates reach for. If analysts can adjudicate 50 items a day and each takes fifteen minutes, then the operating point that produces 500 alerts a day is not conservative, it is useless - the surplus is silently dropped, and which items get dropped is arbitrary. Set the boundary to fill the queue and no more. **Known base rates from outside the model.** Incident counts from previous quarters, a defect or failure rate reported by the business, a regulator's published estimate. These are noisy and usually undercount, since they only include what was caught before, but they bracket the plausible range. **The shape of the score distribution.** Plot the scores. A visible thin right tail separated from the bulk suggests a natural cut; a smooth unimodal spread means no cut is natural and the choice is purely economic. This diagnoses the situation - it does not decide it. **A clean training set, when you can get one.** A one-class SVM is at its best fitted only on data you have reason to believe is clean, for example login sessions from accounts later confirmed uncompromised. Then `nu` no longer stands for "how many anomalies are in here" - it stands for how much of that known-good data you are willing to have rejected, in other words your tolerated false-alarm rate on normal traffic. That is a much easier quantity to reason about, and it is the strongest argument for curating a clean fit set. ## The trap of fitting on contaminated data Most detectors are in practice fitted on whatever data you have, anomalies included. If the true rate is far higher than assumed, the anomalies help define what the model considers normal - the boundary stretches to cover them and they stop being flagged. If the assumed rate is far higher than the truth, the cut slices deep into genuinely normal data and the queue fills with false alarms that burn reviewer trust, which is usually the fatal outcome: analysts stop reading the alerts and the system dies whether or not the model was good. ## Making the parameter provisional Treat the first value as a starting point with an explicit revision loop: 1. Ship a threshold set by capacity. 2. Record every review outcome. Those adjudications are the beginning of a labelled set. 3. After enough reviews, measure the confirmed-anomaly rate in the flagged band. If nearly everything flagged is confirmed, you are almost certainly cutting too shallow and missing real cases - loosen it. If almost nothing is confirmed, tighten it. 4. Sample a small number of items from *below* the cut for review too. Without that, you only ever learn about the region you already flag, and you will never discover that the threshold is far too strict. 5. Re-check after any population shift, because the same score quantile corresponds to a different absolute score once the data moves. ## What good sounds like in an interview "It is an assumption, not an estimate. I set it from the review capacity, sanity-check against the historical incident rate and the score histogram, keep the raw scores so I can move the cut without refitting, and build a revision loop from the review outcomes - including a small random sample from below the cut so I can tell whether I am too strict."

  • Does changing the assumed contamination rate change which rows are ranked most anomalous?
    No, not for a score-based detector. The ranking comes from the scores and is fixed; the contamination rate only picks the quantile where you cut it into flagged and not flagged. That is a good reason to persist the scores rather than only the labels - you can move the operating point later without refitting anything. A one-class SVM's `nu` is the opposite case: it changes the fit itself.
  • What goes wrong if the true anomaly rate is far higher than the rate you assumed?
    The detector is usually fitted on the same contaminated data, so a large anomalous population helps define what counts as normal: the boundary stretches to include it and those cases stop scoring high. You get a comfortable-looking, quiet system that is blind to the very thing it was built for. Curating a clean fit set, or holding out suspected-bad rows from fitting, is the defence.
  • How do you find out whether your threshold is too strict?
    Send a small random sample of items from below the cut for review alongside the flagged queue. Reviewing only what you flag teaches you precision and nothing about what you are missing, so the loop can sit at a far too strict threshold indefinitely and look healthy. If the below-cut sample keeps turning up confirmed cases, loosen the boundary.
  • What does raising nu do to a one-class SVM's boundary?
    It permits a larger fraction of the training points to fall outside the boundary and forces at least that fraction to act as support vectors, so the boundary tightens around the bulk of the data and grows more complex. Lowering it produces a looser, smoother boundary that admits nearly all the training data. Either way the model is refitted, unlike a post-hoc score threshold.

saying these in an interview costs you the question

  • Claims the true contamination rate can be estimated from unlabelled data
  • Confuses nu, a training parameter, with a score threshold
  • Leaves contamination at a default and never revisits it
  • Ignores how many alerts reviewers can actually handle
  • Reviews only flagged items and calls the result an evaluation
  • Says a higher nu always detects more true anomalies

context