How do you pick the score threshold for a content-moderation classifier?
answer
- it is a policy decision, not a default
- two errors, unequal costs
- imbalanced data breaks one common curve
- bands beat a single cutoff
- the queue you can staff constrains it
basics
~20 sNot from a default. Score a labelled sample, plot precision against recall per category, and choose the operating point where the cost of a false positive and the cost of a missed harm balance for that category — then check the resulting flag volume fits your review capacity.
solid answer
~50 sA threshold is a business decision dressed as a number. Every setting trades two errors against each other: false positives, where you punish a legitimate user, and false negatives, where harmful content ships. Those costs are not symmetric and not the same across categories, so there is no universally correct value and 0.5 is only a default, not a recommendation. The method is to label a representative sample of real traffic, sweep the threshold, and plot the precision-recall curve per category — precision-recall rather than ROC, because harmful content is rare and ROC flatters a classifier under heavy class imbalance. Then pick the point that matches the policy. In a shooter's party chat, moving the harassment threshold from 0.5 to 0.82 cut false positives by about 70% while missed abuse rose 4% — a good trade when wrongful mutes were driving churn. One more constraint binds the choice: the flag volume you produce must fit the human review queue's capacity, or the threshold is theoretical.
code
python · 11 linesdef sweep(scored, thresholds):
for t in thresholds:
tp = sum(1 for s, y in scored if s >= t and y)
fp = sum(1 for s, y in scored if s >= t and not y)
fn = sum(1 for s, y in scored if s < t and y)
precision = tp / (tp + fp) if tp + fp else 1.0
recall = tp / (tp + fn) if tp + fn else 1.0
print(f"t={t:.2f} precision={precision:.2f} recall={recall:.2f} flagged={tp + fp}")
labelled = [(0.95, True), (0.88, True), (0.60, False), (0.55, True), (0.20, False)]
sweep(labelled, [0.50, 0.82])go deeper
Know that a classifier returns a score and that you choose a cutoff, and that raising it means fewer wrongly flagged users but more harmful content getting through.
Explain precision and recall in moderation terms and show how sweeping the threshold moves along that curve, including why 0.5 is a default rather than a recommendation.
Demonstrate you have made the call: per-category thresholds set from unequal error costs, precision-recall over ROC under class imbalance, bands with a human-review middle, and re-tuning after drift or a classifier change.
Own the threshold as policy — who signs off on the tradeoff, how enforcement rates are reported, what regulatory or platform exposure a missed harm carries, and how much review headcount the chosen operating point commits you to.
## What a threshold actually is A moderation classifier returns a continuous score per harm category, usually in [0, 1]. The threshold is the cutoff at which you take action. Everything above it is treated as a positive; everything below is allowed through. The classifier's quality is fixed by the model; the threshold is entirely yours, and it is where nearly all the product behaviour lives. ## The two errors and their costs - **False positive**: legitimate content is flagged. Cost = a wrongly muted player, a support ticket, an appeal, churn, and erosion of trust in the system. - **False negative**: harmful content ships. Cost = harm to a real user, possible regulatory or platform exposure, reputational damage. These costs differ by orders of magnitude across categories, and they are not symmetric within one. Wrongly muting someone for salty banter is a bad afternoon; missing a credible threat is a safety incident. That asymmetry is the entire input to threshold choice. If a candidate answers "maximise F1", press them: F1 weights the two errors equally, which is almost never the actual policy. ## Reading the precision-recall curve To choose an operating point you need labelled data — a sample of real traffic that humans have adjudicated per category. Sweep the threshold across that sample and compute, at each value: - **precision** = of what you flagged, the fraction that was truly harmful - **recall** = of what was truly harmful, the fraction you caught Plotting precision against recall gives the achievable frontier. Use the **precision-recall curve rather than ROC** for moderation: harmful content is typically well under 1% of traffic, and ROC's false-positive rate has that enormous negative class in its denominator, which makes a mediocre classifier look excellent. Precision, which has flagged-volume in its denominator, tells you the truth a reviewer will experience. ## A worked move A shooter's party-chat harassment classifier ran at the default 0.5. Reviewers reported that most of what reached them was ordinary competitive trash talk, and the support queue was full of wrongful-mute appeals. Moving the threshold to 0.82 cut false positives by roughly 70% while missed abuse rose 4%. That was the right call for that product, because the dominant cost was wrongful enforcement on paying players — but it is a *policy* judgement, not an optimisation result. On a category like child safety, a 4% rise in misses to buy a cleaner queue would be indefensible, and the same team would move that threshold in the opposite direction. ## Bands, not a single cutoff Mature systems rarely use one cutoff per category. They use bands: - above a high cutoff — automated enforcement, no human in the loop - between a low and high cutoff — allow, but escalate to human review - below the low cutoff — allow silently Bands convert a binary decision into a routing decision and let you keep automated action precise while still catching the uncertain middle. They also give you a knob that changes reviewer load without changing enforcement. ## Operational constraints on the choice A threshold you cannot staff is not a threshold. If the middle band produces 3,000 escalations a day and the review queue can absorb 900, the excess is either silently allowed or piles into a growing backlog — both of which quietly change your effective policy. Compute expected flag volume at each candidate threshold from your traffic rate and the score distribution, and treat queue capacity as a hard constraint alongside the precision and recall targets. ## Calibration and drift Two further traps: - **Scores are not probabilities.** A 0.82 does not mean an 82% chance of harm unless the classifier was calibrated. Thresholds are ordinal operating points on *this* classifier's score distribution and do not transfer across classifiers or across model versions. Any classifier upgrade invalidates every tuned threshold and requires a re-sweep. - **Distributions drift.** Slang, in-game events, new game modes and coordinated abuse campaigns all shift the score distribution. Re-measure precision and recall on a fresh labelled sample on a schedule, and alert on flag-rate changes that are not explained by traffic changes. ## What interviewers listen for Strong answers name the two error types and their asymmetric costs, choose the precision-recall curve over ROC and can say why, set thresholds per category rather than globally, mention bands with a human-review middle, and treat review capacity as a real constraint. Weak answers say "we tuned it until it felt right" or reach straight for F1.
- Why prefer the precision-recall curve over ROC when tuning a moderation threshold?Because harmful content is rare. ROC plots true-positive rate against false-positive rate, and the false-positive rate divides by an enormous negative class, so it barely moves even when most of what you flag is wrong. Precision divides by flagged volume instead, which is exactly what a reviewer experiences. Under 1% prevalence, ROC can look excellent while precision is under 10%.
- You upgrade to a new moderation classifier. What happens to your tuned thresholds?They are invalid. A threshold is an operating point on one classifier's score distribution, not a portable probability, so the same numeric cutoff means something different on a new model. Re-label or reuse a held-out sample, sweep again per category, and pick the point that reproduces your target precision and recall before switching enforcement over.
- How would you set a threshold for a brand-new harm category with no labelled data?Start in shadow mode: score traffic, take no enforcement action, and sample across the score range for human labelling. That gives you both the score distribution and a labelled set within days. Until then, run the category escalation-only with a deliberately conservative cutoff so it generates review signal rather than enforcement you cannot yet justify.
- What signal tells you a previously well-tuned threshold has drifted?Flag rate per category moving without a matching change in traffic volume or mix, reviewer overturn rate on flagged items climbing, and a rise in appeals or in user reports about content that was allowed. Any of those triggers a fresh labelled sample and a re-sweep rather than a reflexive nudge to the number.
A moderation threshold is like the sensitivity dial on a smoke alarm: turn it up and the burnt-toast alarms stop, but you also find out about a real fire a little later.
saying these in an interview costs you the question
- Uses 0.5 because it is the default
- Optimises F1, weighting both error types equally
- Sets one global threshold across all harm categories
- Reads scores as calibrated probabilities of harm
- Ignores whether review capacity can absorb the flag volume
- Never re-tunes after a classifier upgrade