You are assembling the benign half of a test corpus for a content-moderation guard. Where do the benign items come from, and what makes a benign request a hard negative rather than an easy one?
answer
- real logs plus constructed near-misses
- hard = close to the boundary
- professional, creative, meta contexts
- zero over-blocks means the set is too easy
- label review, freeze, version
basics
~20 sPull benign items from the product's real traffic and from topics sitting right next to the blocked ones: safety questions, clinical or legal wording, security research, fiction, non-English phrasing. A hard negative looks like an attack on the surface but is a request the product must answer. Label them, review the labels, freeze the set.
solid answer
~60 sTwo sources, and you want both. **Real traffic** — sampled and scrubbed logs from the product itself — anchors the set to what users actually send, including the messy phrasing no one would invent. **Constructed near-misses** cover the adjacency the logs are thin on: requests that use the vocabulary of a blocked category for a legitimate purpose. What makes a negative *hard* is surface similarity to the positive class. A greeting is benign and the guard will pass it, so it measures nothing. A clinician asking about overdose thresholds, an analyst asking how a malware family behaves, a novelist writing a violent scene, a moderator quoting the abuse they are reporting — these sit close enough to the decision boundary that a cautious guard refuses them, which is precisely the failure you are trying to price. The hard part is labelling: someone with product and domain context has to agree each item genuinely deserves an answer. If your team labels a borderline item benign and the product's own policy would refuse it, you will report an over-block that is actually policy working as intended.
code
json · 1 line{"id": "benign-clinical-014", "label": "benign", "near_category": "self-harm", "source": "prod-log-sample-2", "expected": "answer", "rationale": "clinician-facing dosage safety question", "labeled_by": "product-policy-reviewer", "set_version": "benign-v3"}go deeper
Should say benign items must resemble real user requests and come from the product's own traffic, not invented small talk.
Should define hardness as surface similarity to blocked content, name the professional, creative and meta near-miss veins, and mention labelling and freezing the set.
Should treat the set as a versioned measurement asset, use a zero over-block rate as a signal the set is too easy, and handle label disagreement as a policy finding.
Should raise who owns and funds the benign corpus over time, the privacy handling of sampled traffic, and budgeting labelling effort at engagement start.
A benign hard-negative set is a measurement asset. Build it like one — sourced, labelled, versioned, reused — and above all build it so its items sit near the guard's decision boundary, because items far from the boundary cannot move any number you care about. ### Hardness is measurable, not a vibe Most guards expose something continuous underneath the verdict. OpenAI's moderation response carries `category_scores` alongside the `flagged` boolean. Azure AI Content Safety returns a per-category severity of 0, 2, 4 or 6 and leaves the block cut-off to the operator. Llama Guard's verdict is a generated token, so the probability it assigned to `unsafe` is available if you request logprobs. Record that continuous value for every benign item, not just pass/fail. The distribution of those values is the diagnostic. A benign set whose scores all pin at the floor is an easy set, and its 0% over-block rate is a property of the set, not of the guard. A well-built set puts real mass just under the cut-off: items that pass today and will be refused the first time anyone tightens by one step. That is what "hard" means operationally — distance to the boundary, which you can see, rather than a subjective judgment that an item "looks edgy". ### Where the items come from Two veins, and you want both. **Sampled production traffic** anchors the set to the distribution users actually produce: truncated messages, pasted context, mixed languages, phrasing nobody would invent. It needs a privacy path — scrubbing identifiers, and usually a review before the corpus may leave the engagement — which is calendar time rather than compute time, so start it on day one. **Constructed near-misses** fill the adjacency the logs are thin on. For each category the guard screens, write legitimate requests that use that category's vocabulary. The reliable veins are professional (clinical, legal, security research, journalism), creative (fiction, historical writing) and meta (reporting, quoting or moderating abusive content in order to act on it). ### Labelling and versioning "Benign" is a judgment about this product's policy, not a universal fact, so the label needs a reviewer who knows that policy. Budget one to two minutes per item plus adjudication: 500 items is about a reviewer-day plus a half-day of disputes, and that human time, not the endpoint bill, is what makes benign sets get skipped. Keep the disagreements rather than resolving them by fiat — an item two reviewers split on is an item the guard will also split on, and it belongs in the report as a policy question, not as a guard defect. Then freeze the set, stamp it with a version, and grow it by appending a new slice rather than editing items in place. Rerun value comes entirely from comparability: if the benign half drifts between runs, a change in the over-block rate is unattributable. Keep the per-item decisions, not only the aggregate, so you can diff exactly which items flipped verdict. ### How the number misleads **Easy-set zero.** An over-block rate of zero reads as "the guard is permissive". Far more often it means the items sit far from the boundary and the set could not have detected anything. Look at the score distribution before you publish the zero. **Generator collusion.** If you generate benign items with the same model family that powers the guard — or with the same model that generated your attack half — you are drawing from a distribution that model finds unremarkable. The set is optimistically biased in a way no amount of volume fixes, because the bias is in the sampler, not the sample size. **Survivorship in the logs.** Production traffic is the traffic that survived the guard's reputation. Users refused three months ago stopped asking; whole professional use cases quietly moved to another tool. The requests most likely to be over-blocked are systematically thinnest in the logs, so a log-only benign set understates over-blocking precisely where it is worst. That is why constructed near-misses are not optional garnish — they are the correction for a bias the logs cannot show you. ### What to check Tabulate the guard's continuous score across the benign set and confirm real mass near the cut-off. Compute inter-annotator agreement on a labelled sample and report it, because an over-block rate is only as trustworthy as the labels beneath it. Confirm every benign item got a recorded decision and none errored out silently. Confirm the set version and the guard configuration are both stamped in the run record. And re-read ten items at random as a product owner would: if you would not be annoyed to see them refused, they are not hard negatives.
- Your benign set produces an over-block rate of zero. What is your first hypothesis?That the set is too easy — items sit far from the decision boundary — not that the guard is permissive. Add near-misses in the blocked categories and re-measure.
- Two reviewers disagree on whether an item is benign. What do you do with it?Keep it, flag it, and surface it as a policy ambiguity in the report rather than scoring it silently either way — the guard will be ambiguous about it too.
- Why version the benign set instead of improving it in place?Because rerun comparisons are only meaningful against a fixed set; editing items in place makes a change in over-block rate unattributable.
Sampling today's traffic for benign examples is like surveying a shop's current customers about the queue: the people who gave up and left are exactly the ones you needed to hear from, and exactly the ones missing.
saying these in an interview costs you the question
- Filling the benign half with greetings and small talk and calling it done
- Generating benign items from the same source that generated the attacks, with no product context
- Editing the benign set in place between runs, then comparing over-block rates across runs
- Treating a zero over-block rate as good news without checking whether the set was hard enough
- Putting unscrubbed production traffic into a corpus that leaves the engagement