skip to content

Why add negative exemplars showing what a classifier must not flag?

level: seniorimportance: should knowfreq 36%

answer

  1. positives give a direction, not a border
  2. the useful negatives are near misses
  3. state why it was not flagged
  4. corrections overshoot the other way
  5. track both error rates, separately

basics

~20 s

Positive-only demonstrations teach where the rule applies but never where it stops, so the model over-generalizes and flags look-alikes. Near-miss negatives — cases that resemble a hit but are not one — draw the boundary the positives leave undefined.

solid answer

~50 s

A set of positive examples defines a direction, not a border. The model infers whatever pattern the positives share, which is usually broader than the actual rule, and the result is over-flagging on inputs that merely look like hits. The fix is contrastive: include exemplars that are deliberately close to the positive cases but correctly labeled negative. In an export-control screening prompt, that means shipments whose product descriptions read like controlled goods but fall outside the control list, each shown with the negative label and, ideally, a one-line reason. The reason matters — a bare negative label tells the model "not this one", while a stated reason tells it *which feature* was decisive. Watch the balance: too many negatives push the model the other way into under-flagging, so track false-positive and false-negative rates together on a held-out set after every change to the set.

go deeper

for a junior

Know that showing only cases that should be flagged teaches the model to over-flag, and that examples of what should not be flagged are part of a complete exemplar set.

for a middle

Explain that positives define a direction rather than a boundary, and that near-miss negatives — close to a positive but correctly negative — are what actually mark the edge.

for a senior

Demonstrate the operating discipline: source near misses from reviewer-cleared production flags, carry a stated reason on each, change one thing at a time, and read false positives and false negatives separately after every change.

for a principal

Own the error-cost position — which direction you are willing to be wrong in and by how much — and insist that the exemplar pool is not used to patch an ambiguous policy that should be rewritten and re-adjudicated instead.

## What positives alone actually teach Give a model six examples of documents that should be flagged and it will extract whatever those six have in common. That commonality is almost never the real rule. It might be a vocabulary, a document length, a phrasing style, or a source system — any correlate that happens to be shared. The model then applies that over-broad pattern to new input, and the visible symptom is a flood of false positives on things that superficially resemble the training examples. This is the quiet failure mode of exemplar selection: nothing looks wrong in the prompt. The examples are all correct. The set is simply one-sided, and one-sidedness is invisible until you look at the errors. ## Near-miss negatives do the work Not all negative examples are equally useful. A negative that is wildly unlike any positive — a random unrelated document — teaches almost nothing, because the model was never going to flag it. The valuable negatives are the **near misses**: inputs that share most surface features with a positive but land on the other side of the rule. In an export-control screening prompt, a useful negative is a shipment whose description uses the same technical vocabulary as a controlled item but describes a variant below the control threshold, or a destination that looks sensitive but is not on the restricted list. Each such pair — a positive and its near-miss negative — pins one dimension of the boundary. Three or four well-chosen pairs teach more than a dozen easy negatives. ## State the reason, not just the label A negative demonstration labeled only "not flagged" tells the model that this input is out, but leaves it to guess which feature was decisive. If the exemplar carries a short reason — naming the specific attribute that put it outside the rule — the model has a much better chance of applying the same distinction to a new case. This also has a review benefit: a human auditing the exemplar pool can check whether the stated reason matches the actual policy, which is far harder when the exemplar is a bare label. Keep the reasons in the same shape across all exemplars, positive and negative alike. Inconsistency between how positives and negatives are presented gives the model a shortcut — it can learn to predict the label from presentation rather than content. ## The over-correction risk Negatives are a correction, and corrections overshoot. A set that becomes negative-heavy pushes the model toward not flagging, which in a screening context is the more dangerous direction: a missed control violation costs far more than a false alarm a human clears in thirty seconds. So the balance is not "equal counts"; it is "whatever balance produces the error profile you want", and that requires an explicit position on the relative cost of the two error types. Practically: - Decide the acceptable false-negative rate first, since it is usually the constrained one. - Measure false positives and false negatives separately on a held-out set after every change to the exemplar set. A single accuracy number hides a swap of one error type for the other. - Change one thing at a time. Adding three negatives and rewording the instruction in the same commit makes the resulting shift uninterpretable. ## When negatives are not the right fix Over-flagging has other causes, and reaching straight for negative exemplars can paper over them: - **An ambiguous rule.** If two reviewers disagree about the correct label on borderline cases, the model cannot do better than the humans. Fix the definition and the labeling guidance first. - **A miscalibrated threshold.** If the system emits a score and flags above a cutoff, the cutoff may simply be wrong; that is a tuning change, not a prompt change. - **Mislabeled positives.** If some of the flagged examples should not have been flagged, the model is faithfully learning a wrong rule. Audit the positives before adding anything. ## What good looks like in an interview A strong answer names the mechanism (positives define a direction, not a boundary), specifies *near-miss* negatives rather than negatives in general, mentions carrying a stated reason, and closes with the measurement — tracking both error types, with an explicit stance on which one you are willing to trade. A weak answer just says "add counter-examples" and stops.

  • How many negatives is too many in a screening prompt?
    There is no fixed ratio; the right balance is whichever one produces the error profile you have committed to. In screening, false negatives usually cost far more than false positives, so you set an acceptable miss rate first, then add negatives only while the false-negative rate stays inside it. Measure both rates after each change — a set that quietly halves false positives by doubling misses has made the system worse, not better.
  • Where do you get near-miss negatives from in the first place?
    From the system's own errors. Sample production inputs the model flagged that a human reviewer then cleared, and read the reviewer's reason — those are near misses by construction, already adjudicated. Adjudicated disagreements between two reviewers are the second-best source, because they sit exactly on the boundary. Both need the same PII and provenance handling as any other exemplar drawn from real records.
  • Would adding negatives fix over-flagging caused by an ambiguous policy definition?
    No, and it can hide the problem. If human reviewers disagree on the borderline cases, the exemplars themselves will encode that disagreement and the model will reproduce it inconsistently. Measure inter-rater agreement on a sample first; if it is poor, tighten the written definition and re-adjudicate the labels. Only once humans agree does adding near-miss negatives reliably move the boundary where you want it.

Teaching only with hits is like marking a hazard zone with flags inside it and none along the edge — people can tell where the danger is, but not where it ends.

saying these in an interview costs you the question

  • Adds only obvious, unrelated negatives that were never at risk
  • Uses a bare negative label with no stated reason
  • Assumes equal positive and negative counts is always right
  • Reports one accuracy number after changing the balance
  • Reaches for negatives when the real problem is an ambiguous rule

context