A team spot-checks 1% of incoming labelled training data - what does that bound?
answer
- review is sampled, the attack is chosen
- coverage is not detection
- who picks how many rows get written?
- 1% of the forty, not the million
- two ways to miss: draw, then read
basics
~20 sSpot-checking 1% bounds how obviously wrong a typical incoming row is, not whether the corpus was poisoned. Whoever writes poisoned rows chooses how many, and can write few enough that a 1% draw almost never lands on one.
solid answer
~50 sSampled review is a quality control built against accidental damage - vendor drift, a confused rater, a bad import - which spreads roughly evenly through the intake, so a small random sample meets it in proportion. An adversary with write access to the labelling queue is not a defect rate: they choose the volume. Two separate things must both happen for the spot-check to stop them. The review has to *draw* a poisoned item - with forty poisoned rows in a corpus of a million and 1% coverage, the expected number drawn is about 0.4 - and the reviewer has to *recognise* it once drawn, which is a different problem, because each item was written to be individually plausible. So `1%` is a coverage figure: it bounds what share of items got a human verdict. It is not a detection rate, and quoting it as a poisoning control confuses the two.
go deeper
Be ready to say what a spot-check rate measures - the share of items a human looked at - and to note that whoever writes poisoned rows picks how many to write, so coverage and detection are different numbers.
Explain the two conditions that must both hold: the sample has to draw a poisoned item, and the reviewer has to recognise it once drawn. Work the expected-count arithmetic out loud rather than asserting the conclusion.
Show that you would not accept a coverage figure as a poisoning control in a design review. Say what the number bounds, what it leaves entirely unbounded, and what you would ask about instead.
Own the framing: sampled review is priced against accidents, not against an adversary who sizes writes against it. Decide what assurance you are willing to claim from it in public, and where the next unit of spend actually goes.
## What a spot-check was designed to catch Sampled review of incoming training data is an inherited quality control, and it was designed against *accidental* damage: a labelling vendor drifting off the guideline, a rater who misread an edge case, a mis-configured import that shifted a column. Those defects have a **rate**. They are spread roughly evenly through the intake, so a small random sample meets them in proportion to how common they are, and reviewing 1% gives you a usable estimate of a defect rate for 1% of the cost. That is a good deal, and it is why the practice exists. An adversary is not a defect rate. Someone with write access to the queue that feeds training - a vendor-side position on a labelling pipeline, an insider on an intake process, an upstream source that is trusted by default - chooses three things an accident does not: **how many** items to write, **which** items, and **what each one looks like on its own**. ## Two independent conditions, and the sample rate only touches one For a sampled review to stop a poisoning attempt, both of these have to happen: **1. The sample has to draw a poisoned item.** This is arithmetic, and it is the arithmetic that gets skipped. The expected number of poisoned items drawn is the poisoned count multiplied by the reviewed share. Forty poisoned items in a corpus of a million, reviewed at 1%, gives an expected 0.4 draws - the review usually sees none of them at all. The number that feels reassuring is the wrong one: 1% of a million is ten thousand items reviewed, which sounds like a lot of work, and it is a lot of work. But the quantity that decides the outcome is 1% *of the forty*. **2. The reviewer has to recognise it once drawn.** This condition is untouched by the sample rate. A reviewer is handed one item and asked a per-item question - is this labelled correctly, is it plausible, does it match the guideline? An item chosen to survive exactly that question survives it. The reviewer is not shown the set the item belongs to, and is not asked to compare it against the other items written by the same source in the same week. Because the conditions are independent and multiply, doubling the sample rate at best doubles a small number. It does nothing at all to the second condition. ## Why the attacker's count can be so small - the goal decides the scaling Poisoning splits on goal, and the two halves behave completely differently under corpus growth. An **availability** or degradation attack wants the model measurably worse across the board. That genuinely scales with the share of the corpus the adversary controls, so it is diluted by more data, and because it needs volume a random sample meets many of its items. Sampled review is a real control against this. A **targeted** attack wants one chosen kind of input read a particular way, or a conditional behaviour keyed to something the adversary controls at inference. That behaves like an approximately **absolute number of examples** and is barely diluted by corpus size at all. A corpus that grows tenfold does not make such an attack ten times more expensive; it makes the defender's quoted review percentage ten times smaller. ## What the coverage number does legitimately bound It is a real number and worth having. It bounds the share of items that received a human verdict. It estimates labelling quality and catches vendor drift early. It deters lazy, bulk, obviously-wrong writes, because those show up in proportion. And it gives you the honest denominator for any assurance statement you later have to make. What it cannot do is convert into a statement about the unreviewed remainder, because the thing you are worried about chose its own size against precisely that remainder. ## The sentence to have ready "We reviewed X% of intake item-by-item this quarter and rejected nothing individually implausible" is true, defensible, and worth saying. "Our training data is not poisoned, because we spot-check 1%" is not a claim the number supports, and an interviewer asking this question is checking which of the two sentences you reach for.
- Does raising the spot-check rate from 1% to 5% change the picture?It multiplies expected draws by five - about 0.4 becomes about 2 for a forty-item attack - and multiplies the review bill by five. The second condition still holds: those two drawn items were written to be individually plausible and pass, and an adversary can absorb a couple of losses by writing a few more. You buy a linear improvement against someone who can re-size linearly and far more cheaply than you can staff.
- Which kind of poisoning would a 1% spot-check plausibly catch?A crude, high-volume one. An attack that wants the model measurably worse overall must control a real share of the corpus, so a 1% sample meets many of its items, and bluntly mislabelled items are individually reviewable - a rater looking at one can say the label is wrong. Sampled review is a genuine control against volume and against obviously wrong labels. It is weakest exactly where the attack is small and each item is defensible on its own.
- What should you ask for instead of the sample rate?Who can write into the training queue at all, and what the reviewer was actually asked to decide. The first bounds the population of people who could size an attack against your review; the second tells you whether any part of the process ever looks across items rather than at one item at a time. Both say more about targeted poisoning than a percentage does.
A random bag check at an exit is priced against forgetful people. It does very little to someone who counted the checks first and carried out only what a rare check would wave through.
saying these in an interview costs you the question
- Quotes a sample rate as if it were a detection rate
- Assumes poisoning must be a large fraction of the data
- Treats drawing a poisoned item as catching it
- Says a bigger corpus dilutes any poisoning attack
- Trusts that reviewers would simply notice bad data