skip to content

Passing the Spot-Check

Human review is sampled and poison is targeted, so the arithmetic settles it before anyone reads a row. Interviewers use it to see whether a stated review rate is treated as a number.

on this pageshow

explore

questions

4

Why does fixed-hours human review cover less of a training corpus each year while a targeted poisoning attack needs no more rows?

level: middleimportance: must knowfreq 55%

answer

  1. a flat numerator over a growing denominator
  2. fraction versus absolute count
  3. which poisoning goal does growth dilute?
  4. hours cap the count, never the share
  5. same items reviewed, smaller slice of intake

basics

~20 s

Reviewed items are capped by staffed hours, so the reviewed share is a flat count over a growing intake and falls yearly. A targeted attack needs roughly a fixed number of rows, not a fixed fraction, so its cost stays flat.

solid answer

~50 s

Two quantities move in opposite directions. The reviewable count is staffed hours times items per hour - a headcount decision, not a property of the data - so as intake grows, the reviewed share is that flat count over a rising denominator and shrinks. On the attack side the goal decides the scaling. An availability attack, which wants the model measurably worse overall, does need a real share of the corpus and is genuinely diluted by growth. A targeted attack - one chosen kind of input read the wrong way, or a conditional behaviour keyed to something the adversary controls - behaves like an approximately absolute number of examples and is barely diluted at all. So corpus growth strengthens you against the crude attack, does almost nothing against the sized one, and quietly shrinks the review percentage you have been quoting as a control.

go deeper

for a junior

Know that the number of items a review can cover comes from staffed hours, so a growing intake shrinks the percentage even when the team is working just as hard as before.

for a middle

Explain both curves: the flat reviewed count over a rising denominator, and the split between a poisoning goal that needs a share of the corpus and one that needs roughly a fixed number of rows.

for a senior

Be able to spot the trend in a QA report - stable items reviewed, falling share - and say what it means for an attack sized against the sampling rate rather than against the corpus.

for a principal

Decide what to do when the arithmetic says restoring last year's share means a multiple of headcount forever, and be clear about which poisoning goal your growth is genuinely defending against.

## The defender's side: a staffing line, not a data property The number of items a human review can process is **staffed hours times items per hour**. Both factors are operational: how many reviewers are funded, and how long a careful judgment on one item takes. Neither has anything to do with how much data arrives. So the *reviewed share* - the number everybody quotes - is a flat numerator over a denominator that grows with the business. An intake that doubles halves the share with no change in review quality, no change in effort, and no incident to notice. The count of items reviewed looks completely stable in the operations report, which is precisely why the erosion is easy to miss: the work did not get worse, the world got bigger. There are only three ways to move it, and all three are unattractive. Fund a multiple of headcount, which grows with intake forever. Shorten the time spent per item, which trades coverage for exactly the per-item care the review exists to provide. Or review a chosen slice rather than the whole intake, which changes what the number means and needs a separate argument about which slice. ## The attacker's side: goal decides whether growth dilutes them This is the half candidates get wrong, and it is the whole point of the question. **Availability / degradation poisoning** aims to make the model measurably worse in general. Its effect is roughly proportional to the share of the training distribution the adversary controls, so a larger corpus really does dilute it: the same number of bad rows moves a bigger pile less. It also needs volume, which is exactly what a proportional sample is good at meeting. **Targeted poisoning** - make inputs of one specific kind come out a chosen way, or install a conditional behaviour that fires on something the adversary supplies later - behaves very differently. What matters is having enough examples for the model to fit that specific, narrow association. That is much closer to an **absolute count** than to a fraction, and it is only weakly sensitive to how much unrelated data sits alongside it. Growth from a hundred thousand items to a million does not multiply the attack's cost by ten. ## Putting the two curves together Over a few years of ordinary growth: - items reviewed: flat - share reviewed: falling - rows needed for a broad degradation attack: rising with the corpus - rows needed for a targeted attack: roughly flat - expected poisoned rows drawn by the review: falling, because it is the attack's flat count times the falling share The organisation experiences this as good news. More data, better model, same QA function, same clean review reports. The security property has moved the other way the whole time. ## Why "we review the same number of items as always" is the wrong reassurance A flat count of items reviewed says the QA function is stable. It does not say what fraction of writes had a chance of being seen. Against an adversary who sizes writes against a sampling rate, the fraction is the thing that decides how likely any of their rows is drawn - and even a drawn row still has to be recognised, which per-item review is poorly placed to do. ## The honest framing Sampled review is priced against **accidents**, whose rate is a property of the process, so a proportional sample estimates it well. It is not priced against an actor who reads the review rate and chooses a row count under it. Corpus growth is a genuine defence against the poisoning goal that needs volume, and close to no defence at all against the one that needs forty well-placed rows.

  • Where does the reviewable count actually come from?
    Staffed hours times items per hour. It is a budget and staffing line, so it moves only when someone funds more reviewers or shortens per-item review - and shortening per-item review buys coverage by spending the per-item care that is the second condition for catching anything. Nothing about the volume of arriving data changes it on its own.
  • If growth dilutes availability poisoning, is more data a defence?
    Against that goal only, and it is a real effect worth stating. A broad degradation attack has to move a bigger pile with the same rows, so its required volume rises with the corpus and its footprint becomes easier to meet with a proportional sample. It gives you nothing against a targeted attack or a conditional behaviour, whose row count barely responds to corpus size.
  • Does the adversary need to know your exact sampling rate?
    No, only its order of magnitude, which is often inferable from published QA practice, vendor contracts or headcount. Getting it wrong costs them a margin rather than the attack: writing somewhat more items than strictly needed is cheap, so a drawn item is a loss they priced in rather than a failure of the attempt.

saying these in an interview costs you the question

  • Assumes poison must be a constant fraction of the corpus
  • Claims more training data dilutes every poisoning goal
  • Confuses items reviewed with share of intake reviewed
  • Treats the review rate as a property of the dataset
  • Assumes review headcount scales with intake automatically

context

open as a page

A team spot-checks 1% of incoming labelled training data - what does that bound?

level: juniorimportance: should knowfreq 60%

basics

~20 s

Spot-checking 1% bounds how obviously wrong a typical incoming row is, not whether the corpus was poisoned. Whoever writes poisoned rows chooses how many, and can write few enough that a 1% draw almost never lands on one.

open as a page

A sampled QA review of a training-data queue rejected zero items - what does that bound?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A zero-rejection sampled review bounds the reviewed items only: each was individually plausible to one reviewer. It says nothing about the unreviewed remainder, and against a small targeted attack the review most likely drew no poisoned item at all.

open as a page

An auditor asks what share of your training corpus a human reviewed, and intake tripled while reviewer headcount did not - what do you tell them?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Give the real share, say it fell because intake grew rather than because review got worse, and state what it bounds: those items got a per-item human verdict. It is not a claim that the corpus is unpoisoned.

open as a page