skip to content

A sampled QA review of a training-data queue rejected zero items - what does that bound?

level: seniorimportance: should knowfreq 40%

answer

  1. what was drawn, and what was asked
  2. equally likely under both hypotheses
  3. the effect lives across items, not in one
  4. per-item verdicts never sum to a set claim
  5. it bounds how carefully they wrote

basics

~20 s

A zero-rejection sampled review bounds the reviewed items only: each was individually plausible to one reviewer. It says nothing about the unreviewed remainder, and against a small targeted attack the review most likely drew no poisoned item at all.

solid answer

~50 s

Read the result as a conjunction of two weak statements. First, coverage: a stated share of intake was drawn, so a zero-rejection quarter is entirely consistent with a sized attack whose expected number of drawn items was under one. Second, the unit of judgment: each reviewer was asked whether *this item* is correctly labelled and plausible, so a clean pass bounds per-item implausibility among drawn items and nothing else. A pattern that exists only across a set of items - the same narrow association reinforced by two hundred separately reasonable items - is never in any reviewer's field of view, so no number of clean per-item verdicts accumulates into a statement about the set. The two questions worth asking of the report are what share of intake was reviewed and what the reviewer was asked to decide. Neither is a rejection count.

code

text · 10 lines
text
QA review of the labelling queue
                            FY-1        FY-0
items ingested           240,000     520,000
staffed review hours       1,600       1,600
items reviewed            16,000      16,000
share of intake reviewed     6.7%        3.1%
items rejected in QA             0           0
...
review unit: one item, judged against the labelling guideline
(no cross-item or per-source comparison recorded)

go deeper

for a junior

Know that a review with no rejections tells you about the items that were looked at, and that a percentage of intake was looked at at all - not that everything else is fine.

for a middle

Explain why zero rejections is equally consistent with a clean corpus and with a sized attack, and why the expected number of poisoned items drawn is the quantity that decides it.

for a senior

Given a coverage report, say which rows support which claim, why the share fell without anyone doing anything wrong, and what you would ask about the unit of review rather than the rejection count.

for a principal

Be prepared to stop a clean QA report being used as assurance in a document that leaves the organisation, and to say what evidence would be needed instead.

## Read the result as two separate claims, both weak "Zero rejections" is a conjunction, and each half bounds much less than it sounds like. **Half one: what was drawn.** A review sees a share of intake. If the share is a few percent and a targeted attack needs a few dozen items, the expected number of poisoned items appearing in front of any reviewer is below one. A zero-rejection quarter is exactly what you would observe if nothing was wrong, and also exactly what you would observe if a sized attack had landed. An outcome that is equally likely under both hypotheses distinguishes nothing. **Half two: what was asked.** A reviewer receives one item and answers a per-item question - is this labelled correctly, is it plausible, does it match the guideline? A clean pass therefore bounds *per-item implausibility among drawn items*. It is a statement about individual items, and it is the only statement the process is structured to produce. ## The unit of judgment is the deeper limit The coverage half can, in principle, be bought with money. The second half cannot, because it is about what question is being answered. A targeted poisoning effect is a property of a **set**. It exists because a hundred or a few hundred items jointly reinforce a narrow association that would not otherwise be there. No single member of that set carries the effect. Handed any one of them, a reviewer applying the guideline gives the correct answer for that item - which is *pass* - and does so honestly and competently. The pattern lives in the relationship between items: their common source, their arrival window, their shared unusual feature combination, the way they cluster in a region of the input space that was previously sparse. So per-item verdicts never accumulate into a set-level statement, however many you collect. Ten thousand clean per-item verdicts are ten thousand answers to a question that was not about the attack. ## Reading the coverage report Given a QA line like the one above, the useful reading is: | what the row says | what it bounds | | --- | --- | | items reviewed: flat | the QA function is staffed and stable | | share reviewed: falling | the chance any given write is ever seen, falling too | | items rejected: zero | drawn items looked fine one at a time | | review unit: one item | no part of the process compares items to each other | The row that changed year on year is the share, and it changed for a reason with no defect in it: intake grew and hours did not. Nobody did anything wrong, and the property got worse. ## What a triager should do with this Stop treating the rejection count as evidence and treat it as a description of the process. Then ask the questions the report does not answer. What share of intake was reviewed, and which share is it - all sources, or the ones easiest to route? What was the reviewer asked to decide, and has anything in the pipeline ever looked at items *together* by source or arrival window? Which sources have write access at all, and did any of them change this year? None of those is a better sampling scheme; they are questions about what the existing number means. The one thing not to do is convert a zero into an assurance sentence. ## The direction of the claim A clean review bounds the per-item anomaly that whoever wrote the data was willing to leave visible - that is, how carefully they wrote - not whether they were there. Getting that direction backwards is the whole failure mode: it turns a description of your own review procedure into a claim about an adversary who read that procedure first.

  • Would a non-zero rejection count be better news?
    It is better evidence that the review is functioning, and that is worth something: a QA process that never rejects anything may not be exercising judgment at all. But rejections are dominated by ordinary labelling defects, so the count tells you about vendor quality, not about a targeted attack. A sized attack expects to lose the occasional item and writes with slack for exactly that.
  • What is the honest sentence to put in the audit record?
    Something of the form: this share of intake received a per-item human verdict against the labelling guideline, and nothing individually implausible was found among those items. Then say explicitly what it does not support - no statement about the unreviewed remainder, and no statement about patterns spanning items, because nothing in the process examines those.
  • Does the share reviewed falling from 6.7% to 3.1% represent a regression by the QA team?
    No, and saying so would misdirect the response. Items reviewed and hours are identical; intake more than doubled. The team's output is unchanged and the coverage property degraded anyway, which is why this belongs on a risk register rather than in a QA performance review.

saying these in an interview costs you the question

  • Reads zero rejections as evidence the corpus is clean
  • Ignores what share of intake was actually reviewed
  • Assumes a per-item verdict can reveal a set-level pattern
  • Treats a falling reviewed share as a QA quality problem
  • Quotes items reviewed rather than share of intake

context