skip to content

A labelling QA pass re-checks every training label against its sample — which poisoning does it miss?

level: middleimportance: must knowfreq 64%

answer

  1. ask what predicate the control evaluates
  2. the predicate is satisfied honestly
  3. consistency is not integrity
  4. one row cannot show a distribution

basics

~10 s

It stops label-flipping, where the label contradicts a sample anyone can re-verify. It misses clean-label poisoning: those rows carry correct labels, so the predicate QA evaluates is true for every one.

solid answer

~50 s

The control is real, and it is worth naming what it bounds. It bounds rows whose label contradicts their content, which is the whole of the flipping build — on a corpus like a malware sample feed, where an analyst can re-open a file and re-derive its family, that contradiction is recoverable and the control catches it. It bounds nothing else. A clean-label contributor submits benign files that really are benign, labelled benign; the QA predicate is satisfied honestly on every row, so there is nothing for a reviewer to flag. Two further limits are worth stating: where ground truth is subjective rather than re-derivable, even flipped labels become arguable; and QA is a statement about the label column, never about what the corpus's feature distribution now encodes. Reporting it as "poisoning is covered" overstates it by one whole attack family.

go deeper

for a junior

Know that a label check compares the label with the sample, and that poison whose label is genuinely correct passes it. Be able to say that in one sentence without hedging.

for a middle

Explain the mechanics: the predicate the control evaluates, why a clean-label row satisfies it honestly, and why examining rows one at a time cannot reveal an effect that only exists across a set.

for a senior

Show you would scope the control in writing — which adversary it covers and which it does not — and that you notice when subjective labels shrink even its intended coverage.

for a principal

Be ready to argue against reporting a bound on one attack family as coverage of the class, and to say what evidence you would want instead and what it would cost.

## What the control actually asserts A labelling QA pass asks one question of each row: *does this label match this sample?* Everything it can bound follows from that predicate, and nothing else does. Being precise about the predicate is the entire skill being tested here, because the sentence an interviewer is listening for — "we have labelling QA, so poisoning is covered" — is a category error, not a small overstatement. ## What it bounds well It bounds label-flipping. In that build the contributed sample is ordinary and the label is a lie: on a classifier retrained from a vendor sample-sharing feed, a file that belongs to a known malware family submitted as benign. The row is internally contradictory, and the contradiction is *recoverable* — an analyst can re-open the binary and re-derive what it is, independently of what the submitter claimed. That is a genuine and useful property of this domain, and it is what makes the cheap build the reviewable one. An adversary who has only a submission channel and no model knowledge is limited to this build, and against them the control does real work. ## What it does not bound, and why It does not bound clean-label poisoning, and the reason is not that the control is weak or badly run. It is that the attack was constructed to satisfy the predicate. The contributed rows are genuinely what their labels say: benign files, labelled benign, submitted by a party who is entitled to submit files. What was chosen is *which* rows — content that, in the feature space the model reads, sits where it will pull the learned boundary. A reviewer re-deriving the label gets the same label back. There is no anomaly in the row, because nothing about the row is wrong. Perfect QA and zero QA give the identical verdict on every one of those rows. This is the difference between a **consistency** control and an **integrity** control over the corpus. Consistency asks whether the label agrees with the sample. Integrity would ask what training on this corpus produces — a different question, answered by different evidence, and not answered by inspecting rows one at a time at all. ## Two further limits worth naming **Recoverable ground truth is a property of the domain, not of the control.** A binary can be re-analysed and its family re-derived. A row labelled "abusive", "relevant", or "high risk" often cannot be re-derived to a single answer, and a flipped label there is indistinguishable from a defensible judgement call. On such corpora the control bounds noticeably less than the flipping family. **A per-row control cannot see a distributional effect.** Each poisoned row is unremarkable on its own; the effect exists only in what the set of them does to the fit. A process that examines rows individually is structurally unable to observe the thing that makes the attack work, however well it examines them. ## What to say instead The honest report is scoped: *this control catches rows whose label contradicts their content, in a domain where content re-derives to a label. It does not address correctly-labelled rows contributed to move the boundary.* That is a useful statement — it tells a reader which adversary you are covered against (one with a submission channel and no model knowledge) and which you are not (one who also holds a stand-in model to predict what a contributed row does). The direction of the claim matters. A clean QA pass is evidence that *the flipping build was not used or did not get through*. It is not evidence that the corpus is clean, and it is not evidence about the trained model's behaviour. Reporting a bound on one attack family as coverage of the class is the specific defect this question exists to catch, and it is the most common one in real reviews because the control is genuinely good at its own job. ## The follow-on an interviewer usually asks "Then what would you look at instead?" The honest answer is that no single control replaces it, and that the alternatives all price differently: they cost model knowledge, retrain compute, or behavioural testing against inputs you care about, rather than a per-row read. What you should *not* do is claim a control you have and then quietly widen its scope in the write-up.

  • Does the control bound the flipping family completely?
    Only where ground truth is re-derivable. On binaries an analyst can re-open the file and recover its family independently of the submitted label, so a contradiction is visible. On subjective labels — abusive, relevant, high risk — a flipped label reads as a defensible judgement call, and the control degrades to catching only the obvious ones.
  • How would you write up this control honestly in a review?
    Scope it to the family it bounds: it catches rows whose label contradicts content that can be re-derived, and it does not address correctly-labelled rows contributed to move a boundary. That tells a reader which adversary you are covered against — one with a submission channel and no model knowledge — and names the one you are not.

saying these in an interview costs you the question

  • Reports label QA as covering poisoning generally
  • Thinks a careful enough reviewer would spot clean-label rows
  • Confuses a consistency check with a corpus integrity claim
  • Assumes ground truth is always re-derivable from the sample
  • Expects a per-row review to reveal a distributional effect

context