Reviewers sample collected training rows before each retrain — why is that not a bound?
answer
- compare what each side scales with
- hours funded versus accounts times time
- a sample bounds the sample
- honest labels leave nothing to flag
- ask about concentration, not rejections
basics
~20 sBecause review capacity is fixed reviewer-hours and therefore a sampled fraction, while the contribution scales with accounts and elapsed time and repeats every cycle. A clean review bounds what was in the sample, not what is in the batch.
solid answer
~50 sHuman review is a rate, not a gate. Reviewers can look at whatever their hours allow, which in a weekly batch of a million collected rows is a fraction of a percent — so a clean result says the sampled rows contained nothing a reviewer recognised as bad, and nothing about the rest. Meanwhile the other side scales with the number of accounts and the length of the window, and gets a fresh attempt at every retrain. Worse, the rows in question need not look wrong: when the label is generated by the user's own behaviour, genuinely accepted content is genuinely labelled accepted, so there is no inconsistency for a reviewer to spot. The design question is therefore what share of a batch any one uncontrolled channel is allowed to supply, and how concentrated contributions are — not whether anybody looks at samples.
code
text · 9 linescollection window 2026-03-02 .. 2026-03-09 (retrain: weekly)
accepted-completion rows collected 1,284,000
rows sampled for human review 12,400 (0.97%)
rows rejected by review 310
distinct seats contributing 41,900
rows from top 1% of seats (419 seats) 218,000 (17.0%)
rows from seats opened this window 96,300 (7.5%)
...
review verdict: PASSgo deeper
Know that reviewing a sample of collected training data tells you about the sample only, and that a clean result is not a statement about the whole batch.
Explain the asymmetry precisely: review is bounded by funded hours while contribution grows with accounts and window length, and repeats each cycle.
Demonstrate you would replace the review question with composition measurements — per-account share of a batch, share of a narrow context slice, share from new accounts — and know why per-row inspection is the wrong unit.
Be ready to decline the headcount ask and redirect the spend into collection design, and to say what you will report to leadership in place of a review pass rate.
## The answer that sounds like a control and is not Ask a team how they handle poisoning of a product that trains on its users, and the common answer is "we review the data before it is used." It is worth taking seriously, because review is genuinely useful — and then showing exactly what it does and does not establish. ## Review is bounded by hours; the other side is bounded by accounts and time Review capacity is a headcount multiplied by a throughput: how many rows a person can meaningfully judge per hour, times hours funded. That is a fixed number per week. Divide it by the batch size and you get a sampled fraction, and in any product with real traffic that fraction is small. The contribution side scales differently. It grows with the number of identities the contributor holds — a term priced in seats, not in effort — and with the length of the retrain window, and it resets at every retrain. Two quantities growing on different axes do not meet: doubling reviewer headcount doubles the sampled fraction once, while the other side is multiplied by whatever the account cost allows. This asymmetry is the substance of the answer. Anyone can say "review does not scale"; the credible version names *what* each side scales with. ## What a clean review actually establishes Be precise about the direction of the claim, because this is where candidates get marked down. A clean review establishes: - the **sampled** rows contained nothing the reviewers **recognised** as bad, - under whatever criteria they were given, - for **this** batch only. It does not establish that the batch is free of rows placed to change a behaviour, that the retrained model will behave like the last one, or that contributions were spread evenly across accounts. Each of those is a different measurement. ## The part that has nothing to flag Even a hypothetical full review runs into a harder problem. Reviewers catch rows whose label contradicts their content — that is what a labelling queue is good at. But when the product derives the label from behaviour, no contradiction exists. Content that was genuinely produced and genuinely accepted is genuinely labelled as accepted. Each row is individually unobjectionable; what carries the effect is which rows there are and how many, not any single one being wrong. This is the general shape of correctly-labelled contributions: there is no anomaly at the row level to detect, so a per-row inspection process is looking at the wrong unit. ## The measurements that actually say something If review is a rate rather than a bound, what does bound it? Quantities about the *composition* of the batch rather than the quality of individual rows: - what share of the batch came from any single account, and from the top handful of accounts; - what share of a **narrow context slice** — the rows touching the behaviour you care about — came from one channel; - how much of the batch came from accounts younger than one retrain interval; - whether contribution concentration changed between batches. Those are questions a data platform can answer, and each of them has an obvious design lever behind it: a cap on how much any identity may supply to one batch, a floor on how many distinct sources a context slice must draw from, and a cost attached to opening an account. None of that is a scanner; it is collection design. ## How to deliver this in an interview Do not dismiss review — say what it buys. It catches sloppy, obvious and accidental contamination, it gives you labelled examples of what bad looks like, and it is how you learn the criteria in the first place. Then state the limit plainly: it covers a sampled fraction, it is priced in hours, and it is blind by construction to rows whose labels are honest. Finish on the design question, which is the one the interviewer wants: *what fraction of a batch may one uncontrolled channel supply, and do we measure it?*
- A weekly collection report shows a 0.97% review sample and a PASS verdict. What do you ask next?How concentrated the batch is: what share came from the top handful of accounts, and what share of the narrow context slices came from one source. Also how many rows came from accounts younger than one retrain interval. The rejection count tells you about reviewer criteria; the concentration figures tell you whether any single channel could have supplied enough to matter.
- So is human review worthless here?No — it catches sloppy and accidental contamination, produces labelled examples of what bad looks like, and is how the criteria get written at all. It is worth funding. It is simply not a bound, because it covers a sampled fraction and is blind to rows whose labels are honest. Report it as coverage, never as assurance.
- Why does doubling the review team not change the picture much?It doubles the sampled fraction once, from small to slightly less small. The other side is multiplied by however many accounts the contributor can afford and by how many retrains happen per quarter. Linear spend against a term that scales with money and time is a losing trade, which is why the lever belongs in collection design instead.
saying these in an interview costs you the question
- Treats a clean review as evidence the batch is clean
- Says more reviewers would close the gap
- Assumes poisoned rows always look wrong to a human
- Quotes the rejection count as the safety metric
- Ignores that the review must be repeated every retrain