skip to content

An auditor asks what share of your training corpus a human reviewed, and intake tripled while reviewer headcount did not - what do you tell them?

level: principalimportance: nice to knowfreq 24%

answer

  1. the question is which claim you will sign
  2. share, not the flat item count
  3. attach the bound before someone else infers one
  4. tripling headcount buys a linear gain
  5. who can write into the queue at all

basics

~20 s

Give the real share, say it fell because intake grew rather than because review got worse, and state what it bounds: those items got a per-item human verdict. It is not a claim that the corpus is unpoisoned.

solid answer

~50 s

Three decisions, and they are yours to own. First, report the share honestly and put the trend beside it, because the flat count of items reviewed is the number that will otherwise get quoted and it hides the erosion. Second, refuse the conversion: a coverage figure supports "this share received a per-item verdict against the guideline and nothing individually implausible was found", and it does not support a statement about the remainder or about patterns spanning items. Third, decide where the money goes. Restoring last year's share means a multiple of headcount forever, for a linear gain against someone who can re-size their writes far more cheaply than you can staff. The defensible position is usually that sampled review is a control against accidents and sloppy bulk writes, that constraining who can write into the queue is where the assurance actually comes from, and that you are willing to say so in the record.

go deeper

for a junior

Know that a review coverage number is a share of intake and that it needs a plain statement of what it covers whenever it is quoted to somebody outside the team.

for a middle

Be able to explain why the item count and the share tell opposite stories, and why a clean sampled review does not license a statement about the data that was not reviewed.

for a senior

Practise writing the supported and unsupported sentences explicitly, and being the person who attaches the bound to the number before it leaves the team.

for a principal

Own the spend decision and the public claim: whether restoring coverage is worth a permanently growing headcount for a linear gain, and what you are prepared to assert on the record when somebody can compel an answer.

## What is actually being asked The auditor's question sounds like a metric request and is really a request for a **claim you will stand behind**. Whatever number you say becomes a sentence in a document you do not control, and the sentence will be read as assurance unless you attach the bound yourself. ## Decision one: which number, and with what beside it There are two candidates and they tell opposite stories. *Items reviewed* is flat and reassuring. *Share of intake reviewed* has fallen with growth. The flat count is the one that gets quoted, because it is stable and it makes the team look consistent - which it is. Report the share, with the trend and the reason: intake grew, staffed hours did not, nothing regressed. Volunteering that framing is what stops the finding from landing on the QA team, where it does not belong, instead of on the risk register, where it does. ## Decision two: the sentence you will and will not sign Write the bound yourself rather than leaving it to be inferred: - **Supported:** a stated share of intake received an item-by-item human verdict against the labelling guideline, and nothing individually implausible was found among those items. - **Not supported:** any statement about the unreviewed remainder. The review draws a share; against writes sized against that share, the expected number of poisoned items appearing in front of a reviewer can be below one, so a clean result is equally consistent with a clean corpus and with a sized attack. - **Also not supported:** any statement about patterns spanning items. Reviewers answer a per-item question, so a pattern that exists only across a set of individually reasonable items is not in the scope of what was checked. The hardest part of this is professional rather than technical. "We cannot assert that this corpus is unpoisoned" is an uncomfortable sentence to put in an audit record, and it is the true one. Saying it early, with the supported claim beside it, is far better than having it extracted later. ## Decision three: where the next unit of assurance spend goes The obvious move is to fund review back up to last year's share. Do the arithmetic in front of whoever is asking. Tripled intake means roughly triple the reviewer hours to hold the percentage, and it means that again next time the business grows. What that buys is a linear increase in the chance of *drawing* a poisoned item - and the drawn item still has to be recognised, which per-item review is structurally poor at, because each item was written to survive exactly that judgment. You are proposing a permanently growing cost for a linear gain against an adversary who can re-size their writes for almost nothing. That is a bad trade and you should say so rather than quietly accept the budget. The better answers point elsewhere. Who can write into the training queue at all is a much smaller and more auditable population than the corpus is, and it is where the exposure actually comes from. Whether anything in the pipeline ever looks at items *together* - by source, by arrival window - is a question about the unit of review, not its rate, and it is the half that a bigger sample does not touch. And there is a real strategic call about whether this deployment faces an actor who would size writes against a published QA rate at all: for many pipelines the honest answer is that review is priced against vendor sloppiness, that this is fine, and that the assurance claim should be scoped to say so. ## What good looks like in the room You give the falling share unprompted, you attach the bound to it in writing, you decline to convert it into a cleanliness claim, and you come with a view on where assurance should actually be bought rather than a request to triple a budget. An auditor can work with that. What they cannot work with - and what tends to become the finding - is a confident percentage with no statement of what it covers.

  • The auditor asks for a single yes-or-no: is the training data clean?
    Answer no in the strict sense and then give what you do have. You can assert that a stated share got a per-item human verdict with nothing individually implausible found, that the write population is enumerated, and what changed this year. A yes would be a claim the process cannot produce, and it is the claim that becomes the finding when someone later asks how it was established.
  • Is there a case for keeping the sampled review at all?
    Yes, and say so plainly rather than seeming to dismiss the function. It measures labelling quality, catches vendor drift early, deters lazy bulk writes because those appear in proportion, and gives you the honest denominator for every claim above. The argument is not that review is worthless; it is that its value is priced against accidents and volume, and should be described that way.
  • How do you keep this off the QA team's performance review?
    Present the two rows together: hours and items reviewed identical, share of intake down because intake grew. The degradation has no defect in it and no individual is responsible for it. Framing it as a capacity-versus-growth property, on the risk register, is what stops a structural finding turning into a personnel one and losing the actual lesson.

saying these in an interview costs you the question

  • Quotes the flat count of items reviewed as unchanged assurance
  • Lets a coverage number stand as a cleanliness claim
  • Proposes tripling review headcount without pricing the gain
  • Refuses to write down what the number does not support
  • Blames the QA team for a share that fell with growth

context