skip to content

A design review asks whether a tenfold bigger training crawl lowered poisoning risk. What do you say?

level: principalimportance: nice to knowfreq 28%

answer

  1. give two answers, not one
  2. concede the half that is genuinely true
  3. the growth number is not the finding
  4. can any question be investigated afterwards?
  5. size is not a mitigation in a register

basics

~20 s

Answer per goal, not overall: growth genuinely raised the cost of degrading the model and did nothing about planting one behaviour. Then say the real finding - with no retained origin and no pinned snapshot, the question cannot be investigated at all.

solid answer

~50 s

I would refuse the single-number answer and give two. For an adversary who wants the model measurably worse, a tenfold corpus really did raise their bill roughly tenfold, and that is a legitimate improvement to record. For an adversary who wants one behaviour on one rare context, the requirement behaves as an absolute document count, so growth bought nothing and scaling capacity alongside it did not help either. Then I would put the actual finding on the table: if per-document origin was not retained and the corpus is re-crawled per run rather than pinned, we cannot investigate any suspicion after the fact and cannot diff one run's data against the next. That is the gap I would fund - retained provenance, a pinned re-fetchable snapshot, tighter admission for the slices anyone can write - and I would ask that the risk register stop carrying corpus size as a mitigation.

code

text · 9 lines
text
corpus_v3 (as presented)
  documents            1.24B      (v2: 0.11B)
  sources              public web crawl, 2 snapshots merged
  near-duplicate pass  applied
  quality filter       applied, threshold unchanged
  per-document origin  not retained after stage 2
  snapshot pinning     none (re-crawled per training run)
  ...
  poisoning assessment "risk reduced: diluted at this scale"

go deeper

for a junior

Know that the answer depends on which attacker is meant, and that a bigger corpus is not a general safety improvement.

for a middle

Be able to give the two answers with their reasons: a share-based requirement scales with the corpus, an absolute-count requirement does not.

for a senior

Show you would go past the question asked to the pipeline gap behind it - no retained origin and no pinned snapshot means no investigation is possible later.

for a principal

Own the register wording and the funding argument, including the costs of provenance, pinning and tighter admission, and be willing to say explicitly that this deployment accepts the risk if that is the right call.

## The situation You have inherited a pretraining corpus assembled from a public crawl. It grew by an order of magnitude between the last training run and the next. Somebody has written, in good faith, that this reduced data-poisoning risk. You are asked in a review whether that is true. The judgment is yours to own, and it has three parts: what you concede, what you correct, and what you ask for. ## Part one: concede the half that is true An attacker whose goal is a globally worse model has a requirement expressed as a **share** of the training signal. A tenfold corpus raises what they must write roughly tenfold, and at some size the attack stops being affordable to anyone without industrial resources. That is real risk reduction and refusing to acknowledge it makes the rest of your argument sound like reflexive negativity. Record it - specifically, against that goal. ## Part two: correct the half that is false An attacker whose goal is one behaviour on inputs they choose has a requirement that behaves as an approximately **absolute document count**, because the planted content competes with the other evidence about that same rare context rather than with the corpus at large. Corpus growth adds documents about other things. Growing model capacity alongside the corpus, which is what happened, makes rare patterns easier to fit rather than harder. So one question - "did growth make us safer?" - has two answers pointing in opposite directions, and any risk statement that reports a single verdict is wrong for one of them. This is the substance of the review contribution. ## Part three: name the finding nobody wrote down The important thing in a review of this kind is usually not the growth number at all. It is what the pipeline cannot tell you: - If **per-document origin was not retained**, then no question asked after the fact is answerable. Not "where did this come from", not "what else came from the same place", not "was this present in the previous run". - If the corpus is **re-crawled per training run rather than pinned**, then two runs cannot be diffed, and a corpus changing under you is invisible by construction. Those two gaps are worth more than any filter tuning, because they are the difference between a suspicion you can investigate and one you can only argue about. State plainly that today the honest answer to "was our training data poisoned?" is "we have no way to find out" - and that this is a decision the organisation made by omission, not a technical impossibility. ## What you ask to fund, and what you refuse to promise Fund, in order: 1. **Retained provenance per document**, carried through every pipeline stage. 2. **A pinned snapshot per training run**, stored so it is re-fetchable and diffable. 3. **Tighter admission for the slices anyone can write into**, which is where near-free write access lives. This is a scoping decision, and it is cheaper than any control applied afterwards. 4. **Behavioural evaluation on the contexts whose corruption would be expensive**, since you can test a model on what matters and you cannot read a billion documents. Refuse to promise: that the corpus is clean, that filtering established it, or that any amount of growth will. Refuse to let dataset size stand in a risk register as a mitigation, because it mitigates one goal and reads as if it mitigates the threat. ## The costs you should surface rather than hide Every item on that list has a price, and a principal-level answer says who pays it. Provenance and pinning cost storage and pipeline complexity, and they slow the crawl-to-run loop. Tightening admission removes real content along with the risk, and the loss lands on sparse topics rather than on the average. Behavioural evaluation costs the time of people who must decide *which* behaviours matter, which is a product question and not an engineering one. None of these is free, and presenting them as free is how they get cut in the next planning round. ## The line to actually say "Growth raised the cost of degrading the model roughly in proportion and did not change the cost of planting one behaviour. The risk register should carry those separately, and it should stop listing corpus size as a mitigation. What I want funded is retained per-document origin and a pinned snapshot, because right now, if someone reports that our model behaves oddly on some rare input, we cannot investigate it at all."

  • The review pushes back that provenance and snapshot storage are too expensive. How do you argue it?
    By pricing the alternative. Without retained origin and a pinned snapshot, the answer to any future report about odd model behaviour is 'we cannot find out', which converts every suspicion into an unbounded investigation or an unexamined risk. I would rather argue for the smallest version - origin retained, one snapshot pinned per run - than lose the capability entirely to a cost objection.
  • Is there a case for deciding this deployment simply does not face this adversary?
    Yes, and it is a legitimate call to make explicitly. If the model is internal, the outputs are reviewed by people, and no rare context would be expensive to corrupt, the honest decision is to accept the risk and spend elsewhere. What is not acceptable is reaching that conclusion by accident, via a dataset-size line nobody examined.
  • How would you word the risk register entry?
    As two entries. One for indiscriminate degradation, marked reduced by corpus growth with the reasoning stated. One for a targeted behaviour on a chosen context, marked unchanged by growth, with the mitigation listed as provenance, snapshot pinning and behavioural evaluation. One combined entry cannot be true for both.

saying these in an interview costs you the question

  • Gives one verdict for two attacker goals
  • Leaves corpus size listed as a mitigation
  • Treats filtering results as corpus assurance
  • Never says the pipeline cannot be investigated after the fact
  • Proposes controls without naming who absorbs their cost

context