Candidate generation is nearly free for an extraction attacker — so what actually binds the run?
answer
- The cheap stage is not the constraint
- Every flagged candidate costs a human
- The threshold is priced against review capacity
- Precision per confirmed hit, not volume
- Recall has no denominator here
basics
~20 sThe filter binds, not the generation. What limits the run is the false-positive rate at the chosen score threshold and the human verification effort each flagged candidate consumes, since only a holder of the genuine source can confirm a hit.
solid answer
~50 sThroughput is not the constraint; precision is. Producing candidate spans costs almost nothing, so the pipeline can always make more, and the scarce resource is downstream: every candidate the score filter flags must be checked against the genuine artefact by someone who holds it, and that check is slow and human. Lower the threshold and you flag ten times more candidates, precision falls and the verification cost per confirmed hit rises; raise it and you confirm fewer things you could have proved. The run is priced in verification, so the threshold is chosen against the verification budget rather than against the candidate pool. Two other limits bite. Recall has no denominator, so you can never say what you missed. And if the endpoint stops returning per-token log-probabilities, the reference comparison loses its input and the shortlist collapses back into unranked generated text.
go deeper
Remember the inversion: making candidate text is cheap, and telling real recall from invention is the expensive part. That alone explains why extraction runs are judged on precision.
Explain the pipeline shape and what the threshold trades: a lower cutoff flags more candidates, drops precision and raises verification cost per confirmed hit. Be able to say why the score filter exists at all — to make the human verification stage affordable.
Show you would size the run from review capacity backwards, name who holds the verification oracle, and refuse to quote recall. Be ready to say what withholding confidence scores does and does not change about the underlying exposure.
Frame it as spend: what precision the exercise must reach to be worth funding, what access the verifier needs, and what removing a product feature like log-probabilities costs legitimate users against the attacker cost it imposes.
## The cost structure is upside-down from the intuition Asked what limits a training-data extraction campaign, most people reach for compute: how many candidate spans can you generate. That is the cheap half. Generation against a paid endpoint costs a small amount per span and parallelises trivially, so the candidate pool can be made as large as anyone wants. The expensive half is deciding which candidates are real. A well-run extraction pipeline therefore has a very particular shape: an enormous, nearly free generation stage; a scoring stage that ranks candidates by how much better the target fits them than an independent reference does; and a small, slow, expensive verification stage where a human with access to the genuine source confirms a shortlist. The middle stage exists only to make the last stage affordable. ## The threshold is a price, not a hyperparameter Everything about the run is set by where the score threshold sits. - **Lower it.** More candidates are flagged. Precision falls, because the additional candidates are drawn from the part of the distribution where the target's advantage over the reference is small. Each confirmed hit now costs more verification time, since the reviewer wades through more misses per success. - **Raise it.** The shortlist shrinks and precision climbs, but genuine hits with a modest confidence gap are discarded unseen and never appear in the report. The run is priced in verification, so the threshold is set from the verification budget backwards: how many candidates can a reviewer with archive access actually check, and what precision do they need for that to be worth their time. Choosing it from the candidate pool forward — flag the top ten thousand because we have ten thousand — is how these exercises turn into an unreviewable pile. ## Who can verify is part of the limit Verification requires the genuine artefact. That means the run's ceiling depends on the chair you are sitting in. An internal reviewer or an auditor granted archive access can confirm candidates and produce findings. An outside party holding no copy of the source can rank candidates but confirm none, so their strongest honest deliverable is a ranked hypothesis list with the confidence gap attached and no claim of leakage. Failing to say which of those you produced is the most common defect in an extraction write-up. ## Recall has no denominator The pipeline can tell you what it confirmed. It cannot tell you what fraction of the archive is recoverable, because it never enumerated the archive's presence in the corpus and never exhausted the generation space. A run that confirms forty spans confirms forty spans; it does not measure a leak rate, and a document with no confirmed hit has not been shown absent. Anyone reading the output as a coverage percentage is reading something the run never measured. ## The vantage can be taken away, at a price The reference comparison needs the target's own confidence on the candidate. If the endpoint stops returning per-token log-probabilities, that input disappears, and the adversary is left ranking on the reference alone — which ranks how intrinsically likely the text is, precisely the wrong quantity. The shortlist collapses back into unranked generated text and the verification stage becomes unaffordable. That is worth stating carefully, because it is a cost control rather than a boundary. Removing scores does not make anything unmemorized and does not make text unrecoverable; it removes one cheap discriminator and forces the adversary onto more expensive ones, which for a campaign that lives or dies on precision is a substantial bill. It is also not free to the model owner: log-probabilities are a legitimate product feature for evaluation and calibration work, so removing them is a trade against real users, not a costless hardening step. ## How to answer this in an interview State the inversion first — generation is free, discrimination is not — then name the two quantities that actually bind: the filter's false-positive rate at the chosen threshold, and verification effort per confirmed hit. Then add the two limits that shape what may be claimed: recall is unmeasured, and verification requires whoever holds the source. Someone who answers in terms of GPU hours or rate limits has not understood where the work is.
- You have review capacity for 200 candidates. How does that change how you run the pipeline?It fixes the threshold. You set the score cutoff so that roughly 200 candidates survive, and you spend your remaining effort raising precision inside that slice rather than generating more text: a better reference model, several references, or holding capability constant with perturbed variants of the same span. Generating more candidates without moving the threshold just deepens a pile nobody will read. The report then states the threshold, the reference and the confirmed count, not the pool size.
- The endpoint stops returning log-probabilities. Is the risk gone?No. Memorization is a property of the weights, fixed at training, and nothing about it changed. What changed is the price of the adversary's cheapest discriminator: without the target's scores the reference comparison has no input, so candidates come back unranked and verification becomes unaffordable at scale. Treat it as a cost control that raises the bill, not a boundary that removes the exposure — and price it against the legitimate evaluation uses of those scores.
- Why can this run never report a leak rate for the archive?Because there is no denominator. The pipeline sampled a generation space it did not exhaust and verified a shortlist it did not enumerate, so it measures confirmed hits and nothing about coverage. A document with no confirmed hit was not shown to be absent; it was not found by this run at this threshold. Reporting the confirmed count with the threshold and reference attached, and explicitly declining to state recall, is the honest shape.
saying these in an interview costs you the question
- Names compute or rate limits as the binding constraint
- Sets the threshold from the candidate pool rather than review capacity
- Assumes anyone can verify a candidate span
- Reports confirmed hits as a leak percentage
- Treats withheld log-probabilities as eliminating the risk
- Believes more candidates always means more findings