skip to content

Learning From Its Users

Production input folded back into training makes an outsider's submissions training data by design, and the retrain cadence is the window. Interviewers ask because sampled review will not scale.

on this pageshow

explore

questions

4

Why is a code assistant that retrains on accepted completions a poisoning surface with no breach?

level: juniorimportance: must knowfreq 55%

answer

  1. no breach is involved anywhere
  2. the product collects this on purpose
  3. the ordinary interface is the write channel
  4. accounts, throughput, and the retrain interval

basics

~20 s

Because the product collects its own users' behaviour as training data on purpose. An ordinary paying seat's accepted completions become rows in the next retrain, so an outsider writes into the training set through the normal interface, without touching any system.

solid answer

~50 s

Poisoning needs write access to training data, and this product hands that out as a feature. A code-completion assistant that fine-tunes weekly on completions its users accepted has turned the ordinary usage path into a write channel: whatever an enrolled seat produces and accepts is collected, labelled by the user's own behaviour, and folded into the next batch. The attacker needs no credential, no gradient, no view of the weights, and no defect to exploit — only accounts and time. That is what separates this from evasion, where the adversary crafts an input for a finished model and the model is unchanged afterwards. Here the effect lands in the weights, so it applies to every user who hits the same context, and it renews at each retrain. The real bound is not "can they get in" but how much of a batch that uncontrolled channel supplies.

go deeper

for a junior

Be ready to say, in one sentence, that a product which trains on user activity has given ordinary users a write channel into its training data, and that no break-in is required.

for a middle

Explain the mechanics: which interactions are collected, what generates the label, and why an effect that lands in the weights reaches every user rather than one request.

for a senior

Show you would ask for the collection rule and the retrain cadence before assessing anything, and that you frame the risk as a share of a batch rather than as a yes-or-no access question.

for a principal

Own the design tension: the collection loop is the product's quality engine, so the call is how much of a batch one uncontrolled channel may supply, not whether to keep learning from users.

## The claim being made A product that learns from its users has an open write channel into its own training data. That is not a bug in the collection pipeline; it is the collection pipeline working as designed. Take a concrete setting: a code-completion assistant sold as paid seats, which fine-tunes on a weekly cadence using telemetry about which suggested completions developers accepted and kept. The design intent is obvious and good — the product gets better at the code its customers actually write. The security consequence is that an outside party who buys a seat is, by design, an author of next week's training rows. ## Why "no one breached anything" is the wrong frame Engineers reach first for the access-control question: who can write to the training bucket, who can push to the feature store, who signed the dataset. Those questions matter, and they are answered by other controls. They are not the question here, because the adversary in this scenario never touches any of those things. They use the product exactly as a legitimate customer does. Every row they contribute arrives through the front door, correctly authenticated, correctly billed. So the threat model shifts. The interesting quantity is not "is the pipeline protected" but "what fraction of a training batch can one uncontrolled channel supply, and how often does that channel get another turn?" ## How this differs from attacking the deployed model It is worth being precise about the two families, because interviewers use exactly this contrast to sort candidates. - **Evasion** happens at inference. The adversary sends a crafted input to a model whose weights are already fixed, and gets that one input read the wrong way. The model is identical afterwards. Nobody else is affected. - **Poisoning** happens at training. The adversary influences what the model learns. The change lives in the weights, so it is present for every user who reaches the affected behaviour, and it persists until something retrains it away. A product that learns from its users is a poisoning surface specifically because it grants ordinary users a training-time position. It is also worth separating this from a system that degrades by consuming its own outputs — a feed retrained only on what it previously showed narrows over time — because there is no actor in that story. Take the adversary out of this one and there is nothing left to talk about but data hygiene. That presence of an outside party who *chooses* what enters the batch is the whole subject. ## What the adversary actually needs Not much, and that is the point: - **Standing:** one or more ordinary accounts. If seats are sold, the cost of participation is a price list entry. - **Knowledge of the collection rule:** what gets logged, and what the product treats as a positive label. In this example the label is generated by the user's own behaviour — a completion that was accepted and left in place looks like a good example to the collector. - **Patience:** the loop closes on a schedule. Nothing lands until the next retrain. Notably absent: the weights, the architecture, gradients, any score the endpoint returns, any privileged read. This is a training-time position held by someone with a purely black-box relationship to the model. ## Why the label matters as much as the content A reviewable poisoning attempt is one where a row's label is obviously wrong for its content — that is something a labelling queue can flag. When the product derives the label from behaviour, the attacker does not need to lie about anything. Content that is genuinely accepted is genuinely labelled accepted. There is no inconsistency for anyone to notice, which is why "we look at the data" is a much weaker statement here than it sounds. ## What to say in an interview State the mechanism in one line — the interface is the write channel — then name the limit the adversary works under rather than the attack: throughput per account, number of accounts they can hold, and the retrain interval. Naming the limit is what turns this from a scary story into a design conversation, because every one of those three terms is something the product team can actually change.

  • How is this different from a model that degrades because it keeps training on its own outputs?
    That story has no outside party in it. A system consuming its own logged output drifts as a data-quality and feedback problem, and it would happen with no attacker present. Here an outsider decides what enters the batch and what behaviour they want back, which makes the size of their contribution — not the loop's existence — the thing you have to bound.
  • Does the attacker need to know anything about the model itself?
    No. They need no weights, no gradients, and no scores. What they need is knowledge of the collection rule: which interactions get logged, and what the product treats as a positive label. That is usually inferable from the product's own documentation and observable behaviour, which is why the black-box position costs so little here.
  • Why does the effect persist rather than being a one-off?
    Because it lands in the weights rather than in a single request. Once a behaviour is learned, every user reaching that context sees it, and the next retrain gives the same channel another turn. The honest description is a recurring window on a schedule, not a single event you can close.

A suggestion box that is emptied straight into the policy manual each week. Nobody has to break into the office; they just have to keep posting, and to post more than everyone else.

saying these in an interview costs you the question

  • Says poisoning requires breaching the training pipeline
  • Calls this prompt injection or a jailbreak
  • Assumes the attacker needs weights or gradients
  • Treats it as identical to crafting one bad input at inference
  • Says access control on the data store closes it

context

open as a page

In a weekly-retrained code assistant, what bounds how much one account can contribute?

level: middleimportance: should knowfreq 42%

basics

~20 s

Three quantities multiply: accepted items per account per day, how many accounts the contributor can hold and afford, and how long the retrain interval is. The product of those is their share of one batch, and it renews every cycle.

open as a page

Reviewers sample collected training rows before each retrain — why is that not a bound?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because review capacity is fixed reviewer-hours and therefore a sampled fraction, while the contribution scales with accounts and elapsed time and repeats every cycle. A clean review bounds what was in the sample, not what is in the batch.

open as a page

What can you honestly promise about a code assistant that retrains weekly on user activity?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A statement about the model that shipped, not about the one shipping next week. Because every retrain re-opens the same channel, the honest answer is a trajectory with named limits on contribution, not a status of clean or safe.

open as a page