In a weekly-retrained code assistant, what bounds how much one account can contribute?
answer
- write it as a limit, not a story
- a rate, an identity count, a cadence
- the interval is both window and repeat rate
- then ask: a share of which slice?
basics
~20 sThree quantities multiply: accepted items per account per day, how many accounts the contributor can hold and afford, and how long the retrain interval is. The product of those is their share of one batch, and it renews every cycle.
solid answer
~50 sExpress the limit as a rate times a count times a cadence. The rate is how many collected items one seat can produce per day, which the interface itself caps — a developer can only accept so many completions. The count is how many seats the contributor holds, bounded by whatever an account costs and by any identity checks. The cadence is the retrain interval: a weekly loop lets a rate accumulate for seven days before it lands, and hands out another window immediately after. Multiply and you get rows contributed per batch; divide by batch size and you get a share. The important refinement is the denominator: a share of the whole batch understates things badly, because the relevant slice is the part of the batch that covers the specific context the contributor cares about, and that slice can be tiny.
go deeper
Know that the amount someone can contribute is limited by how fast the product collects from one account and by how many accounts they have, not by anything security-specific.
Be able to write the limit as a rate times an identity count times a retrain interval, and to say why the interval is both an accumulation window and a repeat rate.
Show you would ask for batch composition before quoting any share, because a per-context slice is the denominator that decides whether a contribution is negligible.
Frame the three terms as collection-design levers with costs attached, and be ready to say which one your product can move without damaging the feature the loop exists to serve.
## Turning a scary story into a number "Users can poison the model" is not yet a threat model. It becomes one when you state the limit the contributor works under. For a product that folds collected traffic back into training, that limit is naturally written as **a rate, times a count of identities, times a cadence**. Work it through for a code-completion assistant that fine-tunes weekly on accepted-completion telemetry sold as paid seats. ### The rate — what one identity can emit per day The product's own interface is the throttle. A seat produces collected rows only when a developer works: suggestions are shown, some are accepted, and only accepted-and-kept ones become training examples. That gives a ceiling per seat per day that is set by the interaction design, not by any security control. It is a real bound and it is worth measuring, because it is the only term that is free — you did not have to build anything to get it. It is also the term most often overestimated in the defender's favour. Automation of the client, or simply a heavy usage pattern, can sit far above a typical developer's rate without looking anomalous in any single day. ### The count — how many identities they can hold If seats are sold, this term has a price. A hundred seats is a purchase order, not an exploit. If sign-up is free, the term is bounded only by whatever identity friction exists. Either way the honest way to write it is as a cost: *what does it cost to multiply the rate by N?* An interviewer is listening for you to convert the count into money or effort rather than treating it as unbounded. ### The cadence — how often the loop closes The retrain interval does two things at once, and people usually see only the first. 1. It sets the **accumulation window**: a weekly loop lets seven days of rate pile up into one batch. 2. It sets the **repeat rate**: every retrain is a fresh attempt with a fresh batch. A monthly cadence gives a bigger window but fewer of them; a daily cadence gives smaller windows but twelve times as many chances per quarter. That second effect is why the honest description of this surface is a recurring window rather than an event. ## The denominator problem Multiplying rate by count by cadence gives rows per batch. To say whether that matters you divide by something — and the choice of denominator is where most analyses go wrong. Dividing by the whole batch is the flattering choice. A few thousand rows against a million looks like a rounding error, and for a goal of making the model generally worse it more or less is: that kind of degradation genuinely dilutes as the corpus grows. But a contributor who wants a specific behaviour in a specific context is not competing with the whole batch. They are competing with the part of the batch that touches that context — a particular library idiom, a particular function name, a particular file pattern. That slice may be a few hundred rows in a corpus of a million. Against that denominator, an ordinary seat's weekly output is not a rounding error at all. So the sizing question is always: *share of which slice?* Answering it requires knowing the composition of the batch, which is a data question the platform team can actually answer, and usually has not. ## What the arithmetic is for Each of the three terms is a design lever, which is the practical payoff of writing the limit this way: | Term | What changes it | | --- | --- | | rate per identity | how much any one account may contribute to a batch, regardless of activity | | identity count | what an account costs, and what it takes to open one | | cadence | how long a window is, and how many windows there are per quarter | None of these is free, and none of them is a security feature bolted on afterwards — they are collection-design decisions that were already made implicitly when the loop was built. ## What to avoid saying Do not answer with an attack recipe, and do not answer with "it depends." The expected answer is the three quantities and the denominator question, plus the observation that the whole thing renews on a schedule. If you can also say which of the three terms is cheapest for your product to change, you are answering as an engineer who owns the system rather than one describing a paper.
- Does shortening the retrain interval from weekly to daily reduce the exposure?It cuts the accumulation window per batch but multiplies the number of windows. Less can land in any one model, and a bad batch ages out faster, but the channel gets seven times as many turns per week. Whether that is better depends on whether you can act between retrains at all; if you cannot, faster cadence mostly means faster propagation.
- Why is a share of the whole batch the wrong number to quote?Because a contributor aiming at one context competes only with the rows covering that context, which can be a few hundred out of a million. A whole-batch share makes any single channel look negligible and hides concentration. Quote the share of the slice you actually care about, and say how you computed the slice.
- What does per-seat throughput actually depend on in a product like this?The interaction design: how many suggestions are shown, what counts as acceptance, whether acceptance is confirmed by later retention, and whether the client can be driven programmatically. None of that was chosen as a security control, which is why measuring it usually produces a higher number than the team expects.
saying these in an interview costs you the question
- Treats the contribution as unbounded rather than priced
- Quotes a share of the whole corpus and calls it negligible
- Ignores that every retrain re-opens the window
- Assumes seat limits are enforced when they are only billed
- Confuses the retrain interval with a rate limit