skip to content

An attacker can write to your training corpus but cannot query the model — what does that buy them?

level: juniorimportance: must knowfreq 72%

answer

  1. two surfaces, not one
  2. the model was built from writes
  3. no request ever reaches the gateway
  4. authorization upstream is usually weaker

basics

~20 s

A durable change in the model itself. Everything the model knows arrived through writable paths, so an endpoint's authentication and rate limits bound only the inference-time surface, not what a later training run reads and turns into weights.

solid answer

~50 s

There are two different surfaces. The inference-time surface is what the endpoint exposes: requests, and whatever the reply carries. The train-time surface is everything a training run reads — the corpus, the label or disposition queue, an ingested third-party feed, the job that assembles the training table. An adversary holding a write on any of those needs no queries at all. Their change lands in the weights, so it applies to every caller, persists across restarts, and arrives through an ordinary authorized workflow: no request-level control could have refused it and no request log records it. That is why "it is behind an authenticated API" is not a threat model. The endpoint bounds who can *ask* the model something; it says nothing about who could shape what it answers. Ask instead who holds write upstream, and whether anyone would notice.

go deeper

for a junior

Be ready to name the two surfaces out loud: what the endpoint exposes, and what a training run reads. Knowing that a write upstream needs no queries at all is most of the answer at this level.

for a middle

Expect to walk the retraining loop and say which steps accept writes and who holds them, and why a change there is durable in a way that a single crafted request never is.

for a senior

Show that endpoint telemetry is the wrong evidence for this class, and say what evidence would bound it: the writer set on each upstream path, what reviews a change before a training run reads it, and how long a write sits unexamined.

for a principal

Own the framing when a design review calls the endpoint the trust boundary. Argue that every write path feeding a retrain is a production authorization surface, and be ready to say what you would fund to shrink it and what that costs in freshness.

## Two surfaces, and only one of them is the endpoint A trained model in production has two attack surfaces that have almost nothing in common, and confusing them is the most common scoping error in an AI threat model. The **inference-time surface** is the one everybody draws. A request arrives at an endpoint, the model returns something, and the controls sit on that path: authentication, authorization, rate limiting, input validation, request logging. An adversary working here must send requests. Their work is metered, priced and recorded, and what they can achieve is bounded by what the reply gives back and how many requests they can afford. The **train-time surface** is everything a training run reads. For a model that is periodically retrained, that includes the corpus of examples; the labels or dispositions attached to them, wherever those come from; any third-party or community data the pipeline ingests; and the job that assembles the final training table out of those pieces. Each of these is a write path, and each has its own set of principals who hold write on it. An adversary on the train-time surface sends no requests at all. They write one artefact that a later training run reads. Delete the adversary from that picture and you are left with an ordinary data pipeline; add them back and you have an access class that most threat models never wrote down. ## Why a write upstream is worth more than a query Four properties make this vantage distinctive, and they are the ones to say out loud. **It is durable.** A crafted request affects one prediction. A write that a training run consumes becomes a property of the weights: it applies to every caller, it survives restarts and redeploys, and it is still there after the incident that first drew attention to it has been closed. **No request-level control could have refused it.** The pipeline read exactly what it was designed to read, from exactly the place it was designed to read it. There is no malformed input to reject and no rate to limit. The workflow behaved correctly. **No request-level log records it.** Gateway telemetry is a record of requests. This adversary made none. The only record that could bound them is a write record on the upstream path, and that is frequently the record nobody kept. **Authorization upstream is usually weaker than at the endpoint.** The endpoint has a gateway, credentials per caller and a rate limit, because it faces the internet. The disposition queue, the ingestion job for a shared feed and the notebook that builds the training table often run under shared accounts, accept writes from a wide set of principals, and have nothing standing between a write and the next training run. ## What the API boundary actually bounds Get the direction of the claim right. An authenticated, rate-limited endpoint establishes which requests arrived and who sent them. It establishes nothing about how the weights came to encode what they encode. Saying "the model is behind an API, so the attack surface is the API" states a true fact about one surface and then generalises it to a surface it never touched. The useful reframe: everything the model knows came in through a writable path, and the threat model owes an account of each of those paths. ## Three things this is not **Not insider risk.** That names a category of person and leads to personnel controls. The contributor to a shared community feed your pipeline ingests is not an insider at all, and neither is whoever obtained the credential the ingestion job runs under. Access classes are stated over authorization, not over employment. **Not a compromise of the checkpoint artefact.** Tampering with a weight file after it is built, or with the registry it is fetched from, is a supply-chain problem with supply-chain answers around provenance and integrity of the file. Here nothing is tampered with. The artefact is exactly what the pipeline produced from exactly what it was given, and provenance controls will verify it perfectly. **Not the same as holding the weights.** An adversary handed weights and gradients has a much stronger vantage and a completely different set of moves available. The one described here may never see the model at all. ## How to state the vantage A usable statement has the same shape as any other access class, and fits in a few lines: which principals hold write on this path; what they cannot do (here: no queries, no weights); what consumes the write and on what schedule; and what stands between the write and the next training run, plus how long a write can sit before anyone would look. Written that way, the class is enumerable, testable and shrinkable, which is what separates a threat model from a worry. ## What an interviewer is checking That you scope before you attack. A candidate who answers every AI-security question at the endpoint has learned one surface. The tell of someone who has actually threat-modelled a trained model is that they ask where the training data comes from and who can write to it before they ask what the endpoint returns.

  • If they never query the model, how do they confirm the write worked?
    They do not need confirmation for the change to exist, and a threat model that assumes they require a feedback channel is granting them a limit they do not have. Where they do want confirmation, the deployed system's ordinary observable behaviour is usually enough. Treat query access as something that makes their work easier to tune, not as a precondition for the attack.
  • Is this not just supply-chain risk on the checkpoint file?
    No. A supply-chain attack tampers with the artefact after it is built: the file you fetch is not the file that was produced, and provenance and integrity controls are the answer. Here nothing is tampered with. The pipeline read exactly what it was supposed to read, and someone with a legitimate write chose the content it read. Provenance on the artefact will verify it perfectly.
  • Does an endpoint rate limit help against this at all?
    It prices the attacks that need queries, by making each probe cost time and money. It does nothing to an adversary whose entire interaction with you is one write to a corpus that a later training run reads. Different surface, different control, and quoting the rate limit as mitigation for this class is a scoping error, not a defence.

Guarding the front counter tells you who walked in and what they asked for. It tells you nothing about who had access to the recipe the kitchen cooks from.

saying these in an interview costs you the question

  • Says the API is the model's whole attack surface
  • Assumes an attacker must be able to query the model
  • Treats the training corpus as internal, therefore trusted
  • Confuses a write before training with a tampered artefact
  • Calls it insider risk and stops there

context