skip to content

Vantage and Reach

An attack claim means nothing until the access it assumed is stated: weights and gradients, a score vector, a bare label, or a write path upstream of training. Interviewers check you scope first.

on this pageshow

explore

questions

20

An attacker can write to your training corpus but cannot query the model — what does that buy them?

level: juniorimportance: must knowfreq 72%

answer

  1. two surfaces, not one
  2. the model was built from writes
  3. no request ever reaches the gateway
  4. authorization upstream is usually weaker

basics

~20 s

A durable change in the model itself. Everything the model knows arrived through writable paths, so an endpoint's authentication and rate limits bound only the inference-time surface, not what a later training run reads and turns into weights.

solid answer

~50 s

There are two different surfaces. The inference-time surface is what the endpoint exposes: requests, and whatever the reply carries. The train-time surface is everything a training run reads — the corpus, the label or disposition queue, an ingested third-party feed, the job that assembles the training table. An adversary holding a write on any of those needs no queries at all. Their change lands in the weights, so it applies to every caller, persists across restarts, and arrives through an ordinary authorized workflow: no request-level control could have refused it and no request log records it. That is why "it is behind an authenticated API" is not a threat model. The endpoint bounds who can *ask* the model something; it says nothing about who could shape what it answers. Ask instead who holds write upstream, and whether anyone would notice.

go deeper

for a junior

Be ready to name the two surfaces out loud: what the endpoint exposes, and what a training run reads. Knowing that a write upstream needs no queries at all is most of the answer at this level.

for a middle

Expect to walk the retraining loop and say which steps accept writes and who holds them, and why a change there is durable in a way that a single crafted request never is.

for a senior

Show that endpoint telemetry is the wrong evidence for this class, and say what evidence would bound it: the writer set on each upstream path, what reviews a change before a training run reads it, and how long a write sits unexamined.

for a principal

Own the framing when a design review calls the endpoint the trust boundary. Argue that every write path feeding a retrain is a production authorization surface, and be ready to say what you would fund to shrink it and what that costs in freshness.

## Two surfaces, and only one of them is the endpoint A trained model in production has two attack surfaces that have almost nothing in common, and confusing them is the most common scoping error in an AI threat model. The **inference-time surface** is the one everybody draws. A request arrives at an endpoint, the model returns something, and the controls sit on that path: authentication, authorization, rate limiting, input validation, request logging. An adversary working here must send requests. Their work is metered, priced and recorded, and what they can achieve is bounded by what the reply gives back and how many requests they can afford. The **train-time surface** is everything a training run reads. For a model that is periodically retrained, that includes the corpus of examples; the labels or dispositions attached to them, wherever those come from; any third-party or community data the pipeline ingests; and the job that assembles the final training table out of those pieces. Each of these is a write path, and each has its own set of principals who hold write on it. An adversary on the train-time surface sends no requests at all. They write one artefact that a later training run reads. Delete the adversary from that picture and you are left with an ordinary data pipeline; add them back and you have an access class that most threat models never wrote down. ## Why a write upstream is worth more than a query Four properties make this vantage distinctive, and they are the ones to say out loud. **It is durable.** A crafted request affects one prediction. A write that a training run consumes becomes a property of the weights: it applies to every caller, it survives restarts and redeploys, and it is still there after the incident that first drew attention to it has been closed. **No request-level control could have refused it.** The pipeline read exactly what it was designed to read, from exactly the place it was designed to read it. There is no malformed input to reject and no rate to limit. The workflow behaved correctly. **No request-level log records it.** Gateway telemetry is a record of requests. This adversary made none. The only record that could bound them is a write record on the upstream path, and that is frequently the record nobody kept. **Authorization upstream is usually weaker than at the endpoint.** The endpoint has a gateway, credentials per caller and a rate limit, because it faces the internet. The disposition queue, the ingestion job for a shared feed and the notebook that builds the training table often run under shared accounts, accept writes from a wide set of principals, and have nothing standing between a write and the next training run. ## What the API boundary actually bounds Get the direction of the claim right. An authenticated, rate-limited endpoint establishes which requests arrived and who sent them. It establishes nothing about how the weights came to encode what they encode. Saying "the model is behind an API, so the attack surface is the API" states a true fact about one surface and then generalises it to a surface it never touched. The useful reframe: everything the model knows came in through a writable path, and the threat model owes an account of each of those paths. ## Three things this is not **Not insider risk.** That names a category of person and leads to personnel controls. The contributor to a shared community feed your pipeline ingests is not an insider at all, and neither is whoever obtained the credential the ingestion job runs under. Access classes are stated over authorization, not over employment. **Not a compromise of the checkpoint artefact.** Tampering with a weight file after it is built, or with the registry it is fetched from, is a supply-chain problem with supply-chain answers around provenance and integrity of the file. Here nothing is tampered with. The artefact is exactly what the pipeline produced from exactly what it was given, and provenance controls will verify it perfectly. **Not the same as holding the weights.** An adversary handed weights and gradients has a much stronger vantage and a completely different set of moves available. The one described here may never see the model at all. ## How to state the vantage A usable statement has the same shape as any other access class, and fits in a few lines: which principals hold write on this path; what they cannot do (here: no queries, no weights); what consumes the write and on what schedule; and what stands between the write and the next training run, plus how long a write can sit before anyone would look. Written that way, the class is enumerable, testable and shrinkable, which is what separates a threat model from a worry. ## What an interviewer is checking That you scope before you attack. A candidate who answers every AI-security question at the endpoint has learned one surface. The tell of someone who has actually threat-modelled a trained model is that they ask where the training data comes from and who can write to it before they ask what the endpoint returns.

  • If they never query the model, how do they confirm the write worked?
    They do not need confirmation for the change to exist, and a threat model that assumes they require a feedback channel is granting them a limit they do not have. Where they do want confirmation, the deployed system's ordinary observable behaviour is usually enough. Treat query access as something that makes their work easier to tune, not as a precondition for the attack.
  • Is this not just supply-chain risk on the checkpoint file?
    No. A supply-chain attack tampers with the artefact after it is built: the file you fetch is not the file that was produced, and provenance and integrity controls are the answer. Here nothing is tampered with. The pipeline read exactly what it was supposed to read, and someone with a legitimate write chose the content it read. Provenance on the artefact will verify it perfectly.
  • Does an endpoint rate limit help against this at all?
    It prices the attacks that need queries, by making each probe cost time and money. It does nothing to an adversary whose entire interaction with you is one write to a corpus that a later training run reads. Different surface, different control, and quoting the rate limit as mitigation for this class is a scoping error, not a defence.

Guarding the front counter tells you who walked in and what they asked for. It tells you nothing about who had access to the recipe the kitchen cooks from.

saying these in an interview costs you the question

  • Says the API is the model's whole attack surface
  • Assumes an attacker must be able to query the model
  • Treats the training corpus as internal, therefore trusted
  • Confuses a write before training with a tampered artefact
  • Calls it insider risk and stops there

context

open as a page

In a merchant-onboarding risk model, what separates fields an applicant rewrites for free from ones they cannot?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Price to the applicant, not distance. Self-declared fields cost only a retype. Fields such as trading tenure cost real money or real waiting. Fields bound to a third-party attestation cost the applicant a fraud against that party.

open as a page

Why are the fields a prediction API returns part of its threat model?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Every field handed back is signal an attacker gets for free. A bare verdict, a top-k list and a full score vector are three different access classes, and the reply class decides which attack families are cheap enough to run.

open as a page

Your banking app ships the face-matching model inside its install bundle — what access does that hand an attacker?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Anyone who installs the app has the weights. Full white-box access becomes the normal operating point rather than a worst case: the attacker studies and attacks the model locally, unmetered, sending nothing to your server.

open as a page

In an adversarial ML evaluation, what does it mean to grant an attacker white-box access?

level: juniorimportance: must knowfreq 78%

basics

~20 s

White-box access is an assumption that the attacker holds the model's weights, architecture and gradients. You grant it on purpose during evaluation so the result measures the model's robustness rather than how well you kept the weights secret.

open as a page

Your moderation API returns only a verdict, no scores. Which attacks does that stop?

level: middleimportance: must knowfreq 62%

basics

~20 s

None outright. Removing confidences deletes a cheap signal and forces the adversary into label-only families such as decision-based search and verdict-based membership tests, at far higher query cost. Coarsening a reply moves a family's price; it closes none.

open as a page

What makes a write path into a training corpus an access class rather than "insider risk"?

level: middleimportance: should knowfreq 54%

basics

~20 s

An access class states a vantage and a limit: who holds write, what later reads it, what stands in between, and how long a write goes unexamined. It is a property of authorization, not of a person's motive.

open as a page

Why does a perturbation radius fail to describe a loan applicant editing their own application?

level: middleimportance: should knowfreq 50%

basics

~20 s

An applicant does not nudge a stated income by a fraction of a percent; they type a different number. A radius assumes a true input to stay near, and a self-authored form has none. The limit is per-field cost.

open as a page

The face encoder ships on-device but the match threshold stays server-side — what does that split still bound?

level: middleimportance: should knowfreq 55%

basics

~20 s

It bounds only what the server computes for itself: the enrolled template it holds and the accept decision it makes. Everything the shipped encoder computes is the attacker's, unmetered, and offline search leaves no trace at your endpoint.

open as a page

Why is a white-box attack result treated as a ceiling and a query-only result as a floor?

level: middleimportance: should knowfreq 62%

basics

~20 s

A white-box adversary holds everything a weaker one could obtain, so its measured success is the most any adversary achieves — a ceiling. A query-only run measures one particular under-informed adversary, so its success can only be raised by more budget or better technique — a floor.

open as a page

Your network-flow alert-triage model stopped flagging one traffic class and request logs look normal — what does that rule out?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Almost nothing. Request logs establish which queries arrived, not what the weights encode. A change written into the disposition queue or the training table produces no request at all, so the evidence that bounds it is the upstream write record.

open as a page

A merchant-risk model's most predictive columns are all self-declared - what do you report?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Accuracy was measured on applicants with no reason to misstate those fields. Report the share of score sitting on columns the subject retypes for free, and what the cheapest application that flips a decline costs.

open as a page

Your moderation API now rounds scores to two decimals and returns only the fired policy. What changed?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Two separate changes. Rounding coarsens the signal and creates ties, so probes must be larger or more numerous. Dropping the non-fired categories is the bigger cut: the reply now describes one boundary instead of several.

open as a page

The model shipped inside a past mobile app release was extracted — what does shipping a replacement buy you?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Very little on its own. Copies already taken cannot be recalled, and old installs keep the old file working for months. A replacement helps only once the server stops honouring what the old model produces.

open as a page

Your white-box robustness evaluation was capped at contracted GPU-hours — what does that cost the bound?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The ceiling is only as tight as the search you paid for. A granted-access run that ran out of compute reports that those attacks failed, not that no adversary succeeds, so the bound holds only against adversaries whose own effort is smaller than the search you funded.

open as a page

An outsourced analyst's triage clicks become your model's labels — do you accept that write path?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Usually yes, but only once it is written down as a production authorization surface. The decision is not whether the vendor is trustworthy; it is what attribution, gating and dwell you will fund, and what freshness that costs.

open as a page

Onboarding wants to drop bank verification to lift signup conversion - what do you own in that call?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Naming what the step buys: it holds one column the applicant cannot set for free. Removing it moves score mass onto retypeable fields and drops the cheapest successful application to near zero. Price that shift; the conversion appetite is the business's.

open as a page

A product owner wants to drop confidence scores from a paid moderation API as a security control. What do you say?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Support it as a repricing with a number attached, never as a boundary. Say which families move to which cost, insist the multiplier is measured, and be explicit that paying integrators lose a field some will rebuild.

open as a page

Product wants the whole liveness check on-device for latency and offline use — what do you require stays server-side?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Whatever authorizes a consequential action. Ship the model if latency demands it, but the decision it feeds must be made server-side on evidence the client cannot mint, with the device result treated as a hint rather than a credential.

open as a page

A vendor asks your evaluation lab to test its model without receiving the weights — how do you decide?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Refusing the grant makes the report partly a measurement of the vendor's secrecy, which is not a safety property and cannot support a clearance decision. Take the query-only run only as a supplementary realism datapoint, and say in writing what the absent grant means.

open as a page