What makes a write path into a training corpus an access class rather than "insider risk"?
answer
- authorization, not personnel
- state a vantage and a limit
- what reads the write, and when
- a shared feed's contributor is no insider
basics
~20 sAn access class states a vantage and a limit: who holds write, what later reads it, what stands in between, and how long a write goes unexamined. It is a property of authorization, not of a person's motive.
solid answer
~50 sThreat models are written over access, not over people. "Insider risk" names a category of person and leads to personnel controls; it cannot tell you what somebody with that vantage can achieve. An access class is stated the same way white-box access or query-only access is stated: what the adversary can touch — write to the disposition queue, to an ingested community feed, to the job that assembles the training table — what they cannot touch, which here is usually no queries and no weights, what consumes the write and on what schedule, and what stands between the write and the next training run. Stated that way it is enumerable, testable and shrinkable. It also covers the outsider contributing to a shared feed, who is not an insider at all. Reach is set by whether a later training run reads the write, not by the writer's job title.
go deeper
Know that a threat model names access, not people, and that write access upstream of training is one such access. Being able to say why "a trusted insider" is not a scope is enough here.
Be ready to state one upstream path in full: writer set, what they cannot do, what consumes the write and when, and what stands in between. That four-part statement is what the question is testing.
Demonstrate that you have actually enumerated writer sets on a real pipeline, including shared accounts and service identities, and can say which paths a training table genuinely draws from as opposed to which ones the diagram claims.
Be prepared to argue for owning these paths as production authorization surfaces with named owners and a review expectation, rather than leaving them to a personnel process that cannot describe their reach.
## Why the framing matters Threat models are statements about access. Every other vantage in adversarial ML is written that way: an adversary with weights and gradients, an adversary with a full probability vector, an adversary with a bare top-1 label, an adversary with a query budget they must pay for. Each one names what can be touched and what cannot, and every attack claim is only meaningful relative to one of them. Write access upstream of training deserves exactly the same treatment, and it usually does not get it. It gets filed under "insider risk", which is a different kind of statement entirely: a category of person, handled by a personnel process — background checks, joiners-and-leavers, awareness training. Those are worth having and they answer a different question. They cannot tell you what somebody holding a write on the label queue could achieve, or whether anybody would notice. ## What an access class has to say Four parts, and none of them mention a person's motive. **Who holds write.** Not "the team", but the enumerable set of principals with write on this specific path: the named accounts, the shared account an outsourced queue authenticates as, the service identity an ingestion job runs under, anyone who can open a change against the notebook that assembles the training table. If the answer is a shared credential, then the writer set is "whoever holds that credential", and that is the honest statement. **What they cannot do.** The limit is the other half, and leaving it out turns the class into a vague fear. Here the limit is usually stark and worth stating plainly: no query access, no weights, no ability to see the model's replies at all. That limit is what makes it a distinct class rather than a weaker version of something else. **What reads the write, and when.** A write only has reach if a training run consumes it. Which run, on what schedule, and does the write reach it directly, or through an assembly step that filters or aggregates first? A path whose writes are never read by any training run is not an access class into the model at all, and saying so is a real result. **What stands in between, and for how long.** Does anything compare one training table to the previous one before a retrain consumes it? Does a change to the assembly job take a second pair of eyes? How long can a write sit before anybody would look at it? This is the part that decides whether the class is merely present or effectively unbounded. ## Why "insider" is the wrong word specifically The most likely writers on these paths are frequently not insiders in any employment sense. A community or partner feed your pipeline ingests accepts contributions from a population you do not employ and cannot enumerate. An outsourced triage or annotation queue is clicked through by people who work for someone else, often authenticated as one shared identity. An ingestion job runs under a service account that is nobody, so a credential for it is access with no person attached at all. And a genuine employee's stolen credential produces exactly the same write as the employee would. Every one of those is the same access class with the same reach. Grouping them by employment status splits a single technical vantage across four unrelated processes and loses the thing they have in common: they can all write one artefact that a later training run reads. ## What it buys you to state it properly Once written as a vantage and a limit, the class behaves like any other line in a threat model. You can enumerate it, so you know how many principals hold it. You can test it, by asking whether a change you make to an upstream path actually reaches a training table and whether anyone sees it on the way. You can shrink it, by narrowing the writer set, adding attribution, or putting something between the write and the retrain. And you can be honest about it afterwards, because the claim it supports is a scope — these principals could have written, over this window, with this much review — rather than an assurance that nothing happened. ## The interview signal An interviewer asking this is checking whether you can carry the discipline of stated access assumptions past the endpoint. Candidates who have only ever reasoned about queries answer with personnel controls, because that is where "someone with access to our data" lives in their mental model. The stronger answer treats the write path exactly as it would treat white-box access: name the vantage, name the limit, and then ask what that adversary can achieve.
- Name the things you would write down for one such path.The principals who hold write on it; what they cannot do besides, which is usually no queries and no weights; what consumes the write and on what schedule; and what stands between the write and the next training run, plus how long a write can sit before anyone looks. That is a complete vantage-and-limit statement and it fits in four lines.
- Why does "the whole team is trusted" not close this?Trust is not a scope. The question a threat model answers is what somebody with that access could achieve and whether anyone would know, and the answer is identical whether the write came from a trusted colleague, from their stolen credential, or from an automated ingestion job running under a shared service identity. Enumerating principals is the work; grading character is a different process.
- Is a write path that no training run reads still an access class?Not into the model. Reach comes from consumption: if nothing a training run reads is derived from that path, a write there cannot become a property of the weights, and saying so is a legitimate result rather than a dodge. The check worth doing is whether an assembly step quietly pulls it in anyway, which is common and is exactly the thing people get wrong.
saying these in an interview costs you the question
- Answers with personnel controls instead of an authorization scope
- Assumes every upstream writer is an employee
- Names the attack but never says who can write
- Treats a service account's write as nobody's access
- Omits the limit, so the class becomes a vague worry