How do you validate an LLM-drafted threat model that may omit components or assert controls you lack?
answer
- the draft cannot flag its own gaps
- validate scope before threats
- unseen module, unmentioned data store
- a control needs an owner and evidence
- unverified reverts to assumption
basics
~20 sCheck the inventory before the threats: reconcile every component and data store in the draft against the real system, then make an owner point at evidence for each control the draft claims exists. Anything unverified reverts to a stated assumption.
solid answer
~50 sBoth failure modes are invisible from the draft alone, so the review has to attack them from outside it. First, validate scope: build the component and data-store list independently — every infrastructure module, every deployed resource, every egress — and diff it against what the draft discusses. A draft generated from one module of a research cluster's infrastructure will happily produce a confident model that never mentions the object store defined in a second module it was never shown, with no hint that anything is missing. Second, validate claims: a drafted line like "requests are authenticated at the ingress" is an inference from the design document's aspiration, not an observation. In the walk-through, the engineer who owns that component points at the actual configuration or test, or the line is rewritten as an assumption and every threat it silently mitigated goes back on the list.
go deeper
Be ready to say that a generated model can leave things out without warning, and that a control it mentions is not proof the control exists. Show that you check the draft against the real system.
Explain the mechanics of both failures: the model has no representation of what it was not shown, and it repeats the design document's intentions as implemented state. Describe the inventory diff and the evidence rule.
Demonstrate the review you would actually run: an independently derived component and flow inventory diffed against the draft, then a walk-through with the builders where every claimed control gets an owner and a pointer or is downgraded to an assumption.
Own the standard: what a threat model must contain before it counts as reviewed, how verified controls and assumptions stay visibly separate over time, and how you keep drafting cheap without letting unverified claims accumulate across many teams' models.
## Two failures with one shape A drafted threat model fails in two ways that a reader cannot detect by reading it: - **Silent omission** — the model says nothing about a component it was never shown, and nothing in the output marks the gap. - **Asserted-but-absent control** — the model states that a protection exists, because the design document said it should, and then quietly treats threats as handled on that basis. Both are unfalsifiable from inside the artifact. A missing section looks exactly like a section that was correctly judged unnecessary; a confident sentence about authentication looks exactly like a verified fact. So validation cannot be "read the draft carefully". It has to bring in information the draft did not have. ## Failure one: what it never saw Suppose a team drafts a model from the infrastructure code for a genomics research cluster and pastes in one module: compute nodes, the scheduler, the internal network. The draft is competent about those. It says nothing about the object store holding subject data, because that bucket is defined in a second module nobody pasted. The output contains no "I did not see storage" caveat, because the model has no representation of what it was not given. The most sensitive asset in the system is therefore absent from a document that looks complete. The counter is an **inventory reconciliation that runs before any threat is discussed**: 1. Derive the component list from a source the assistant did not touch — the deployed resources, the full set of infrastructure modules and their state, the network egress list, the identity system's list of principals, the data catalogue. 2. Diff that list against the entities the draft actually names. 3. For every entity in the diff, decide explicitly: in scope and needing threats, or out of scope with a written reason. The same reconciliation should be done for **flows**, not just components. A component list that matches can still hide the backup job, the analytics export or the support tool that reads the same store by another path. Walking the data-flow diagram and asking "what else reads or writes this" catches those; reading the draft never will. A useful discipline is to require the draft to state its own inputs at the top — which files or documents were provided, dated. That does not make the model self-aware about gaps, but it turns the omission into an auditable fact: *this model was generated from module A only*. A reviewer six months later can then see the boundary of what was ever considered. ## Failure two: controls that exist only in prose The second failure is more insidious because it *reduces* the threat list. A design document for a municipal parking-permit service says the ingress terminates authentication. The draft repeats it as fact and, on that basis, rates spoofing threats against the permit-issuance endpoint as mitigated. If in reality the ingress passes anything through and the service trusts a header, the model has documented the system as safe in exactly the place where an anonymous internet user can issue themselves permits. The rule that fixes this is simple to state and needs enforcement: **a control is a claim with an owner and a piece of evidence.** In the walk-through, for each control the draft names, the engineer responsible either: - points at the concrete thing — the ingress configuration, the policy that rejects unauthenticated requests, the test that proves the endpoint returns 401 without credentials; or - says they believe it is true but cannot show it, in which case the line is rewritten as an **assumption**, flagged, and given an action to verify; or - says it is not implemented, in which case every threat the draft treated as handled goes back to open. Keep verified controls and assumptions in visibly different columns in the finished model. The most common way a threat model goes stale and dangerous is that an assumption from one review is read as a fact in the next. ## The walk-through itself The validation event is a session with the people who built the system, not a solo desk review by whoever ran the prompt. Practical shape: project the diagram, not the list. Go element by element and flow by flow. For each, ask what the draft said, and then ask the two questions the draft cannot answer — *is anything here missing from what it saw*, and *is every protection it credits us with actually running*. Record decisions as you go, including the deletions and their reasons. ## Signals a draft needs more scepticism - It never says "unknown" or "not provided" anywhere, despite an input that plainly lacked detail. - Its component names do not match your real resource names. - Its mitigated-versus-open ratio is high on a system nobody has hardened yet. - Controls appear in the same confident register as threats, with no distinction between observed and assumed. None of these prove an error; each one tells you where to spend the review's attention.
- What would you change in the process so the missing object store is caught next time?Two changes. Feed complete inputs — all infrastructure modules, not one — and require the draft to list, at the top, exactly which inputs it received. Then keep the independent check anyway: derive the component and data-store inventory from deployed resources or the data catalogue and diff it against the draft. Better inputs reduce the gap; only the external diff detects one.
- How do you record the difference between a verified control and one the draft assumed?Separate them structurally, not by tone. Verified controls carry an owner and a pointer to the evidence — a configuration, a policy, a test. Assumptions live in their own list with a named owner and a date to confirm by, and every threat resting on an assumption stays open until it is confirmed. Mixing them is how a later reader inherits a claim as a fact.
- The draft rates most threats as already mitigated. What does that tell you?On a system nobody has hardened, it usually means the model is echoing the design document's intentions back as implemented state. Treat a high mitigated ratio as a review signal, not as good news: sample the mitigated items first and demand evidence for each, because that is where a false sense of safety concentrates.
A surveyor who only walked the rooms you unlocked will still hand you a confident floor plan of the whole house — with the flooded basement simply not on it.
saying these in an interview costs you the question
- Assumes the draft would say so if it lacked information
- Accepts a listed control as an implemented control
- Reviews only the threats the draft printed
- Marks threats mitigated on the design document's wording
- Does the validation alone instead of with the builders