skip to content

What is an LLM-drafted threat model good for when its threats read plausible but generic?

level: middleimportance: should knowfreq 54%

answer

  1. coverage, not analysis
  2. beats a blank page
  3. same list for two unlike systems
  4. shapes are learnable, worth is not
  5. seed the session, then argue

basics

~20 s

It is a first-pass checklist, not an analysis. The draft reliably names the obvious threat for each component type, so a review starts from a filled page instead of a blank one. Ranking and business impact stay human work.

solid answer

~50 s

Treat the draft as coverage, not analysis. Given a pasted design, a language model returns what it has most often seen written about components of that shape: session handling on a web tier, injection at a database edge, secrets sitting in a config store. That is genuinely useful. It beats an empty page, it gives a junior the vocabulary, and it names the categories a tired reviewer skips at 5pm. What it does not do is know what your system is worth. The same fourteen-item list comes back for a print-shop order manager and for a trading platform, because nothing in it is derived from what an outage or a leak actually costs either business. So use it to seed a per-element pass, then make the engineers who built the thing argue with every line, add the threats that come from this design's own oddities, and do the rating yourselves.

go deeper

for a junior

Be ready to say what the draft is: a starting checklist of common threats per component. Show that you read every line with the design open beside it rather than accepting the list wholesale.

for a middle

Explain the mechanism — the model completes patterns over component shapes, and shapes are learnable while your business impact is not. Describe how you use the draft to seed a per-element pass instead of replacing it.

for a senior

Demonstrate the working habit: run the draft, then run the session anyway, forcing a real, handled, or not-applicable verdict on every line and continuing past the end of the list to the threats the design's own oddities create.

for a principal

Own the question of what a drafted floor does to a team's habits. Argue whether cheap drafts raise the baseline across many teams or quietly replace the argument that finds the threats that matter, and say how you would measure which is happening.

## What a drafted model actually is When you paste a design document into a general-purpose language model and ask for a threat model, what comes back is a **plausibility artifact**: a list of the threats most commonly written about systems that look like the one you described. It is not the output of reasoning over your architecture; it is the output of pattern completion over descriptions of architectures. That distinction explains every strength and every weakness discussed below. ## Why the output is generic The draft is conditioned on component *shapes* — "a public API", "a relational database", "a message queue", "an object store" — because those are the tokens the design document supplies in bulk. Threats attach to shapes very reliably: an internet-facing endpoint attracts spoofed identity and denial of service, a database edge attracts injection and tampering, a config store attracts credential disclosure. Those associations are dense in the training material, so the model produces them with high confidence and good coverage. What is *not* dense in that material is your specific business. Nothing in a design document says that a two-hour outage of a print-shop order manager annoys forty customers, while a two-hour outage of a matching engine loses money by the second and draws a regulator. Because impact is absent from the input and unlearnable from the shape, the draft either omits it or fills it with hedged filler ("could result in reputational damage"). That is why two systems with no business resemblance receive near-identical lists — a good diagnostic to run deliberately: draft a model for a second, unrelated design and compare. The overlap is the generic floor; the delta is the only part that responded to your input. ## Where the generic floor still earns its keep Dismissing the draft because it is generic is the opposite mistake. Real value: - **It removes the blank page.** Most threat-modeling sessions die from not starting. A page of candidate threats gets people arguing immediately, and arguing is the activity that finds threats. - **It supplies vocabulary.** A junior engineer who has never run a session gets the standard categories phrased in the language of their own components, which is a much better teacher than an acronym table. - **It covers the boring.** The categories a human reviewer skips because they are always the same — repudiation, weak logging, denial of service on an unauthenticated endpoint — are exactly the ones a model never gets bored of listing. - **It normalizes phrasing.** Threats written in a consistent "actor / action / asset / effect" form are easier to deduplicate and track than free-form notes from six engineers. A reasonable frame: the draft gives you the *floor* of coverage that any competent reviewer would eventually reach, delivered in seconds. The ceiling — the threat that exists only because of a decision made in this design — is still yours. ## What humans must add 1. **Asset and impact.** Attach to each threat what is actually lost: money, customer data, availability of a service someone depends on, credentials, audit truth, unreleased intellectual property. This is where a threat becomes a risk, and it is the step the draft cannot take. 2. **Design-specific threats.** The interesting ones come from the design's oddities: the flow that reuses one service account across two trust levels, the batch job that bypasses the API and writes to the database directly, the reconciliation step that trusts a partner's file. Read the draft, then read the design again looking for what the draft *could not have known*. 3. **Adversary realism.** The draft's implied attacker is usually an anonymous internet user. Force the question of the authenticated low-privilege tenant, the insider with legitimate access, the compromised operator, the neighbour on the same network segment. 4. **Rating and sequencing.** Comparability of ratings across a backlog is a human discipline; a model asked to score items one at a time produces numbers that do not survive being put side by side. ## How to use it in practice Generate the draft, then run the session anyway. Project the draft on the wall, walk the data-flow diagram element by element, and for every drafted line demand one of three verdicts from the people who built the system: *real and unaddressed*, *real and already handled — here is where*, or *not applicable in this design because…*. Delete the third category rather than leaving it to pad the count. Then keep going past the end of the list, because the list ended for reasons that have nothing to do with your system. One last discipline: length is not coverage. A longer draft usually means more restatement of the same category across more components, not deeper analysis. Judge a draft by how much of it is design-specific, not by how many rows it has.

  • How do you tell a design-specific threat in the draft from a plausible-generic one?
    A design-specific line names something only this design has: a particular flow, a named boundary crossing, an assumption the document states. A generic line would survive being pasted into any other system's model unchanged. A cheap test is to draft a model for an unrelated design and diff the two: the intersection is the generic floor, and the difference is the only part that actually responded to your input.
  • Does a longer draft mean better coverage?
    No. Length usually comes from restating the same category against more components, not from finding more. Judge a draft by the fraction of lines that reference something specific to this design, and by whether anything in it surprised the people who built the system. A twelve-line draft with three surprises is worth more than an eighty-line one with none.
  • The draft lists a threat that is impossible in this design. How do you handle it?
    Delete it in the session and record why in one line, because the reason is usually an assumption worth capturing — "not applicable, that service never accepts external traffic" is a claim that may stop being true next quarter. What you must not do is leave it in the list as low severity: padding the backlog with impossible items teaches everyone to skim it.

It is a stock inspection checklist, not an inspection. The checklist tells you which cupboards every house has; it does not know which cupboard this house keeps the valuables in.

saying these in an interview costs you the question

  • Treats the generated draft as the finished threat model
  • Assumes a longer list means better coverage
  • Cannot point to which lines are design-specific
  • Expects the model to know what a compromise costs the business
  • Reads the draft's confident tone as evidence it is grounded

context