skip to content

Keeping the Model Alive

A model nobody updates becomes a confident lie about the system: drift goes unnoticed and findings never close. Interviewers ask how you prove a model still matches what is running.

on this pageshow

explore

questions

13

How do you check that a threat model still matches the system that is actually deployed?

level: middleimportance: must knowfreq 55%

answer

  1. compare against reality, not memory
  2. inventory, routes, grants, telemetry
  3. two passes, not one
  4. the reverse pass finds the most
  5. date it and attach the evidence

basics

~20 s

Reconcile the model against independent evidence of the running system - deployment descriptors, service and datastore inventories, routing and access config, traffic telemetry - and record every element, flow or trust boundary that reality has and the model does not.

solid answer

~50 s

Drift is found by comparing the model with evidence from the running system, not from anyone's memory. I take the model's inventory - processes, stores, external entities, flows, trust boundaries - and check each entry against deployment descriptors, the service and datastore inventory, routing config, identity grants, and telemetry showing who actually calls whom. Then I run the reverse pass, which finds most drift: what exists in production that appears nowhere on the diagram. Walking one real request end to end catches the rest. A ride-hailing dispatch model still showing a single primary database says nothing about the analytics read replica the data team stood up in a second region - a new store of rider location history, a new operator population that can read it, and a boundary nobody drew. I write it up as a dated diff with the evidence, so drift lands as a finding rather than an impression.

go deeper

for a junior

Be able to say what drift is and why a wrong model is worse than none, and to name two places you would look for the truth about a running system, such as the deployment configuration and the datastore inventory.

for a middle

Explain the mechanics: the forward pass over drawn elements, the reverse pass from production inventory, and tracing one real request. Say which evidence sources you would read and why observed traffic beats an owner's recollection.

for a senior

Show judgment about which differences matter. Rank a moved trust boundary and a new copy of sensitive data above renames, and turn a difference into a stated consequence - who can now reach what - rather than a diagram correction.

for a principal

Own how reconciliation is resourced and trusted across an estate: which signals run continuously, which services get a deep manual check, and how a model's verification date and evidence are published so others can judge how much to trust it.

## What drift is A threat model is a claim about a system: these are its components, these flows cross these trust boundaries, these are the threats we accepted and the controls we rely on. The system keeps changing after the model is drawn. **Model drift** is the gap that opens between the drawn system and the running one. The danger is not that the model becomes useless - it is that it stays *persuasive*. A confident diagram that no longer matches production is worse than no diagram, because reviewers, responders and auditors all reason from it. Detecting drift is therefore a reconciliation exercise: compare the model against evidence produced by the running system rather than by the people who drew it. ## Sources of evidence Use sources that are generated by, or read out of, the live environment: - **Deployment and infrastructure descriptions** - what services, queues, buckets, databases and functions are declared for this environment, and in which regions and accounts. - **Service and datastore inventories** - the registry of what is running and what data each store holds, plus its classification. - **Routing, network and firewall configuration** - which paths exist between components, which are exposed externally, and where segmentation actually sits. - **Identity and access grants** - which principals, human and machine, can reach each store or API. Grants often reveal a caller the diagram never drew. - **Telemetry** - call graphs, egress destinations, queue topics with live publishers and consumers. Observed traffic is the strongest single answer to "does this flow exist?". - **Data catalogue or privacy register** - new data classes landing in an existing store are drift even when no new box appeared. A claim from the owner ("nothing architectural changed") is a hypothesis to test, not evidence. ## The two passes **Forward pass - does everything in the model still exist as drawn?** Walk the model's element list. Each process, store, external entity, flow and trust boundary is checked against the evidence above. Phantom elements matter: a component removed from production leaves behind threats marked "mitigated by" a control that nothing now implements, and it inflates the sense of how much of the estate is covered. **Reverse pass - what exists in production that the model never mentions?** This is where most drift is found, because additions rarely announce themselves to the person who owns the model. Take the inventory of what actually runs and subtract what the diagram accounts for. New consumers on an existing queue, a second region, a vendor integration, an added replica, a debugging endpoint, a data export job. **Trace one real request end to end.** Pick a representative transaction and follow it through the actual hops, including retries, caches and async fan-out. This catches structural drift that neither list-based pass surfaces - for example an intermediary that terminates TLS, or a queue that has quietly become the front door for a second caller. ## What counts as a material difference Not every difference is worth a finding. Rank by security consequence: 1. **A trust boundary moved, appeared or vanished.** New region, new account, a component now reachable by a wider population, a call that used to be internal now crossing a vendor edge. 2. **A new store or flow carrying data of an existing or new classification.** New copies of sensitive data are new targets even when the logic is unchanged. 3. **A change to who can reach something** - a new principal, a broadened role, a queue anyone in the cluster may publish to. 4. **Cosmetic differences** - renames, layout, a helper split into two deployments with no boundary change. Note them; do not treat them as risk. ## A worked example A ride-hailing dispatch service was modeled with one primary database inside the service's own boundary. Six months later the data team stood up an analytics **read replica in a second region** so their queries would stop competing with dispatch. Nothing in the service repo changed, so the model did not change. Reconciliation against the datastore inventory and the region list surfaces it immediately. What drifted is not a box. Rider location history now lives in a second jurisdiction, inside a boundary owned by a different team, readable by that team's operators and by whatever tooling they attach. A compromised operator of that replica now reaches the same asset the original model spent its effort protecting, along a path the model never assessed - no rate limits, no access review, likely different logging. The finding is not "diagram out of date"; it is "an unassessed copy of rider location history exists in region B, reachable by an operator population we never analysed". ## Recording the result Write drift down the way you would any other finding: the date the check ran, what evidence was used, the concrete diffs, and their consequence. Attach the model's own metadata - which version was checked and which release of the system it was checked against. Without that pairing, the next review cannot tell what changed since the last one, and "models v1 through v7" becomes a pile of files no one can diff. A check with no recorded evidence is indistinguishable from no check at all.

  • Your models are numbered v1 through v7 with no record of which release each was drawn against. Why does that hurt the next re-review?
    Because reconciliation is a diff, and a diff needs two anchored endpoints. If v5 is not pinned to a release tag or a date, the reviewer cannot say which system changes happened after the last verified state, so they must redo the whole model from scratch every time instead of examining a delta. Pin each model version to the release or commit it describes and to the date and evidence of its last verification.
  • The reconciliation finds a component in the model that no longer exists in production. Is that worth acting on?
    Yes. Phantom elements carry threats marked as handled by controls that nothing now implements, and they make coverage look better than it is. Remove the element, re-home any threat that migrated with the functionality into the component that absorbed it, and note where the data went - decommissioned code frequently leaves its data behind.
  • What would you refuse to accept as proof that a model is still accurate?
    An owner asserting nothing changed; a diagram that renders cleanly; a review date with no scope or evidence recorded; a green build. None of those observe the running system. I want artefacts read out of the live environment - inventory, configuration, grants, telemetry - and a note of which of them was checked, when, and against which release.

It is a stock count, not a re-read of the ledger: you walk the warehouse and see what is on the shelves, then compare that with what the book says should be there.

saying these in an interview costs you the question

  • Says the model is current because nobody remembers changing anything
  • Treats a tidy, well-rendered diagram as evidence of accuracy
  • Only checks that drawn elements still exist, never the reverse
  • Ignores infrastructure stood up outside the service's own repo
  • Reports drift with no date, evidence or version it was checked against
  • Counts renames and layout changes as risk-relevant drift

context

open as a page

When one threat recurs across thirty-eight threat models, what changes about how you fix it?

level: middleimportance: must knowfreq 60%

basics

~20 s

A threat repeating across nearly every model is a missing platform capability, not thirty-eight product defects. Build the control once where every service inherits it, then each model names that control instead of carrying its own ticket.

open as a page

Which metrics show whether a threat modeling practice is actually working?

level: middleimportance: must knowfreq 57%

basics

~10 s

Four families: coverage of in-scope designs and services against a stated denominator, findings closed and mean time to close, time-to-model from trigger to owned findings, and escaped issues found after the design passed.

open as a page

Your threat model produced 60 threats: how do you turn them into backlog items teams actually close?

level: middleimportance: must knowfreq 62%

basics

~20 s

One ticket per threat, in the team's normal delivery backlog, each carrying the threat statement, the agreed control, a named individual owner and a due date. A single 'security' epic hides ownership and lets the whole batch stall.

open as a page

Why require the threat-model diagram update in the same pull request as the code change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It turns drift into a review-time question. The reviewer sees the model diff next to the code diff and can refuse a change whose new flow, store or boundary is missing, instead of discovering the gap at the next audit.

open as a page

Twenty threat models all assume CI runners cannot reach production databases, and nobody has tested it. What do you do?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Turn the repeated assumption into a test the platform owns and runs continuously. An unverified inherited control is a shared belief, not a control, and twenty models depend on it, so one wrong belief invalidates all twenty at once.

open as a page

Your dashboard reports 94% of services threat-modeled: how do you check that number is honest?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Interrogate the three hidden choices behind the percentage: which population forms the denominator, what test earns a service a tick in the numerator, and how recent the model must be to count. Then sample and reconcile against an independent inventory.

open as a page

How do you run an escaped-issue review when an incident hits a system you threat-modeled?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Walk a fixed chain and stop at the first no: was it in scope, on the diagram, enumerated, rated correctly, accepted deliberately, controlled, and was the model current. Each stop names a different part of the practice to fix.

open as a page

What is your definition of done for a threat-model finding before its ticket can be closed?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Done means the agreed control is merged and running in the environment the threat targets, and evidence exists that fails if the control is removed. A design note, a ticket comment or a reviewer's approval is not done.

open as a page

How do you retire a threat model for a decommissioned service so it stops being cited as evidence?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Give every model a status and a verification date, mark this one retired with the date, reason and successor, and generate assurance packs from the live index rather than circulated copies, so a retired model can never read as current coverage.

open as a page

A team's threat model for a new service covers only its two bespoke flows and inherits the platform by reference — when is that delta model legitimate?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Legitimate when the inherited part is named control by control, each of those controls has evidence behind it, and the team can show why its two flows fall outside the inherited pattern. Lazy when 'the platform handles it' names nothing.

open as a page

Why does reporting raw threat count as a threat modeling metric backfire?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Threat granularity is the analyst's choice, so the count has no fixed unit and is not comparable across teams. Once it becomes a target, the cheapest way to raise it is splitting one threat into several.

open as a page

How do you record a threat's closure so the ticket re-opens when the design that justified it changes?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Record the invariant the closure depends on, not just the fix. Name what must stay true, attach a check that fails when it stops being true, and re-open the original ticket rather than filing a new one so the closure history travels with it.

open as a page