skip to content

How do you run an escaped-issue review when an incident hits a system you threat-modeled?

level: seniorimportance: should knowfreq 51%

answer

  1. ask which step would have had to differ
  2. a fixed chain, stop at the first no
  3. in scope, drawn, enumerated, rated, accepted
  4. each stop has its own remedy
  5. the output is a change to the method

basics

~20 s

Walk a fixed chain and stop at the first no: was it in scope, on the diagram, enumerated, rated correctly, accepted deliberately, controlled, and was the model current. Each stop names a different part of the practice to fix.

solid answer

~50 s

I ask which step of the method would have had to work differently, not who missed it. The chain is fixed: was the component in scope for modeling, was it on the diagram, were threats enumerated against it, was it rated correctly, was it deliberately accepted, was a control chosen and shipped, and was the model newer than the system. Stop at the first no, because each branch has a different remedy. In one case an outage came from an unauthenticated internal admin endpoint drawn on the diagram but never enumerated - the element was skipped because *internal* was treated as trust. That is an enumeration-discipline failure, fixed by a rule that every administrative interface is enumerated as an entry point regardless of network position. The output is one or two checkable changes to the practice, kept blameless and recorded, so the branch that recurs shows where the method is weak.

go deeper

for a junior

Know what an escaped issue is - a design-level problem found in production on a system that had a model - and that the review asks which step of the method missed it rather than who is at fault.

for a middle

Be able to walk the chain in order and say what each stop implies: scope and triggers, diagram completeness, enumeration discipline, rating calibration, acceptance authority, follow-through, freshness. Explain why you stop at the first no.

for a senior

Show that you produce specific, checkable changes rather than exhortations, and that you can diagnose the classic shape - an element drawn but skipped because its network position was read as trust. Be ready to run this alongside an operational post-incident review.

for a principal

Own the feedback loop across many incidents: which branch recurs, where to invest scarce facilitation effort, and how to keep escapes reportable at all. Defend a metric that makes your own practice look fallible against pressure to stop counting it.

An **escaped issue** is a design-level security problem that reaches production on a system the practice had already modeled - found by an incident, a penetration test, a bug report or a later review. The escaped-issue review is the practice's own retrospective: not *who broke it*, but *which step of our method would have had to work differently for this to have been caught*. Take the concrete case. An outage was traced to an internal administrative endpoint that required no authentication; the endpoint could be reached from a neighbouring service that an attacker had a foothold on, and the actions it exposed both took the service down and left the audit trail unable to attribute what happened. The service had a threat model. The element was drawn on the diagram. No threat had ever been enumerated against it. ### Walk the chain, and stop at the first no The review is a fixed sequence of questions, each of which localises the failure to a different part of the practice: 1. **Was this component in scope?** If the system or change never triggered a model at all, the failure is in coverage and triggers, not in the method. Fix the rule that decides when modeling happens. 2. **Was it on the diagram?** If the element or the flow was missing, the model was an incomplete picture. Fix how the diagram gets built and reviewed - who confirms it matches the deployment. 3. **Were threats enumerated against it?** This is where the example fails. The element was drawn and then skipped, because it was *internal* and the analysis quietly treated the network position as trust. The remedy is enumeration discipline: every element and every flow that crosses a boundary gets its categories walked, and *internal* is a claim about the network, not a granted identity. A trust boundary that is assumed rather than drawn is the single most common shape of this branch. 4. **Was the threat enumerated but rated too low?** Then the failure is in rating: the consequence, the reachability or the attacker position was misjudged. Fix calibration - typically by re-rating the same threat with the review group and asking what input was wrong. 5. **Was it rated correctly and accepted?** Then the practice worked exactly as designed and produced a decision that turned out badly. That is a question about who is allowed to accept which risks and on what evidence, not a defect in the modeling. 6. **Was a control chosen but never shipped, or shipped without a test?** Then the failure is in follow-through from finding to change, and it belongs to how threats become tracked work. 7. **Was the model simply older than the system?** Then it is a drift failure, and the fix is in re-review cadence rather than in enumeration. Stopping at the first *no* matters. Each branch has a different remedy, and a review that concludes *we should threat model better* has not localised anything. ### The product is a change to the practice A good escaped-issue review produces one or two specific, checkable changes. For the example above, that is a rule that every administrative or operational interface is enumerated as an entry point regardless of network position, plus a check in the review step that no element on the diagram was skipped during enumeration. It may also produce a change to who reviews models, or to the trigger that requires re-modeling. It should not produce: a blame assignment, a quietly rewritten model that now shows the threat as though it had always been there, or an abandonment of the method on the strength of one escape. Rewriting the model without recording that the threat was missed destroys the only evidence the review runs on. ### Framing and cadence Run it blamelessly and on the practice, not the team - the same discipline as an operational post-incident review, and often as an appendix to one. A sensible trigger is every incident with a design-level security cause and every penetration-test finding that a model should plausibly have predicted. Keep the record: over a year, the branch that keeps coming up is the strongest evidence you have about where the method is weak. If three escapes in a row stop at step 3, the enumeration step is where to invest, not the diagramming step. ### Feeding the metric Escapes are the only lagging outcome measure a threat modeling practice has, so this review is also how the number gets its meaning. Count escapes as a rate over modeled systems, but read them as cases. And treat a period with zero escapes with suspicion: escapes are only counted when someone traces an incident back to the design, so a zero more often means the tracing habit is missing than that nothing escaped. Building the review into incident handling is what turns the metric from a decoration into a feedback loop.

  • The team says the endpoint was safe because it was internal. What do you conclude?
    That a trust boundary was assumed rather than drawn. Network position is not an identity, and a neighbouring compromised service sits on the same side of that imaginary line. The review finding is an enumeration gap: internal administrative interfaces must be treated as entry points and have their threat categories walked like any other, with the boundary made explicit on the diagram.
  • What if the threat was found, rated low and formally accepted?
    Then the practice worked as designed and the decision was wrong, which is a different problem. I take it to who is allowed to accept which risks and on what evidence, and I re-rate the threat with the group to find which input was misjudged - usually reachability or attacker position rather than consequence. Nothing about the enumeration step needs changing.
  • How do you keep the review from becoming a blame exercise?
    Frame every question about the method rather than the person, run it as an appendix to the blameless operational post-incident review, and forbid the quiet fix - nobody rewrites the model so the threat appears to have always been there. The record of which step failed is the evidence the whole feedback loop runs on, so destroying it costs more than the embarrassment saves.
  • What do you do with a year of these reviews?
    Read them for the branch that repeats. If three escapes in a row stop at enumeration, invest there - checklists, facilitation, review of skipped elements - rather than spreading effort across diagramming and rating. The pattern across escapes is stronger evidence about where the method is weak than any single case.

It is the same move as an air-accident investigation: the question is never whether the crew was careless, it is which checklist item, if it had existed, would have caught this.

saying these in an interview costs you the question

  • Concludes the team should have threat modeled better
  • Rewrites the model so the threat appears to have been found
  • Treats one escaped issue as proof the method is worthless
  • Blames the engineer instead of naming the failing step
  • Skips the review when the incident had an obvious quick fix
  • Assumes internal network position means the element is trusted

context