skip to content

What is risk-storming as an architecture practice, how is a session actually run, and why is the first step done silently and individually?

level: middleimportance: should knowfreq 42%

answer

  1. Sticky notes placed on the architecture diagram
  2. Individual + silent first → no anchoring, no groupthink
  3. Converge, then score probability × impact
  4. Mitigations with owners, not a wall of notes
  5. Repeat as architecture changes; mixed roles in the room

basics

~20 s

Risk-storming is a group technique where people mark risks directly on the architecture diagrams. Everyone first writes risks alone and silently, then all notes are placed on the diagram, discussed, and ranked by probability and impact so the riskiest spots are visible and get mitigation owners.

solid answer

~50 s

Risk-storming (popularised by Simon Brown) is a collaborative, visual risk-identification technique run over architecture diagrams — typically a context and container view. The steps: (1) agree the diagrams everyone will look at; (2) each participant, working **individually and in silence**, writes each perceived risk on its own sticky note; (3) all notes are placed on the diagram element they concern; (4) the group converges — deduplicating, clarifying, and scoring each risk by probability and impact, usually on a low/medium/high grid; (5) the highest-priority risks get named mitigations and owners, feeding the backlog and architecture decisions. The individual, silent first pass is essential: it prevents anchoring on the loudest or most senior voice and preserves independent perspectives, so rarely-voiced risks survive to the board. Density of notes shows hotspots instantly. It is repeated as the architecture evolves, and pairs naturally with risk-driven design: do just enough architecture work to retire the highest risks.

go deeper

for a junior

Describe the mechanics: agree the diagrams, write risks individually on notes, place them on the diagram, discuss, then rank by probability and impact.

for a middle

Explain why the silent individual pass exists (anchoring, authority bias, coverage) and insist that every high-priority risk leaves with a mitigation and an owner.

for a senior

Connect it to risk-driven design — architecture effort is allocated to retire the top risks — and discuss cadence, participant mix, risk-versus-issue discipline, and expressing impact in business terms so mitigation competes fairly with features.

for a principal

Treat it as one input to a standing risk posture: aggregate hotspots across systems, feed them into architecture decision records and investment cases, wire recurring risks into automated checks or monitoring, and be explicit about which risks are formally accepted, by whom, and until when.

## What it is **Risk-storming** is a lightweight, collaborative technique for identifying and prioritising risk in a software architecture, described by Simon Brown (author of the C4 model). Instead of a risk workshop producing an abstract spreadsheet, risks are attached **visually to the architecture itself** — sticky notes (physical or virtual) placed on the diagram element they threaten. **Risk** here means an uncertain future event with a negative consequence — a component that may not meet its throughput target, an integration whose partner may be unavailable, a data store that may not survive a region outage, a technology nobody on the team has run in production. ## Running a session 1. **Agree the canvas.** Put up the diagrams everyone will annotate — commonly a system-context diagram and one or more container/component diagrams. Everyone must understand the notation before risks are collected; ambiguity in the diagram becomes ambiguity in the risks. 2. **Identify risks individually and in silence.** Each participant writes risks on their own, one risk per note, for a fixed period. No discussion. 3. **Place the notes.** Participants attach each note to the diagram element it concerns. Notes that belong to no element usually reveal something missing from the diagram — itself a finding. 4. **Converge.** Walk the board together: merge duplicates, clarify wording, split compound risks. Duplicate notes from independent participants are a signal of shared concern, not noise to discard before scoring. 5. **Prioritise.** Score each risk on **probability × impact**, typically low/medium/high on each axis, often colour-coded — red for high priority, amber for medium, green for low. 6. **Mitigate.** For the top risks, agree concrete actions with owners: a spike, a load test, a fallback design, a bulkhead, a written decision record, or an explicit accept-and-monitor decision. These become backlog items, not workshop artefacts. ## Why the first pass is silent and individual This is the load-bearing design choice of the technique, and interviewers probe it: - **Anchoring** — the first risk voiced frames everyone's subsequent thinking; independent writing prevents it. - **Authority bias** — the architect or most senior engineer speaking first suppresses dissent from those closest to the code or to operations. - **Coverage** — diverse participants (developers, testers, operations, security, product, support) hold non-overlapping risk models; group discussion tends to converge early and lose the tails. - **Signal from duplication** — if five people independently flag the same component, that convergence is evidence, which is only meaningful if the writing was independent. Groupthink is the failure mode being engineered out: a room that discusses first typically produces fewer, more conventional risks. ## Who to invite Mixed roles beat a room of architects. Operations people surface failure and recovery risks; testers surface observability and verifiability risks; support surfaces what actually breaks for users; security surfaces trust-boundary risks; product surfaces risks about assumptions on volume and change. Keep it small enough for everyone to speak during convergence. ## Where it fits - **Risk-driven design** (George Fairbanks): choose architecture *effort* by risk — do the design work that retires your biggest risks and no more. Risk-storming supplies the prioritised risk list that drives that choice. - **Iteration.** Run it early (before committing to structure), again when the architecture changes materially, and after significant incidents. Date the output; a risk board from a year ago describes a system that no longer exists. - **Relationship to other practices.** It complements rather than replaces threat modelling (security-specific, typically a structured method such as STRIDE over data-flow diagrams), failure-mode analysis, and pre-mortems. Risk-storming is broader and cheaper; a hotspot it reveals often triggers one of the deeper techniques. ## Outputs and pitfalls Good outputs: a prioritised risk list tied to diagram elements, mitigations with owners and dates, and often diagram corrections. Common pitfalls: no scoring (a wall of notes with no priority is not a decision); no owner (risks re-identified verbatim next session); scoring theatre (invented numeric probabilities implying precision nobody has — coarse low/medium/high is more honest); treating it as a one-off ceremony; and confusing risks with issues — an issue is already happening and belongs on the backlog now, a risk is uncertain and needs a probability.

  • How do you prioritise the risks once they are all on the board?
    Score each on probability and impact — coarse low/medium/high is usually enough — and treat the high-probability/high-impact quadrant as the immediate work. Impact should be expressed in business terms (revenue, data loss, regulatory exposure, recovery time), not just technical severity, so mitigation can be compared against feature work.
  • What is the difference between a risk identified this way and an issue?
    A risk is uncertain — it may or may not occur, so it carries a probability and is handled by mitigation, transfer, avoidance or explicit acceptance. An issue is already true today and simply needs fixing. Mixing them corrupts prioritisation because certainties always dominate probabilistic scoring.
  • How does risk-storming relate to threat modelling?
    Threat modelling is a security-specific, adversarial analysis, usually structured (for example STRIDE applied to data-flow diagrams and trust boundaries). Risk-storming is broader and cheaper, covering performance, availability, operability, people and delivery risks as well. They are complementary: risk-storming often surfaces a trust-boundary hotspot that then justifies a full threat-modelling session.

It is a fire-safety walkthrough of a building plan rather than a meeting about fire in the abstract: each person walks the blueprint alone marking where a fire could start or spread, then everyone stands back and sees which rooms are covered in markers. The clusters — not anyone's opinion — tell you where to install the sprinklers.

saying these in an interview costs you the question

  • Starting with open group discussion, which anchors the room on the first or most senior voice
  • Producing an unranked wall of sticky notes with no probability/impact scoring
  • Recording risks with no owner or mitigation, so the same notes reappear next session
  • Running it once at project start and never again as the architecture evolves
  • Inviting only architects, losing the operations, testing, security and support perspectives
  • Assigning falsely precise numeric probabilities that imply data nobody has

context