For six months your team has skipped a required pre-deploy verification step because it is slow, and nothing has broken. What is this pattern called, and what do you do about it?
answer
- success is read as permission
- margin consumed, not risk removed
- the document now describes a fiction
- fund it or formally kill it
- expiry dates on every exemption
basics
~20 sNormalization of deviance: a shortcut that keeps working quietly becomes the standard, and the margin it consumed stays invisible until it runs out. Close the gap between written and actual practice — make the step cheap enough to keep, or formally retire it after an explicit risk decision.
solid answer
~50 sThat is normalization of deviance, the pattern Diane Vaughan described in her analysis of the Challenger launch decision. The mechanism is that each successful deviation is read as evidence the rule was unnecessary, so absence of failure gets mistaken for presence of safety — even though what was actually consumed was margin, not risk. The dangerous state is not the shortcut itself; it is the gap between what the procedure says and what the team does, because responders during an incident will assume the step ran. So I force one of two honest outcomes. Either the control is worth keeping, in which case we pay to make it fast enough that nobody wants to skip it, or it is not, in which case we delete it from the procedure with a named owner and a written statement of the risk we are now accepting. What I refuse to do is leave the documented process lying.
go deeper
Know the name and the basic shape: repeatedly getting away with a shortcut makes it feel safe, even though nothing about the underlying risk has changed.
Explain the loop mechanically — success being misread as evidence the rule was unnecessary — and why the gap between documented and actual procedure is more dangerous than the skipped step itself.
Show the forced choice you would make and its cost: fund the control until it is cheap enough to follow, or retire it in writing with the accepted risk named. Bring a concrete detection habit you have used.
Speak to how drift is generated organisationally — pressure, unfunded controls, exemptions without expiry — and how you build a review cadence that surfaces it before an incident does, without turning the disclosure into a punishment.
## The mechanism Normalization of deviance, named by sociologist Diane Vaughan in her study of the Challenger launch decision, describes how an organisation gradually comes to accept a departure from its own standard. The loop is disarmingly rational at every step: 1. A rule is expensive to follow, so someone skips it under pressure. 2. Nothing bad happens. 3. The absence of harm is taken as evidence that the rule was over-cautious. 4. The shortcut becomes routine, then becomes the actual standard, then becomes invisible. The flaw is in step 3. A safety control does not usually pay out on any given day; it pays out on the rare day the other defences fail. Skipping it does not remove risk, it removes *margin* — and margin is invisible until it is gone. Six clean months is exactly what the loop predicts, not evidence against the concern. ## Why the documented gap is the actual hazard The worst part is not the missing step; it is that the written procedure now describes a system that does not exist. That matters most during an incident, when people reason from documents under time pressure. A responder debugging a bad release will assume the pre-deploy verification ran and will therefore rule out an entire class of cause — the exact class that is now unguarded. Drift converts your documentation from an asset into a trap, and it does so silently. The same applies to any control that has quietly stopped functioning: a check that is skipped, an approval that is rubber-stamped, a suppression created "for this week" two years ago, a failover that nobody has exercised since the person who wrote it left. ## How you detect it Drift is invisible from inside, because everyone involved considers the current behaviour normal. Practical detections: - **Audit written versus actual.** Take a procedure and ask when each step was last performed as written. "We always just skip that one" is the whole finding. - **Listen for the phrase "we always just…"** It reliably marks a deviation that has become invisible to the people describing it. - **Look at expiry.** Suppressions, exemptions and temporary approvals with no expiry date are drift generators; anything granted as temporary should die automatically unless renewed deliberately. - **Ask new joiners.** People in their first month still see the gap between the onboarding document and what the team actually does. That window closes fast, so harvest it. - **Watch incident behaviour.** If responders discover mid-incident that an assumed control was not in effect, you have found drift the expensive way. ## The decision, and its cost Once you find it there are exactly two defensible outcomes, and both cost something: **Restore the control.** If the step guards a real failure mode, it stays — but you have to pay for the reason it was skipped. A control everyone dodges because it takes twenty minutes is a design problem, not a discipline problem. Make it fast, make it automatic so it cannot be skipped, or make it run in parallel with something people are already waiting for. Re-declaring the rule without removing the friction simply restarts the loop, and the second time round people also learn that leadership's reminders are noise. **Formally retire it.** Sometimes the step genuinely no longer earns its cost — the failure mode it covered was engineered away, or another control now catches the same class of problem. Then delete it from the procedure, in writing, with a named owner and an explicit statement of the risk now accepted and how you would detect that risk materialising. This is a legitimate outcome, and treating it as one is what makes it safe for people to raise drift in the first place. What you cannot do is the comfortable third option: acknowledge the shortcut, leave the document unchanged, and rely on everyone knowing. That is the state you started in. ## Why this belongs to blameless culture Drift is a systems phenomenon, not a discipline failure, and it is only ever discovered by people admitting what they actually do. If reporting "we haven't run that step since March" gets someone reprimanded, you will never hear about the next one, and the organisation loses its only detection mechanism. The correct reaction to the disclosure is visible gratitude followed by a decision — which is also the strongest evidence a team can offer that its blamelessness is real rather than declared. ## Answering it well Name the pattern, explain why six clean months is predicted rather than reassuring, distinguish margin from risk, and land on the forced choice: fund the control or formally kill it, but never leave the written and actual processes disagreeing. Adding one concrete detection habit — expiry dates on every exemption, or interviewing new joiners — signals that you have actually hunted for this rather than read about it.
- The team argues six clean months proves the step was unnecessary. How do you answer?By separating margin from risk. A control that guards a rare compound failure will look useless on almost every individual day, so a clean streak is weak evidence either way. I ask what specific failure mode the step catches and whether anything else now catches it. If something does, retire the step deliberately; if nothing does, we have been running without a defence and calling it a result.
- How do you stop new exemptions from becoming permanent drift?Give every exemption, suppression and waiver an expiry date and an owner at the moment it is granted, and let it lapse by default. Renewal should require someone to actively restate the justification. Temporary things without expiry dates are how permanent things get created without anyone deciding to create them.
- Is normalization of deviance ever visible from inside the team?Rarely, because the deviation has become the definition of normal for everyone who lives with it. The reliable observers are new joiners in their first weeks, engineers rotating in from another team, and audits that compare a document against what people actually did last quarter. That is why external eyes and written procedures both matter.
saying these in an interview costs you the question
- Nothing has broken, so the step was clearly unnecessary
- We all know we skip it, so the document is fine
- Just remind everyone to follow the process
- The engineer skipping it should be reprimanded
- Add it to onboarding training and move on