In incident analysis, what is the difference between the trigger of an outage and its underlying cause, and why does the distinction change what you fix?
answer
- what started it vs why it was enough
- latent condition, present all along
- the last hundred deploys were fine too
- cheap gate now, structural fix later
- triggers set exposure frequency
basics
~20 sThe trigger is the event that started this outage; the underlying cause is the latent condition that made the trigger dangerous. Removing the trigger prevents this instance, fixing the latent condition prevents the whole class — and they cost very different amounts.
solid answer
~50 sThe trigger is what happened at t0 — a deploy, a config push, a traffic spike, a node dying. The underlying cause is the condition that had been sitting in the system harmlessly and turned that ordinary event into an outage: an unbounded queue, a retry path with no backoff, a leak that only matters after fourteen days of uptime. The distinction matters because it maps to two different fixes with two different price tags. Banning or gating the trigger is usually cheap and fast — freeze that config path, add a validation check, throttle the caller. Fixing the latent condition is usually expensive and slow, but it is the only thing that stops the next trigger you haven't thought of. Good practice is to do the cheap one now to stop the bleeding and be explicit that you are accepting the latent condition until the real fix lands, rather than pretending the incident is closed because the trigger is gone.
go deeper
Be able to state plainly that the trigger is the event that started the outage and the underlying cause is the pre-existing condition that made that event dangerous, and give one concrete pair.
Explain why the latent condition survived the previous hundred harmless deploys, and describe what changes when you fix the trigger versus the condition.
Show the allocation decision: which fix you ship in the first week, which structural item you fund, and how you make carrying the latent condition an explicit, owned risk rather than an oversight.
Own the standard that says a postmortem closing with only a trigger fix is incomplete, and hold the portfolio view of knowingly-carried latent conditions across services so they are decided on rather than forgotten.
## Two different questions "What started it?" and "why was that enough to break us?" are separate questions, and postmortems that answer only the first are the most common weak postmortem in circulation. The **trigger** is the proximate event: a release went out, a certificate expired, a partner sent ten times its usual volume, a hypervisor rebooted. Triggers are usually specific, well-timestamped, and easy to identify — they are what shows up first when you overlay the change log on the SLI graph. The **underlying cause** (also called the latent condition or contributing condition) is the property of the system that had to be true for that trigger to produce user impact. It was there before the incident, it was there during the last hundred deploys that went fine, and it is still there after you revert. ## A worked example A routine config push takes the service down. Overlay the graph: the push is the trigger, timestamps line up, done. Except: - the service parses config at startup and **exits** on a parse failure, rather than logging and keeping the last known-good config - the config is delivered to **all regions simultaneously**, so every replica hit the same bad file within seconds - the health check reports the process as healthy for 30 seconds after start, so the orchestrator kept replacing pods into the same failure The bad config is the trigger. Fail-closed parsing, global simultaneous delivery, and a health check that lies are the underlying causes. Ban the specific bad field and you are safe from that exact push. Any of the three underlying conditions will convert the next imperfect config into the same outage. ## Why triggers still matter The usual overcorrection is to declare triggers uninteresting. They are not, for two reasons. First, **trigger frequency is a real dial**. If the same class of trigger fires weekly and the latent fix is a quarter of engineering time, reducing trigger frequency — validation in CI, a staged config rollout, rate-limiting the noisy caller — buys you most of the risk reduction immediately and cheaply. Second, some triggers are themselves the finding. A certificate expiring is a trigger, but "nothing tracks certificate expiry" is a systemic gap wearing a trigger's clothes. ## The decision the distinction forces After the incident you have to allocate finite engineering time, and the choice is genuinely contested: | Option | Cost | Buys you | |---|---|---| | Remove or gate the trigger | Hours to days | This exact incident cannot repeat | | Reduce trigger frequency | Days | Fewer exposures to the same latent condition | | Fix the latent condition | Weeks | The whole class of triggers becomes survivable | | Reduce blast radius | Days to weeks | Any trigger hurts a fraction of users, not all | The honest answer in most postmortems is a fast trigger fix now plus one structural item, with a named owner and a date, and an explicit statement that the latent condition is being knowingly carried in the meantime. A postmortem that lists only the trigger fix and closes is how the same outage arrives six weeks later with a different first line. ## Naming it precisely Be careful with the phrase "root cause" in this context. Many teams use it to mean the latent condition, some use it to mean the trigger, and the ambiguity produces arguments that are really about vocabulary. Saying "the trigger was X; the conditions that made X an outage were Y and Z" is unambiguous, and it signals in an interview that you have thought about this beyond the label. ## In an interview Use a concrete pair from your own experience. "The trigger was a partner doubling their request rate on a Monday; the underlying cause was that our retry policy had no jitter, so their retries synchronized and we amplified their spike into a thundering herd against the database." Then say what you fixed first and why — that is the part that shows judgment rather than vocabulary.
- The same latent condition has been in production for two years without incident. Does that make it low risk?It makes its trigger rate low so far, which is not the same thing. Risk is exposure times consequence, and the consequence side is unchanged — when it does fire it takes the service down. What actually raises the risk is anything that increases trigger frequency: a new caller, higher traffic, a change in deploy cadence. Two quiet years is evidence about the past environment, not about the next one.
- How do you decide between fixing the latent condition and just reducing the blast radius?By cost and by how many trigger classes each covers. Fixing the condition removes one specific failure mode; reducing blast radius — regional staging, cells, a fail-static fallback — caps the damage from failure modes you have not thought of yet, which is usually the better bet when the latent fix is a multi-week rewrite. In practice I ship the blast-radius change first and schedule the deeper fix.
saying these in an interview costs you the question
- Calls the trigger the root cause and closes the postmortem
- Assumes reverting the change means the risk is gone
- Treats triggers as uninteresting once a latent cause is found
- Believes a latent condition that never fired is not a real defect
- Uses "root cause" without saying which of the two is meant