skip to content

Root Causes & Contributing Factors

Reconstructing what actually happened and why — beyond a single convenient root cause. Interviewers probe your analysis method because '5 whys found the bug' answers miss the systemic factors that let the bug reach production.

on this pageshow

questions

6

In incident analysis, what is the difference between the trigger of an outage and its underlying cause, and why does the distinction change what you fix?

level: juniorimportance: must knowfreq 50%

answer

  1. what started it vs why it was enough
  2. latent condition, present all along
  3. the last hundred deploys were fine too
  4. cheap gate now, structural fix later
  5. triggers set exposure frequency

basics

~20 s

The trigger is the event that started this outage; the underlying cause is the latent condition that made the trigger dangerous. Removing the trigger prevents this instance, fixing the latent condition prevents the whole class — and they cost very different amounts.

solid answer

~50 s

The trigger is what happened at t0 — a deploy, a config push, a traffic spike, a node dying. The underlying cause is the condition that had been sitting in the system harmlessly and turned that ordinary event into an outage: an unbounded queue, a retry path with no backoff, a leak that only matters after fourteen days of uptime. The distinction matters because it maps to two different fixes with two different price tags. Banning or gating the trigger is usually cheap and fast — freeze that config path, add a validation check, throttle the caller. Fixing the latent condition is usually expensive and slow, but it is the only thing that stops the next trigger you haven't thought of. Good practice is to do the cheap one now to stop the bleeding and be explicit that you are accepting the latent condition until the real fix lands, rather than pretending the incident is closed because the trigger is gone.

go deeper

for a junior

Be able to state plainly that the trigger is the event that started the outage and the underlying cause is the pre-existing condition that made that event dangerous, and give one concrete pair.

for a middle

Explain why the latent condition survived the previous hundred harmless deploys, and describe what changes when you fix the trigger versus the condition.

for a senior

Show the allocation decision: which fix you ship in the first week, which structural item you fund, and how you make carrying the latent condition an explicit, owned risk rather than an oversight.

for a principal

Own the standard that says a postmortem closing with only a trigger fix is incomplete, and hold the portfolio view of knowingly-carried latent conditions across services so they are decided on rather than forgotten.

## Two different questions "What started it?" and "why was that enough to break us?" are separate questions, and postmortems that answer only the first are the most common weak postmortem in circulation. The **trigger** is the proximate event: a release went out, a certificate expired, a partner sent ten times its usual volume, a hypervisor rebooted. Triggers are usually specific, well-timestamped, and easy to identify — they are what shows up first when you overlay the change log on the SLI graph. The **underlying cause** (also called the latent condition or contributing condition) is the property of the system that had to be true for that trigger to produce user impact. It was there before the incident, it was there during the last hundred deploys that went fine, and it is still there after you revert. ## A worked example A routine config push takes the service down. Overlay the graph: the push is the trigger, timestamps line up, done. Except: - the service parses config at startup and **exits** on a parse failure, rather than logging and keeping the last known-good config - the config is delivered to **all regions simultaneously**, so every replica hit the same bad file within seconds - the health check reports the process as healthy for 30 seconds after start, so the orchestrator kept replacing pods into the same failure The bad config is the trigger. Fail-closed parsing, global simultaneous delivery, and a health check that lies are the underlying causes. Ban the specific bad field and you are safe from that exact push. Any of the three underlying conditions will convert the next imperfect config into the same outage. ## Why triggers still matter The usual overcorrection is to declare triggers uninteresting. They are not, for two reasons. First, **trigger frequency is a real dial**. If the same class of trigger fires weekly and the latent fix is a quarter of engineering time, reducing trigger frequency — validation in CI, a staged config rollout, rate-limiting the noisy caller — buys you most of the risk reduction immediately and cheaply. Second, some triggers are themselves the finding. A certificate expiring is a trigger, but "nothing tracks certificate expiry" is a systemic gap wearing a trigger's clothes. ## The decision the distinction forces After the incident you have to allocate finite engineering time, and the choice is genuinely contested: | Option | Cost | Buys you | |---|---|---| | Remove or gate the trigger | Hours to days | This exact incident cannot repeat | | Reduce trigger frequency | Days | Fewer exposures to the same latent condition | | Fix the latent condition | Weeks | The whole class of triggers becomes survivable | | Reduce blast radius | Days to weeks | Any trigger hurts a fraction of users, not all | The honest answer in most postmortems is a fast trigger fix now plus one structural item, with a named owner and a date, and an explicit statement that the latent condition is being knowingly carried in the meantime. A postmortem that lists only the trigger fix and closes is how the same outage arrives six weeks later with a different first line. ## Naming it precisely Be careful with the phrase "root cause" in this context. Many teams use it to mean the latent condition, some use it to mean the trigger, and the ambiguity produces arguments that are really about vocabulary. Saying "the trigger was X; the conditions that made X an outage were Y and Z" is unambiguous, and it signals in an interview that you have thought about this beyond the label. ## In an interview Use a concrete pair from your own experience. "The trigger was a partner doubling their request rate on a Monday; the underlying cause was that our retry policy had no jitter, so their retries synchronized and we amplified their spike into a thundering herd against the database." Then say what you fixed first and why — that is the part that shows judgment rather than vocabulary.

  • The same latent condition has been in production for two years without incident. Does that make it low risk?
    It makes its trigger rate low so far, which is not the same thing. Risk is exposure times consequence, and the consequence side is unchanged — when it does fire it takes the service down. What actually raises the risk is anything that increases trigger frequency: a new caller, higher traffic, a change in deploy cadence. Two quiet years is evidence about the past environment, not about the next one.
  • How do you decide between fixing the latent condition and just reducing the blast radius?
    By cost and by how many trigger classes each covers. Fixing the condition removes one specific failure mode; reducing blast radius — regional staging, cells, a fail-static fallback — caps the damage from failure modes you have not thought of yet, which is usually the better bet when the latent fix is a multi-week rewrite. In practice I ship the blast-radius change first and schedule the deeper fix.

saying these in an interview costs you the question

  • Calls the trigger the root cause and closes the postmortem
  • Assumes reverting the change means the risk is gone
  • Treats triggers as uninteresting once a latent cause is found
  • Believes a latent condition that never fired is not a real defect
  • Uses "root cause" without saying which of the two is meant

context

open as a page

During a postmortem your 5 Whys chain ends at "the config parser didn't validate the field." Why do experienced reviewers push back on stopping there, and what analysis do you do instead?

level: middleimportance: must knowfreq 65%

basics

~20 s

5 Whys follows one causal line and stops at the first fixable defect, so it names the bug but not the review, test, rollout and detection gaps that let the bug reach production. Ask "what else" at each step and record multiple contributing factors.

open as a page

How do you reconstruct the timeline of an incident for its postmortem, and which timestamps in that timeline matter most?

level: middleimportance: should knowfreq 60%

basics

~20 s

Rebuild the timeline from machine evidence — deploy records, alert history, chat transcripts, dashboards — normalized to UTC. The load-bearing timestamps are impact start, detection, human engagement, first mitigation and impact end, because their gaps measure detection and mitigation separately.

open as a page

A draft postmortem lists fourteen contributing factors for one outage. How do you decide which ones belong in the final analysis and which are noise?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Keep a factor if changing it would plausibly have prevented or materially reduced the impact, it can recur beyond this incident, and someone has a lever on it. Drop factors visible only in hindsight and factors so generic they would fit any postmortem.

open as a page

Three outages this quarter had different proximate causes, but all three involved the same shared configuration service. How do you approach the analysis across them?

level: principalimportance: should knowfreq 30%

basics

~20 s

Analyze the set, not the incidents. Tag contributing factors across postmortems, find the factor that repeats, normalize for how much traffic that component actually carries, then argue a structural investment using aggregate impact rather than three separate small fixes.

open as a page

A canary deploy caught a fatal bug before any customer saw an error. Do you analyze it like an incident, and what would that analysis look at?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Yes, when the margin was thin or the save was luck rather than design. The analysis asks how close it came, whether the control that caught it was intentional and repeatable, and what would have happened had it fired ten minutes later.

open as a page