skip to content

Tell me about a time a change you shipped broke production.

level: middleimportance: must knowfreq 58%

answer

  1. the change, and what it broke
  2. how you found out, how fast
  3. mitigate first, own it out loud
  4. root cause, then contributing factors
  5. the guardrail and its proof

basics

~20 s

Tests incident instincts and ownership under pressure. Answer with one outage your own change caused: how fast you knew, mitigation before diagnosis, a plain statement that it was yours, the real root cause, and the guardrail that followed.

how to answer

5 beats
  1. the change you shipped and what it broke
    Two or three sentences: what the change was meant to do, where it ran, and who felt the failure. Keep setup and framing to roughly fifteen to twenty percent of your airtime so the incident itself gets the room.
  2. how you found out, and how fast
    Name the signal and the interval between shipping and knowing. If a user or a downstream team noticed before your monitoring did, say so directly; that admission is what makes the alert you added later believable.
  3. what you did first to stop the damage
    This beat and the next are the bulk of the answer, around sixty percent together. Show mitigation before diagnosis: the revert, the rollback, the traffic shift. Say in one plain sentence that the change was yours, then keep moving.
  4. the root cause, and the factors around it
    Explain the actual mechanism in language a non-specialist interviewer can follow, then list the contributing conditions separately: the environment that could not reproduce it, the missing gate, the review nobody who owned that code attended.
  5. the cost, and the guardrail that followed
    Close with a number for impact or recovery and one specific thing that now exists: a test, an alert, a staged rollout. Give roughly the last quarter of your airtime here, and if the guardrail has caught something since, end on that.

your answer

5 story prompts
pick a story
  • Pick an incident your own change caused, not one you merely helped somebody else recover from.
  • Choose one recent enough that you can still name the detection time and the recovery time.
  • Write out the first thirty minutes minute by minute before you draft any of the narration.
  • Have one impact number and one recovery number ready, reconstructed honestly if you must.
  • If the guardrail you added has caught something since, make that your closing sentence.

draft and rehearse your own answer in a learn session

go deeper

This prompt probes ownership and composure under production pressure. The interviewer wants to see whether your reflex is to restore service before satisfying curiosity, whether you can state that the change was yours without either deflecting or spiralling, and whether you distinguish a real root cause from the conditions around it. A strong answer proves you leave the system harder to break than you found it.

at middle level

I maintain the autoscaler in an open-source build farm that a lot of downstream projects run their CI against. I shipped a change that replaced raw queue depth with a smoothed average as the scale-up trigger, because we were churning nodes badly. It passed review and the unit suite and went out with the release we cut that week. About twenty minutes later the queue-wait alert fired. Autoscale delay had gone from a p95 of 38 seconds to just under seven minutes, and 1,463 community jobs were sitting unscheduled. My first move was not to debug it. I reverted the commit, redeployed, and we were draining the backlog inside two minutes. Then I posted in the public incident thread that the regression was mine, that it was already reverted, and which jobs would need re-running. Once it was quiet I worked out why. The smoothing window I chose was five minutes, so a burst had to persist through most of a window before the average crossed the threshold. Raw depth spikes on the very first sample; a five-minute average of it barely moves, and by the time it did move the queue had already built. The contributing factor was that our staging farm only ever sees steady load, so no test we had could have caught it. We were degraded for 71 minutes end to end. I added a load-shape test that replays a recorded burst against the autoscaler and fails if the delay crosses two minutes. It has since caught one regression from another contributor before it reached the fleet.

why this lands

The signal here is sequencing: the revert happens before the diagnosis, and the ownership sentence is said once, in public, without drama. Detection time, backlog size and degraded duration make the blast radius concrete, and the guardrail earns its place by having fired since. Stopping at a config issue instead of naming the window-length mechanism would downlevel this.

at senior level

I own the scaling subsystem in an open-source infrastructure project, and I was driving a node-image rollout across our runner pools using cordon logic I had written myself. That logic drained old nodes faster than the pool could replace them, so during the rollout three of our nine pools could not add capacity fast enough. Scale-up on those pools went from well under a minute to roughly eleven, while the other six looked perfectly healthy, which is why the first reports were confusing. The judgment call was mine because I was running the rollout. I froze it at the third pool rather than pushing forward on the theory that the next one would behave, and put the drained pools back on the previous image. Normal service came back 48 minutes after the first report. While that ran I asked another maintainer to own the public status updates so I could stay on recovery, and I messaged the two downstream projects with pinned runners directly, since their jobs were the ones failing outright. The cause was my assumption that the replacement path was synchronous; it is not. The contributing factors were a rollout with no per-pool gate and alerting only on job failures, which is a lagging signal. What I changed matters more than the fix. Image rollouts now advance one pool at a time behind a synthetic probe that measures scale-up every ninety seconds and halts on a threshold breach. Four rollouts since, no queue impact.

why this lands

Seniority shows in the decision under ambiguity — freeze rather than roll forward — and in handing comms to someone else so recovery keeps its owner. Naming the false assumption is stronger than naming the bug. Prevention that changes how the team ships, with a count of clean rollouts since, is the closing evidence; a personal resolution to be careful would downlevel it.

for a junior

You are not expected to have owned a rollback. Show that you noticed the signal, escalated immediately instead of quietly investigating, and can explain the mechanism of your own bug afterwards in your own words.

for a middle

Own the whole loop on your own change: detection, the revert or mitigation you executed, the diagnosis, and one concrete test or alert you added. The interviewer wants to hear you choose mitigation before curiosity.

for a senior

Show judgment calls, not just steps: roll forward or back, who communicates while you recover, when to declare. Then show the prevention landing in the system itself, and name what the incident cost in time or trust.

for a principal

Treat your incident as evidence about a class of failure. Show the mechanism you introduced so this shape of outage is caught for everyone, who else adopted it, and how you know it works rather than that it exists.

saying these in an interview costs you the question

  • Choosing an incident someone else caused and calling it yours only loosely
  • Debugging for an hour before mitigating, then presenting that as diligence
  • Blaming the reviewer, the pipeline or the on-call rotation for your change
  • No impact number and no recovery time, so the blast radius stays vague
  • A root cause that stops at a config issue or a bad deploy
  • Prevention described as being more careful next time

  • How long was it between your change going out and you knowing something was wrong?
    Give a real interval, even an embarrassing one, and say what surfaced it: your own alert, a dashboard, or someone else complaining. If a user found it before your monitoring did, say so plainly and connect it to what you added afterwards. Detection time is the number interviewers use to judge whether your instrumentation story is genuine.
  • Who did you tell, and how soon?
    Name the audiences and the order: whoever is recovering with you, then whoever is affected downstream. Show that you communicated while mitigating rather than after, and that you said the change was yours without a paragraph of apology. If someone else took comms so you could keep working, that is a strength, not an admission.
  • What would you do differently if the same change went out tomorrow?
    Answer at the level of the system, not your attention span. Point at the missing gate, the environment that could not reproduce the load, or the signal you were not alerting on. One honest personal change is fine alongside it, but a purely personal answer reads as no learning at all.
  • Has the guardrail you added ever caught anything since?
    If it has, this is your strongest closing sentence, so have the instance ready. If it has not, say that honestly and say how you would know it works, or admit it may be theatre. Interviewers ask this specifically to separate real prevention from a line added to a postmortem and never exercised.

## One story, several wordings This prompt arrives in several wordings that all want the same thing: - *tell me about a time you broke production* - *walk me through an outage you caused* - *describe a change that had unintended consequences* - or the softer *what is the biggest thing you have broken*. Treat them as one story with one adjustable opening sentence. A related but distinct wording, *tell me about an outage you were involved in*, lets you use an incident you only helped resolve; if the interviewer says caused or your change, do not quietly substitute that easier story, because the whole point of the question is watching you stand next to your own mistake. ## The evaluation axis is not remorse Interviewers are checking three things in sequence. 1. **First**, whether your instinct under pressure is to restore service or to satisfy curiosity: the strongest answers contain a moment where the speaker says they stopped investigating and reverted. 2. **Second**, whether you can separate the root cause from the contributing factors without either collapsing everything into one bad line of code or dissolving your own responsibility into a fog of process gaps. 3. **Third**, whether anything about the system changed as a result, and whether you can prove it. ## Where weak answers go wrong A weak answer usually fails structurally rather than morally. It spends ninety seconds on architecture background, twenty on the failure, and finishes with we fixed it and I learned to be more careful. Aim for roughly: - fifteen to twenty percent of your airtime on setup, - sixty percent on what you actually did in the incident and in the diagnosis, - and the remaining twenty-odd percent on outcome and prevention. If the interviewer has to ask how bad was it, your result beat was too thin. ## The numbers that carry the story The evidence that carries this story is numerical and small: - time to detection, - time to mitigation, - the size of the affected population, - and one number showing the guardrail works. None of these need to be impressive. An answer that says users noticed before we did, and that is why the first thing I built afterwards was the alert we did not have, is far stronger than one claiming instant detection with nothing to show for it. ## The ownership sentence deserves its own attention Say once, in plain words, that the change was yours. Then stop. Candidates who apologise repeatedly read as anxious rather than accountable, and candidates who narrate in the passive voice, the config was applied, the pool was drained, read as evasive. **First person, active verbs, one admission**, then move to the work. ## How the bar moves by level The bar moves by level in a specific way. - **At the junior end**, nobody expects you to have held the pager or made the rollback call; noticing, escalating fast, and understanding your bug afterwards is a complete answer. - **In the middle band** the interviewer wants the full loop executed by you, including one guardrail you personally added. - **At senior**, the interesting material is the judgment calls with no obviously right answer, roll forward versus back, when to declare, who talks to affected users while you recover, and prevention that changes how the team ships rather than how you personally behave. - **At the top band** the incident is a sample of a class, and the answer is about the mechanism you introduced so the class is caught without you in the room. ## One last trap The humble outage. Picking an incident with no real impact so you can look calm about it backfires, because the interviewer reads the choice as an unwillingness to be evaluated. Pick something that genuinely cost people time, and let the recovery and the prevention carry the impression.

context