What do time to detect and time to repair a broken shared-branch build measure, and how do you shorten each?
answer
- Two clocks, not one number
- Detection includes nobody reading it
- Repair contains another whole run
- Restore green before understanding why
- Read pass rate with two other numbers
basics
~20 sTime to detect runs from a change landing to a human seeing a credible red signal; time to repair runs from that signal back to green. Shorten detection with earlier checks and addressed notifications, repair with revert-first and clear ownership.
solid answer
~50 sTime to detect is not the suite runtime: it is queue wait plus the runtime of the tier that catches this defect class, plus notification delay, plus the delay before a human actually reads it - and that last stage is usually the largest and the least measured. Time to repair covers triage, attribution across whatever changes were batched into the run, finding an owner, the fix or revert, and one more full pipeline cycle to confirm, so a slow suite is charged twice per incident. Instrument the stages separately, and report median with a high percentile, because a handful of multi-hour episodes is what the team remembers and what a mean hides. Shorten detection by moving checks earlier and notifying a person rather than a channel; shorten repair with a revert-first default and unambiguous ownership, so restoring green never waits on understanding the defect.
go deeper
Be able to say where each clock starts and stops, and why a red shared branch blocks the whole team rather than just its author. Knowing that reverting is a normal first response, not an admission of failure, is expected.
Break each clock into stages and say which stage you would instrument first. Explain why a slow suite is paid twice per incident, once in detection and once in the confirming run.
Bring numbers from a real team and a distribution rather than an average. Show that you found the process bottleneck - unread notifications, ambiguous ownership, batched changes destroying attribution - rather than only tuning the pipeline.
Frame the red branch as a shared outage with a cost per minute across every engineer, and defend the conventions you would set: stop-the-line, revert-first, ownership assigned before the incident, and which of the three headline numbers you would put in front of leadership.
### Two clocks, and the gap between them When a change lands on a shared branch and the pipeline goes red, two intervals describe the damage. **Time to detect (TTD)** runs from the moment the change lands to the moment a human sees a credible red signal. It is not the suite's runtime. It is the sum of: queue wait for a worker, the runtime of whichever tier catches this class of defect, the delay before a notification is delivered, and — the part teams forget to measure — the delay before anybody *reads* it. A team with a nine-minute suite and a chat channel nobody watches can have a TTD of several hours. **Time to repair (TTR)** runs from that credible signal to the branch being green again. It contains triage (is this real or unreliable?), attribution (which of the changes in this run did it?), ownership (whose is it, are they awake?), the fix or revert itself, and one more full pipeline cycle to confirm. Note that TTR contains a *whole extra suite runtime* at the end, which is why a long suite is charged twice: once to detect and once to confirm the repair. ### Where the time actually goes Instrument the stages rather than guessing. In a ride-hailing dispatcher team running a 27-minute shared-branch tier, a month of red episodes broke down roughly like this: | Stage | Median | p90 | | --- | --- | --- | | Land to signal produced (queue + run) | 31 min | 1 h 12 | | Signal produced to signal read | 6 min | 2 h 14 | | Triage: real or unreliable | 9 min | 48 min | | Attribution across batched changes | 12 min | 1 h 05 | | Fix or revert authored | 14 min | 3 h 20 | | Confirming run | 27 min | 34 min | The medians look tolerable and the p90s are the real story: the branch was red for over six hours several times, and almost none of that was writing code. The dominant costs were *nobody read it* and *nobody could tell whose it was*. Both are process defects, not engineering ones, and both are invisible if you only track a single end-to-end number. ### Shortening each clock For TTD: move the check that catches this defect class into an earlier and cheaper tier; cut queue wait by capacity or by prioritising shared-branch runs over speculative ones; make the notification land where the author already is and address it to a person rather than a channel; and integrate in smaller pieces, because a smaller change is its own attribution. For TTR: adopt a **revert-first** default so restoring green is decoupled from understanding the defect — the investigation continues on a branch while everyone else is unblocked. Make ownership unambiguous before the incident, so triage does not begin with a search for a person. Separate the unreliability question from the defect question quickly, using recorded history rather than opinion. And keep the confirming run fast, since it sits inside every repair. ### Pipeline pass rate, and why it is the trickiest of the three **Pass rate** is the share of runs on the shared branch that went green. It is the number most often put on a dashboard and the one most often misread, because it is *ambiguous in both directions*. - A low pass rate can mean the suite is catching real defects — which is the suite working — or that it is unreliable, or that people integrate broken work. - A very high pass rate can mean the branch is healthy, or that the gate has been narrowed until it no longer asks anything hard, or that everything interesting was moved to a non-blocking tier. So pass rate is only interpretable in a triple: pass rate, flake rate and time to repair. A low pass rate with low flake rate and short TTR is a healthy team catching real problems fast. A low pass rate with high flake rate is a suite nobody believes. A high pass rate with a rising escape of defects past the pipeline means the gate has stopped asking. Track all three as trends, and report TTD and TTR as distributions — median plus a high percentile — rather than as averages, because a handful of multi-hour episodes is exactly what a mean hides and exactly what the team actually experiences. ### The interviewer's real question Behind this question is usually: *do you understand that a broken shared branch is a shared outage?* While it is red, everyone else's verdict is worthless, so they either stop integrating or integrate on top of a known-bad base. That is why teams adopt an explicit stop-the-line convention, why revert-first exists, and why TTR at p90 is a better health indicator than any coverage number.
- Your shared-branch pipeline pass rate is 97%. Is that a good sign?Not on its own. A very high pass rate is equally consistent with a healthy branch and with a gate that has stopped asking anything hard - checks moved to non-blocking tiers, slow or awkward cases removed, or a suite so narrow it cannot fail. Read it alongside flake rate and time to repair, and alongside how often defects reach production anyway. A pass rate that rises while defects escape is a warning, not an achievement.
- Why prefer reverting a change to fixing forward when the shared branch is red?Because a red shared branch is a shared outage: every other person's verdict is worthless until it is green, so they either stop integrating or build on a known-bad base. Reverting decouples restoring the signal from understanding the defect, which usually takes far longer. The author keeps investigating on their own branch with no clock running on everybody else. Fix forward only when the revert is riskier than the defect, such as when the change is already partially rolled out.
- Which stage of time to detect do teams most often fail to measure?The delay between the failing signal being produced and a human reading it. Teams instrument queue wait and suite runtime because those come from the pipeline for free, then assume detection ends when the notification is sent. In practice a message to a broad channel outside working hours can add hours, and it is often the single largest contributor. Measure signal-produced to first-human-acknowledgement separately, and route the notification to the change author by name.
saying these in an interview costs you the question
- Equates time to detect with the suite's runtime
- Forgets that repair includes one more confirming pipeline run
- Reports mean time to repair and hides multi-hour episodes
- Reads a high pipeline pass rate as proof of health on its own
- Insists on fixing forward while everyone else is blocked
- Sends failures to a broad channel instead of to the change author