skip to content

Your team has stopped trusting red builds because most failures turn out to be flakes. How do you fix that?

level: principalimportance: should knowfreq 42%

answer

  1. Trust is the asset
  2. Decide what red must oblige
  3. Stop the inflow before draining
  4. Keep one gate people believe
  5. Judge by behaviour, not colour

basics

~20 s

Treat the lost trust as the problem, not the red runs. Decide what a red result must oblige, stop new nondeterministic cases entering the blocking path, fund the cleanup as real work, and keep one gate whose failures the team believes.

solid answer

~50 s

The asset you are rebuilding is the meaning of red, so measure success by behaviour — does a failure get read? — rather than by the pipeline turning green. Four moves. **Set the standard**: state explicitly what a red run obliges people to do, and make that obligation achievable by not asking them to triage a hundred meaningless failures. **Stop the inflow**: give authors the seams to control time, randomness, ordering and external systems, and exercise new cases repeatedly before they can block anyone; cleanup without inflow control is a treadmill. **Preserve a believed signal**: rather than unblocking everything, keep a small, reliable subset as the gate so there is always one verdict people act on. **Fund the work**: named owners and real capacity, not goodwill. What you refuse matters too: blanket automatic re-runs and a wholesale non-blocking suite both produce green pipelines and no trust.

go deeper

for a junior

Notice the behaviour when it appears around you: if the habit is to re-run first and read second, that is the symptom. Recognising it, and not adopting it yourself, is the entry-level part of this.

for a middle

Be able to explain the mechanism — a red result that is often meaningless makes every red result cheap — and describe what you personally do with the unstable case in front of you rather than passing it along.

for a senior

Argue for stopping the inflow as well as draining the backlog, and for keeping at least one gate whose failures the team believes rather than unblocking the pipeline wholesale. Be specific about the seams that make new cases controllable.

for a principal

Own the tradeoff and its funding: what standard a red run must meet, what capacity the cleanup gets, what you refuse to do to reach green, and how you evidence the cost to leadership using incidents you actually have.

## Name the failure correctly The visible symptom is a team that re-runs first and reads second. The underlying failure is that the pipeline's verdict has stopped carrying information, and people have adapted rationally to that. This is normalisation: the deviation becomes the working method, and it is not fixed by asking people to be more diligent about failures that are usually meaningless. Diligence is not the scarce resource — signal is. The cost lands in three places. Engineering hours go into re-running and re-reading. Release cadence slows while runs are repeated. And, most expensively, genuine regressions get discounted along with the noise. Consider a 340-case nightly regression pack for a utility billing run, taking about 47 minutes, where roughly nine cases were known to fail sometimes and always passed on re-run. The team's habit became: re-run the pack, then look. A currency-rounding drift of a hundredth of a unit surfaced in that pack and was absorbed by the habit; it reached invoices weeks later. The suite had done its job on the first night. The trust deficit is what made that irrelevant. ## Move 1 — state what red must mean Write down the obligation: a red result on the blocking gate stops the line and someone owns it before anything else proceeds. That standard is only credible if it is achievable, which is why it is stated over a *scope you can defend* rather than over everything at once. ## Move 2 — preserve one signal people believe The instinct under pressure is to make the whole suite non-blocking until it is clean. That reliably fails: an advisory suite is one nobody reads, its failures become nobody's work, and the instability rots in place while genuine regressions accumulate unnoticed. The better move is a carve-out — identify a subset you can stand behind, make that the blocking gate, and hold it to the standard absolutely. A small trusted gate beats a large ignored one, and it gives the cleanup a target to grow rather than a hole to climb out of. Moving unstable cases out of the blocking path is a legitimate tool here, but on its own it only relocates the problem; the trust comes from what remains blocking, not from what was moved. ## Move 3 — stop the inflow before draining the backlog Cleanup without an inflow control is a treadmill, and the team will notice within a quarter. Inflow control is mostly about what authoring a case costs by default: * seams so any case can control time, randomness, identifiers and ordering without heroics; * per-case isolation and cleanup as the default the harness provides, not a discipline each author reinvents; * an expectation at review that a new case names what it waits on and what it controls; * exercising new cases repeatedly before they are allowed into the blocking path, so an unstable one is caught while it is still one team's problem. ## Move 4 — fund it as work, not goodwill Unstable cases are defects, and defects that live in a backlog labelled "when someone has time" do not get fixed. Give the effort named owners, a place in the plan, and enough capacity that the backlog visibly shrinks — visible progress is itself part of restoring belief. Expect to make a call about cases whose cost exceeds their value; the decision to remove coverage is a real one and should be made deliberately, with an understanding of what was covered, rather than by whoever is annoyed at midnight. ## What to refuse * **Blanket automatic re-runs** to make the pipeline report green. They convert a visible problem into an invisible one and let genuinely intermittent product defects pass the gate. * **Turning everything non-blocking** with no end date and no replacement signal. * **Exhortation** — asking people to read failures more carefully while the base rate of meaningless failures stays where it was. * **Declaring victory on green.** Green with re-runs is not trust. ## How you know it worked Judge by behaviour, not by colour. Does a red result on the gate get investigated rather than re-run? Do engineers cite a failure in review as evidence? Does the team stop asking whether the failure is "real"? Those are the observable signs that red carries information again. Pair them with the incidents you no longer have: escaped defects that the suite had actually caught are the strongest argument you will ever make for the investment, and the one that unlocks capacity from people outside the team. ## What interviewers listen for They want the tradeoff owned out loud: what you block, what you tolerate, what it costs, who pays, and what you refuse to do to reach green. An answer that reaches immediately for automatic re-runs or for making the suite advisory has optimised the colour and abandoned the signal.

  • Is making the suite non-blocking while you clean it up a reasonable step?
    Only with an end date and a replacement signal. An advisory suite is one nobody reads, its failures stop being anyone's work, and the instability rots in place while real regressions hide in it. If you must unblock the pipeline, carve out a subset you believe in and keep that blocking, so there is always one verdict the team acts on.
  • How do you stop new unstable cases arriving while you fix the existing ones?
    Make nondeterminism expensive to author and easy to avoid: seams for controlling time, randomness and ordering; per-case isolation as a harness default; a review expectation that a new case names what it waits on; and repeated execution of new cases before they enter the blocking path. Without inflow control the cleanup is a treadmill and the team learns the effort was theatre.
  • How do you make the cost of a distrusted suite legible to leadership?
    Frame it in outcomes they already own: hours spent re-running and re-reading, releases delayed while runs are repeated, and defects that shipped because a real failure was discounted as noise. Attach the incidents you actually have. A cost argument built on specific escaped defects moves capacity where an abstract appeal to quality does not.

saying these in an interview costs you the question

  • Makes the whole suite advisory and calls it fixed
  • Adds blanket automatic re-runs so the pipeline reports green
  • Asks the team to read failures more carefully instead
  • Treats the cleanup as unfunded background work
  • Removes coverage without knowing what it covered
  • Declares success because the pipeline is now green

context