When a pipeline runs your suite, which failures should stop the stage and which should only be reported beside a passing one?
answer
- Blocking is a risk claim
- Decide once, not per run
- Default blocking, exceptions declared
- Every exception needs an owner and expiry
- An unread report is silence
basics
~20 sFailures that say the change is unsafe block the stage. Failures already triaged and accepted, and checks advisory by design, are reported without blocking. Put the split in a written classification the suite applies, not in a per-run argument.
solid answer
~50 sBlocking is a claim about **risk**, not about defect severity. A failure blocks when advancing the change past this stage would carry unacceptable risk - a genuine regression in behaviour the stage is responsible for. It is only reported when the information is useful but the team has already decided not to act now: a new check running advisory while it beds in, or a known open issue with an owner and a date. What matters most is *where the decision lives*. Argued in the moment - “this one looks unrelated, let it through” - the gate erodes one exception at a time. Express it instead as a classification the suite carries: each check declares whether it blocks, the entry point maps only the blocking set to a failing status, and the rest is reported. Changing a classification then becomes a reviewed change with a name on it.
code
pseudocode · 15 linescheck "checkout completes with a stored payment method":
blocking: true
check "page weight under budget":
blocking: false # advisory while the budget is calibrated
owner: "web-platform"
review_by: "2026-11-01"
entry_point:
failures = run(selection)
blocking_failures = filter(failures, f -> f.check.blocking)
write_results_document(failures) # all of them, with classification
report_against_change(failures) # visible where review happens
return EXIT_FAILURES if blocking_failures else EXIT_OKgo deeper
Know that not every red signal in a pipeline stops the work: some failures stop the stage and others are reported while the stage still passes. Be ready to say which of your team's checks are which.
Explain the mechanics: the entry point maps only the blocking set to a failing status while every failure still reaches the run's record. Expect to describe how a check is declared blocking or advisory in the first place.
Show judgement about risk rather than defect severity, and be ready to diagnose a gate people override. Explain how you would move a check between blocking and advisory without the decision being made ad hoc while a change is waiting.
Own the erosion problem. Be ready to argue for blocking by default with declared, dated exceptions, to publish the size of the advisory set as a health number, and to say what a gate people no longer trust costs an organisation.
Every failure a run produces answers two separate questions, and conflating them is where teams go wrong. The first is *what happened* - a case did not pass, and the record says so with its evidence. The second is *what the pipeline should do about it* - stop the change here, or let it continue while telling someone. The first is a fact. The second is a **policy decision**, and it belongs to the entry contract between the suite and the pipeline that invokes it. ## Blocking is a claim about risk A failure should block when advancing the change past this stage carries unacceptable risk: a real regression in behaviour the stage exists to protect, or the discovery that the run itself cannot be trusted. Note what is *not* on that list - how severe the underlying defect is. A cosmetic problem can be a blocking failure if the stage's job is to protect that surface; a serious but long-known issue with an owner and a date can be non-blocking, because everyone has already decided when it will be fixed. **Severity describes the defect. Blocking describes the decision to advance.** A failure earns a report rather than a block when the information is worth having and the team has already decided not to act on it now: - **A check still being calibrated.** New checks routinely fire on things that turn out to be fine; running advisory for a few weeks buys the data to set the threshold honestly. - **A known, owned, dated failure.** The team has looked, agreed and scheduled it. Re-litigating that decision on every run is waste. - **A measurement with no agreed threshold.** Reporting a number is useful; stopping work at a line nobody agreed to draw is not. - **A signal this stage is not responsible for.** If the stage protects one boundary, a failure from elsewhere is information, not a gate. | Signal from a run | Effect on the stage | | --- | --- | | Behaviour this stage protects regressed | Blocks | | The suite could not run, or ran nothing | Blocks | | A check declared advisory while being calibrated | Reported only | | A failure with an open issue, an owner and a date | Reported only | | A measurement with no agreed threshold | Reported only | | An intermittent case awaiting diagnosis | Governed by the team's separate quarantine policy | ## Express the decision, do not argue it The mechanism matters more than the taxonomy. If the split is decided in the moment - *"this one looks unrelated, let it through"* - three things follow. Every exception looks reasonable on its own. Nobody can see the set of exceptions, because there is no set. And the gate drains one change at a time until the stage is decorative. Expressing it instead means: 1. **The classification lives with the check**, in the repository, declared beside what it checks, so it moves with the code, is reviewed with the code, and cannot drift from what it governs. 2. **Blocking is the default.** Advisory is the exception, and an exception has to be written down. 3. **Every advisory declaration carries an owner and a review date.** Without a date, "for now" is forever. 4. **The entry point maps only the blocking set to a failing exit status.** Everything else is reported. 5. **Reported failures still reach the run's record in full detail.** The classification changes the stage's verdict, never the evidence. ## Both extremes fail *Everything blocks* looks rigorous and is not. When a stage stops work over things that carry no real risk, people learn to route around it: overrides become routine, runs are re-invoked until they go green, and the checks that genuinely matter get the same shrug as the ones that never did. Trust is the resource being spent, and it is spent uniformly across the whole gate rather than on the check that deserved it. *Nothing blocks* is the same failure as a run that always exits successfully. Reports fill up, dashboards look busy, and no change is ever stopped. The only difference is that this version was chosen deliberately, which makes it harder to notice. The health signal between the two extremes is the **size and age of the advisory set**. A set that shrinks as checks are repaired or promoted is a working policy. A set that only grows is a gate being drained politely. ## Reading the overrides If you inherit a pipeline and want to know whether its classification is honest, look at where people override it. Overrides are not spread evenly; they cluster on a handful of checks that block without earning it - slow, unreliable, or measuring something the team never agreed to. Those clusters are the classification telling you it is wrong. The repair is one of two things: fix the check so the block is credible, or reclassify it openly with an owner and a date. What you must not do is leave it blocking and let everyone learn that blocking is optional, because that lesson generalises to every other check in the stage.
- A team keeps advancing changes past a red stage with an override. What does that tell you?That the classification is wrong, not that the people are careless. Overrides cluster on checks that block without earning it - unreliable, slow, or measuring something the team never agreed to. Find the checks they concentrate on, then either repair the check so the block is credible or reclassify it openly. An override used weekly is a classification decision being made badly and unrecorded.
- How do you stop the advisory set from growing forever?Give every advisory declaration an owner and a review date, and publish the set's size as a visible number. Review it on a fixed cadence: each entry either becomes blocking, is fixed, or is deleted. Without a date, “advisory for now” becomes permanent, and the stage quietly stops covering what people believe it covers.
- Should a failure that does not block still appear in the run's record?Yes. Its classification changes the stage's verdict, not the evidence, so it is recorded with the same detail as a blocking failure and stays visible to triage and trend analysis. Dropping it from the record is how an advisory check becomes an unmonitored one, and how a real regression hides behind an old exemption.
Deciding at three in the morning whether an alarm should stop the engines is how accidents happen; that wiring is decided once, in daylight, by people who are not under pressure.
saying these in an interview costs you the question
- Decides blocking case by case while the run is red
- Treats every failure as blocking, then routes around the gate
- Marks a check advisory with no owner and no expiry
- Drops non-blocking failures from the run's record entirely
- Confuses defect severity with whether the change is safe to advance