skip to content

Report-Level Gates

A threshold over the whole aggregated report instead of the suite's exit code: new failures against the last build, a flaky ceiling, unclassified failures, and where the verdict is computed.

on this pageshow

explore

questions

4

A CI job consolidates several suites into one Allure report and then has to decide whether the build passes. What is a report-level quality gate, and where in Allure is that verdict computed?

level: juniorimportance: must knowfreq 58%

answer

  1. a verdict taken after consolidation
  2. not the runner's exit status
  3. one Allure major ships a command
  4. quality-gate prints breaches and exits

basics

~20 s

A report-level quality gate is a threshold checked against the whole aggregated report, not against one runner's outcome. Allure computes it in the 3.x line only, through its quality-gate command, which exits non-zero when a rule is breached.

solid answer

~40 s

A report-level gate is a second verdict, taken after results have been consolidated. The runner's own exit status is per-invocation, so a sharded suite produces several of them and none sees the others. The gate reads the assembled report and evaluates thresholds over all of it: a ceiling on total failures, a floor under how many tests actually ran, a minimum success rate, a requirement that every intended environment produced results. In Allure this ships in the 3.x line as the `quality-gate` command. It reads the configured `qualityGate` rules, validates the results, prints every breached rule with its actual and expected value, and exits non-zero. Allure 2 has no such command. The breached rules are also carried into the generated report, so the verdict is readable without digging out the CI log.

go deeper

for a junior

Be able to say plainly that this verdict is computed over the finished, consolidated report rather than by the process that ran the tests, and name the Allure 3 command that computes it.

for a middle

Explain what a gate can express that a single invocation cannot: totals across shards, a floor under how many tests ran, and comparisons against a previous run.

for a senior

Show that you know which results the gate quietly excludes before any rule runs, and that a gate with no rules configured must refuse rather than pass.

for a principal

Be ready to argue where this decision belongs at all, in the report tool or elsewhere, and what each placement costs in ownership and in feedback time.

## Two different verdicts on the same run A pipeline can take two separable decisions about a test run. The **first** belongs to whatever executed the tests. A process ends, and its exit status reports whether that process saw a failure. It is per-invocation and blind past its own boundary: shard a suite across six workers and you get six such verdicts, none of which can see the other five, and none of which knows anything about the run before this one. The **second** is taken *after* the results have been collected into one place. Every shard's output is read into a single store, superseded retry attempts are folded away, results you have already agreed to ignore are set aside, and only then is a threshold evaluated. That second decision is what "report-level quality gate" names. Because it runs over the assembled report, it can express conditions the first structurally cannot: - a ceiling on the **total** number of failures across the whole report, not per shard; - a floor under how many tests actually ran, which catches the empty or half-collected run that would otherwise pass because nothing failed; - a minimum success rate over everything; - a requirement that each environment you meant to cover produced at least one result; - a comparison of a measured value against the same value in a previous run. | | the runner's own exit status | the report-level gate | |---|---|---| | scope | one invocation | the whole consolidated report | | sees other shards | no | yes | | sees the previous run | no | yes, where a rule asks for it | | expressible condition | "something failed" | a threshold over the aggregate | ## Where Allure computes it **Quality Gate is an Allure 3 feature.** The 3.x CLI carries a `quality-gate` command. Allure 2 has no equivalent: its whole command set is `generate`, `serve`, `open` and `plugin`. This matters more than it sounds, because "quality gate" is a phrase that gets attached confidently to results servers that do not implement one. Do not assume a tool has the feature merely because it stores and displays test results; if you need a gate in front of a store that has none, you compute it yourself from what that store's API returns. The command's shape is straightforward: 1. It reads the Allure configuration from the working directory -- an `allurerc` file -- and takes the `qualityGate` section from it. 2. If neither the configuration nor the command line supplies any rule, it refuses to run rather than passing vacuously. A gate that silently does nothing is worse than no gate. 3. It resolves the results directories it was pointed at, defaulting to the usual `allure-results` layout. 4. It validates the results against each rule group. 5. It prints every breached rule with the value it found and the value it expected, then exits non-zero. With nothing breached it exits zero. The command line can also carry a gate on its own, without a configuration file: `--max-failures`, `--min-tests-count` and `--success-rate` each set the matching built-in rule, and when any of them is present it takes priority over what the configuration said. `--fast-fail` makes the command stop at the first breach instead of collecting them all. ## What the gate is allowed to ignore Two categories of result are removed before any rule sees them, and both are worth knowing because they change the numbers a rule reports: - **Superseded retry attempts.** When several attempts of the same test are present, only the surviving one is validated. The earlier ones do not count toward a failure ceiling. - **Failures you have already resolved.** Allure 3 reads a known-issues file, pointed at with `--known-issues` or through the `resolutions` configuration, and each rule there resolves a matching failure into a category. Results resolved as `muted` or `accepted` are excluded from validation; results resolved as `issue` are **not** -- they are still failures, they simply carry a link to the ticket. ## Where the verdict ends up The exit code is the part CI consumes, but it is not the whole output. The breached rules are also carried into the generated report itself, so someone who opens the report sees which rule failed, what it expected, what it actually found and which tests were implicated -- rather than having to reconstruct that from a console log that may already have scrolled away. That is the practical argument for computing the verdict here at all: the decision and the evidence for it end up in the same artefact, and the artefact is the thing a human opens tomorrow morning.

  • If the gate is only computed after the report is assembled, what have you already paid for by the time it says no?
    The full run and the collection step. A report-level gate is deliberately late: it trades early feedback for a decision made over complete information. Keep the cheap, early checks where they are, use the gate for conditions that only make sense over the aggregate, and reach for the early-exit option only when you genuinely cannot afford to wait.
  • The gate refuses to run and reports that it is not configured. What are the two places it looked?
    The `qualityGate` section of the `allurerc` configuration in the working directory, and the command line's own rule options: `--max-failures`, `--min-tests-count` and `--success-rate`. If neither supplies a rule it stops rather than exiting zero, so a typo in the configuration surfaces as a refusal instead of a silent pass.

saying these in an interview costs you the question

  • Assuming every results server ships a named Quality Gate feature
  • Believing Allure 2's CLI has a quality-gate command
  • Treating the gate as a second copy of the runner's exit code
  • Configuring no rules and assuming the gate still guards something
open as a page

In Allure 3's `allurerc` configuration, how is a `qualityGate` assembled from rules, and what do the built-in `maxFailures`, `minTestsCount` and `successRate` rules each measure?

level: middleimportance: should knowfreq 50%

basics

~20 s

The qualityGate section holds rules: a list of groups whose keys are rule ids and whose values are the expected thresholds. maxFailures caps failed results, minTestsCount floors how many ran, and successRate floors passed divided by total.

open as a page

Allure 3's `quality-gate` command validates rules as results are read, and `--fast-fail` makes it exit at the first breach. Why can a suite that retries its failed tests turn that gate red on a run whose finished report is green?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the gate validates in batches while results are still being read. A failed first attempt is only recognised as superseded once its later, passing attempt has been read, and with fast-fail the command has already exited by then.

open as a page

An Allure 3 report marks a test as newly regressed against the previous run. What comparison produces that label, and what does it do when the previous run skipped the test?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

Allure 3 attaches a status transition, one of new, fixed, regressed or malfunctioned, by comparing this run's status against the most recent history entry carrying a real verdict. Skipped and unknown entries are stepped over, not used as the baseline.

open as a page