skip to content

In Allure 3's `allurerc` configuration, how is a `qualityGate` assembled from rules, and what do the built-in `maxFailures`, `minTestsCount` and `successRate` rules each measure?

level: middleimportance: should knowfreq 50%

answer

  1. two keys under one config section
  2. groups of id-to-threshold pairs
  3. leave use out for the built-ins
  4. successRate is a fraction, not a percent

basics

~20 s

The qualityGate section holds rules: a list of groups whose keys are rule ids and whose values are the expected thresholds. maxFailures caps failed results, minTestsCount floors how many ran, and successRate floors passed divided by total.

solid answer

~40 s

`qualityGate` takes two keys. `rules` is a list of rule groups; inside a group each key is the id of a rule and its value is the threshold that rule expects. `use` lists the rule implementations available to those ids, and if you leave it out the built-ins apply, so most configurations write only `rules`. The built-in ids are `maxFailures`, `minTestsCount`, `successRate`, `maxDuration`, `allTestsContainEnv`, `environmentsTested`, `metricMax`, `metricMin`, `metricMaxDelta` and `metricMaxDeltaPercent`. `maxFailures` counts unsuccessful results and requires the count to stay at or below the threshold; `minTestsCount` counts validated results and requires at least that many; `successRate` is passed divided by total, a fraction rather than a percentage. A key that no available rule answers to raises an error rather than being skipped.

code

json · 11 lines
json
{
  "name": "Nightly regression",
  "qualityGate": {
    "rules": [
      {
        "maxFailures": 0,
        "minTestsCount": 1
      }
    ]
  }
}

go deeper

for a junior

Know where the configuration lives, an allurerc file in the working directory, and that each entry inside a rule group pairs a rule name with the threshold it expects.

for a middle

Be able to walk through the two keys and say what each built-in rule counts, including that successRate is passed over total as a fraction rather than a percentage.

for a senior

Point out the rules that catch what a failure ceiling misses: a floor under how many tests ran, and the delta rules that pass by default when no previous value exists.

for a principal

Own the question of which thresholds are worth blocking on at all, and how you stop a per-job command-line override from quietly replacing the gate the repository declared.

## The shape of the configuration Allure 3 reads its configuration from an `allurerc` file in the working directory -- `allurerc.js`, `allurerc.mjs`, `allurerc.ts`, `allurerc.json`, `allurerc.yaml` and a few more spellings are tried in turn, and `--config`/`-c` overrides the search. The gate lives under one top-level key, `qualityGate`, and that key has exactly two parts. **`rules`** is a list of *rule groups*. A group is a plain object whose keys are rule ids and whose values are the thresholds those rules expect. Two groups are two independent sets of conditions, evaluated in order. The reason to write more than one is that a group can carry behaviour as well as thresholds -- most usefully, that a breach in this group should stop the command immediately while a breach in another should only be reported at the end. **`use`** is the list of rule implementations available to the ids in `rules`. It exists so that rules written outside Allure can be plugged in. If you do not write `use`, the built-in rules are used, which is why almost every real configuration contains only `rules`. The binding between the two is by name and it is strict: for each key in a group the gate looks for a rule of that id among the available implementations, and **if it finds none it raises an error**. A misspelled `maxFailure` does not quietly do nothing; it stops the command. That is the right behaviour for a gate -- a silently inert rule is the failure mode you least want. ## What each built-in rule measures | rule id | what it reads | passes when | |---|---|---| | `maxFailures` | count of unsuccessful results | the count is at or below the threshold | | `minTestsCount` | number of validated results | the count is at or above the threshold | | `successRate` | passed divided by total | the fraction is at or above the threshold | | `maxDuration` | the longest single result's duration | that duration is within the limit | | `allTestsContainEnv` | results whose environment is not the named one | there are none | | `environmentsTested` | the set of environments seen | every named environment appears | | `metricMax` / `metricMin` | the average of a named metric | it is within the given bound | | `metricMaxDelta` / `metricMaxDeltaPercent` | that average against its most recent previous value | the change is within the given bound | Three of these repay a closer look. - **`successRate` is a fraction, not a percentage.** It is computed as passed over total and compared directly against what you wrote, so the threshold belongs on a zero-to-one scale. Writing it as though it were a percentage produces a rule that can never pass. - **`minTestsCount` is the rule people forget, and it catches the worst failure.** A run that collected nothing has no failures, so every failure-counting rule is satisfied and the gate goes green over an empty report. A floor under how many tests were actually validated is the only rule in the set that notices. - **`maxDuration` bounds the *slowest single test*, not the run.** It takes the maximum duration across the validated results and compares that one value against the limit. It is not a wall-clock budget for the suite. The two delta rules are the only built-ins that look outside the current run: each takes the average of a named metric now, finds that metric's most recent value in the stored history, and compares the change against a threshold -- as an absolute difference or as a percentage. When there is no previous value to compare against they **pass** rather than fail, which is sensible on a first run but does mean a lost baseline reads as a clean bill of health. ## Rules from the command line The three most common rules can be set without a configuration file at all. `--max-failures`, `--min-tests-count` and `--success-rate` each set the matching built-in, and when any of them is present they are assembled into a single group that **takes priority over whatever the configuration said**. `--fast-fail` attaches the stop-at-first-breach behaviour to that group. This is convenient for a one-off pipeline step and it is a trap for a repository that also has a configuration: the command line *replaces* the gate rather than adding to it, so a job that passes `--max-failures` is no longer running the environment and metric rules the repository declared. ## What never reaches a rule Before any rule is evaluated, two categories of result are dropped: attempts superseded by a later attempt of the same test, and failures already resolved as muted or accepted through the known-issues mechanism. Every count in the table above is therefore a count of *surviving, unresolved* results -- which is usually what you want, and always worth saying out loud when someone asks why the gate's number does not match the number of red rows they can see on the page.

  • Your configuration declares environment and metric rules, and the pipeline step passes `--max-failures`. Which rules actually run?
    Only `maxFailures`. When any of the command-line rule options is present, Allure assembles them into a single rule group and that group replaces the configured gate rather than extending it. The environment and metric rules declared in `allurerc` are not evaluated for that invocation.
  • Why would you split a gate across two rule groups instead of putting every rule in one?
    Because a group carries behaviour as well as thresholds. One group can stop the command at its first breach while another collects every breach and reports them together, and a group can be narrowed to a subset of results so a strict threshold applies only to the tests that warrant it.

saying these in an interview costs you the question

  • Writing successRate as a percentage rather than a fraction
  • Assuming an unknown rule id is skipped silently
  • Reading maxDuration as a budget for the whole run
  • Trusting a failure ceiling on a run that collected nothing