skip to content

An Allure 3 report marks a test as newly regressed against the previous run. What comparison produces that label, and what does it do when the previous run skipped the test?

level: middleimportance: nice to knowfreq 38%

answer

  1. four names, not a boolean
  2. no label when nothing changed
  3. the baseline steps over non-verdicts
  4. failed and broken get different names

basics

~20 s

Allure 3 attaches a status transition, one of new, fixed, regressed or malfunctioned, by comparing this run's status against the most recent history entry carrying a real verdict. Skipped and unknown entries are stepped over, not used as the baseline.

solid answer

~40 s

A result carries a transition drawn from four values: `new`, `fixed`, `regressed` and `malfunctioned`. `new` means the test has no history at all. Otherwise the current status is compared against the most recent *significant* status in that test's history, and entries whose status is `skipped` or `unknown` are stepped over, so the comparison reaches back to the last run that actually produced a verdict. If the status changed, a move to `passed` is `fixed`, to `failed` is `regressed` and to `broken` is `malfunctioned`. Allure keeps `failed` and `broken` separate, so a test that started erroring rather than asserting gets its own label. If the status did not change there is no transition at all: a test that failed last time and failed again is a standing failure, not a regression.

go deeper

for a junior

Recall the four transition names and that a result whose status did not change carries none of them at all.

for a middle

Explain how the baseline is picked: the most recent history entry carrying a real verdict, stepping over skipped and unknown ones.

for a senior

Point out that an empty history makes every test new, so a regression-only rule fails open on exactly the run where you would most want it to bite.

for a principal

Decide whether a regression-only verdict is a policy worth having, given that it never blocks on a failure the team has already grown used to.

## Why "new failure" needs a definition "Gate the build on new failures only" is one of the most-requested rules in test reporting, and it hides two decisions that have to be made before it can mean anything: *new relative to what*, and *what counts as a change*. Allure 3 answers both, and the answers are worth knowing before you build a verdict on top of them. Each test result in a 3.x report can carry a **status transition**, taken from exactly four values: | transition | what it means | |---|---| | `new` | this test has no history at all in the report | | `fixed` | the status changed and this run passed | | `regressed` | the status changed and this run failed | | `malfunctioned` | the status changed and this run broke | There is no fifth value for "unchanged". A result whose status matches its baseline simply carries no transition -- which is the correct shape, because a test that failed last build and failed again is a *standing* failure, not a new one, and a rule that treats the two alike would block on the same problem every day. ## How the baseline is chosen The comparison is not against "the previous run" in the naive sense of the immediately preceding entry. The procedure is: 1. If the test's history is empty, the transition is `new` and nothing else is consulted. 2. Otherwise, walk the history from most recent and take the first entry whose status is *significant*. `skipped` and `unknown` are not significant and are stepped over. 3. Compare the current status against that one. If they are equal there is no transition. If they differ, the transition is named for the current status: `passed` gives `fixed`, `failed` gives `regressed`, `broken` gives `malfunctioned`. Step 2 is the part that surprises people, and it is the right behaviour. A test that was skipped last night because its tag was excluded from the nightly run has told you nothing about the product, and treating that skip as the baseline would make the next real failure look like a fresh regression and the next real pass look like a fix. It also means the label can be comparing against something considerably older than "yesterday", and the report does not spell out how far back it reached. If you are using these labels to decide something, that is a fact to hold in mind. ## `failed` and `broken` are not the same transition Allure separates a test that failed an assertion from one that broke before it could assert, and the transition vocabulary preserves the distinction: `regressed` for the first, `malfunctioned` for the second. That is genuinely useful triage information -- a suite that suddenly reports a wave of `malfunctioned` results is usually telling you about the environment or a fixture rather than about the product -- but it means a rule phrased as "block on new failures" has to say which of the two it means, or it will silently ignore half of them. ## What this does and does not give you for gating Two limits are worth stating plainly. - **The built-in gate rules do not include a transition rule.** The rules Allure 3 ships count failures, count tests, compute a success rate, bound a duration, check environments, and compare named metrics. None of them is "fail on a regression". A gate that acts on transitions is something you narrow a rule group to, or a rule you supply yourself. - **The rules that do look outside the current run are the metric-delta ones**, and they compare a named metric's average now against its most recent stored value. That is a different mechanism from the per-test transition, and one that passes rather than fails when no previous value is found. The second limit generalises. Every one of these labels is a statement about history, and history is something the report has to have been given. A run whose history is empty produces `new` for every test -- technically correct and completely uninformative, and easy to misread either as "the whole suite regressed" or as "nothing regressed". Before you gate on new failures, confirm the report actually has a baseline to compare against, because the failure mode is silent in both directions: a lost baseline makes everything new, and a stale one makes a long-standing failure keep looking fresh. ## A short checklist - Decide whether your rule means `regressed`, `malfunctioned`, or both. - Remember that unchanged failures carry no transition, so "new failures only" will not block on them at all. - Do not assume the baseline is the immediately previous run; skipped and unknown entries are stepped over. - Verify a baseline exists before trusting any of it, and keep a plain failure ceiling in the gate for the run where it does not.

  • A report with no history at all is fed to a rule that blocks on regressions. What does it do?
    Nothing blocks. With no history every test's transition is `new`, and `new` is not `regressed`, so a regression rule finds none, including for tests failing right now. The empty-baseline case fails open, which is why a plain failure ceiling and a floor under how many tests ran still belong in the gate.
  • Why does Allure name a change to `broken` differently from a change to `failed`?
    Because they usually have different causes. A newly `failed` test asserted and got the wrong answer; a newly `malfunctioned` one never got far enough to assert. A wave of the second is normally an environment or fixture problem, and a separate name lets you see that in the report rather than infer it.

saying these in an interview costs you the question

  • Assuming the baseline is always the immediately previous run
  • Expecting a transition value on an unchanged result
  • Treating broken and failed as one regression category
  • Gating on new failures without checking a baseline exists