skip to content

History and Trends

What lets one execution's numbers be compared with the last: the key a result series is filed under, how earlier runs reach a fresh report, and how repeated attempts are shown.

on this pageshow

explore

questions

17

On a test result in an Allure 2 report, what do the `retriesCount` and `retriesStatusChange` fields say, and what does `retriesCount` deliberately not include?

level: juniorimportance: must knowfreq 50%

answer

  1. a count and a boolean
  2. counts the other attempts only
  3. off by one from executions
  4. the flag just compares statuses

basics

~20 s

retriesCount is how many earlier attempts were folded into this result, so it excludes the attempt on screen; a test executed twice reads one. retriesStatusChange is true when any earlier attempt ended in a different status.

solid answer

~40 s

Both fields are written by Allure 2's `RetryPlugin` onto the one attempt it leaves visible. The plugin builds a list of every attempt in the group **except** the survivor, then calls `setRetriesCount(...)` with that list's size — so the field is always one less than the number of attempts the runner made, and a test executed three times reads two. `setRetriesStatusChange(...)` takes the statuses of those same earlier attempts, discards any equal to the survivor's, and sets the flag to true if anything is left. It is a plain boolean: it says the outcome moved across the attempts, not which way it moved and not which attempt moved it. A nonzero count with a false flag therefore means the test was repeated and came out the same way every time.

go deeper

for a junior

Recall that retriesCount counts the earlier attempts only, so a test executed twice reads one, and that retriesStatusChange is a plain true or false rather than a status of its own.

for a middle

Explain where both values come from: one list of every attempt except the survivor, sized for the count, and filtered against the survivor's status for the flag. Say why a count of zero can never accompany a true flag.

for a senior

Be able to say what a reader may and may not conclude from the pair alone - that a nonzero count with a false flag means the outcome never moved, and that neither field says which way it moved or which attempt moved it.

for a principal

Decide what your teams are allowed to read off these fields. A count of repeated attempts is a workload signal, it names no tests, and any summary built on it needs the underlying results reachable beside it.

## Where the two fields come from When a runner repeats a test, Allure 2's generator ends up with several results for it. `RetryPlugin` groups them, keeps the attempt with the newest start time as the visible one, and hides the rest. Two fields are then set on that survivor to describe what was folded into it: `retriesCount` and `retriesStatusChange`. Neither is written by the test runner and neither exists in the results a runner produces — they are **derived at report-generation time** and appear on the reader-side result model. The plugin computes them from one list. It takes the results in the group, orders them, drops the survivor, and maps what is left into the entries of the `retries` block. Everything else follows from that list: - `setRetriesCount(...)` receives the list's **size**. - `setRetriesStatusChange(...)` receives whether the statuses in that list contain anything the survivor's status is not. ## The off-by-one, and why it is not a bug The single most common misreading is treating `retriesCount` as the number of times the test ran. It is not. Because the list excludes the survivor, the field is the number of *repeats*, not the number of *executions*. | the runner executed the test | `retriesCount` | |---|---| | once | 0 | | twice | 1 | | three times | 2 | The naming is consistent once you read it as "how many retries were needed" rather than "how many attempts exist". It also means the field is a workload figure — the amount of repeated execution the run absorbed for this test — and it says nothing on its own about whether the repetition helped. A count of two accompanied by a visible failure means the test was tried three times and failed at the end. ## What the boolean does and does not claim `retriesStatusChange` is computed by collecting the statuses of the earlier attempts, filtering out every status equal to the survivor's, and setting the field to true when the remaining set is non-empty. Three consequences follow directly, and all three are worth being able to state: 1. **It is symmetric.** A test that failed and then passed sets it. A test that passed and was then re-executed into a failure sets it just the same. The flag records that the outcome moved, not the direction of travel. 2. **It is not an eventually-passed marker.** A test whose earlier attempts were all failures and whose survivor is also a failure leaves the flag false, because nothing differed. Reading a false flag as "this test is fine" gets it exactly backwards in the worst case: a test that failed on every attempt has a nonzero count and a false flag. 3. **It does not name anything.** The boolean identifies neither which attempt differed nor what status it carried. That detail lives in the `retries` block beside it, where each entry carries its own status and status message. ## Reading the pair together The two fields are only useful as a pair, and the four combinations divide cleanly: - **count 0, flag false** — the test executed once. Nothing was folded in. - **count > 0, flag false** — the test was repeated and every attempt landed on the same status as the one shown. Repeated work, no disagreement. - **count > 0, flag true** — the test was repeated and at least one attempt disagreed with the visible outcome. This is the combination worth a reader's attention, in both directions. - **count 0, flag true** — cannot occur. With no earlier attempts there are no statuses to compare, so the filtered set is necessarily empty. ## Where they live, and where they do not Both fields sit on the **surviving** result only. The hidden attempts are ordinary results carrying their own status and their own timings; they are not annotated with counts of their own, and reading one of them directly tells you nothing about how many siblings it had. The fields also travel into the report's serialised data for the visible result and into the leaf entries the trees are built from, which is how a tree row can show a retry marker without opening the test's page. One more boundary is worth stating plainly. Neither field says anything about *why* the repetition happened, and neither is a judgement. `retriesCount` counts what the runner did; `retriesStatusChange` compares statuses. Any interpretation beyond that — whether the repetition was policy or accident, whether a moved outcome should be treated as noise or as a defect — is a conversation the report supports rather than one it settles.

  • A test ran three times in one execution. What does `retriesCount` read on the row you see?
    Two. The plugin builds its list from every attempt except the survivor and then sets the count to that list's size, so the field is always one less than the number of executions the runner actually performed.
  • Can `retriesStatusChange` be true when the visible attempt failed?
    Yes. The flag only asks whether an earlier attempt carried a status the survivor does not. A test that passed first and failed on a later attempt sets it exactly as a test that failed first and passed later does — the flag reports movement, never direction.

saying these in an interview costs you the question

  • Reads retriesCount as the total number of executions
  • Thinks retriesStatusChange means the test eventually passed
  • Assumes a nonzero count implies an earlier failure
  • Expects both fields on every attempt, not just the survivor
open as a page

An Allure 2 report is regenerated on every CI build, yet its trend chart always shows a single point and each test case page shows only the current run. What is missing, and where does Allure 2 get trend data from?

level: juniorimportance: must knowfreq 58%

basics

~20 s

The previous report's history directory is missing. Allure 2 keeps no database: it builds trends only from a history folder already present in the results directory it reads, so you copy that folder in from the last report before generating.

open as a page

In Allure's result model, `io.qameta.allure.model.Status` has exactly four values and none of them is flaky, so where does a result's flaky mark actually live and what writes it?

level: juniorimportance: must knowfreq 68%

basics

~10 s

Flakiness is a boolean on StatusDetails, set with setFlaky(...), not a status value. The result keeps one of Status's four values - FAILED, BROKEN, PASSED or SKIPPED - and the flag rides alongside it.

open as a page

In an Allure `-result.json` file, a test result carries both a `historyId` and a `testCaseId`. What does each of those keys identify, and what goes wrong if a tool treats them as interchangeable?

level: juniorimportance: must knowfreq 58%

basics

~20 s

testCaseId identifies the logical test case - one value for a method however many rows it runs. historyId identifies one trend series: that case plus the parameter values it ran with. Two rows share testCaseId but not historyId.

open as a page

In an Allure 2 report, a results directory holds several `-result.json` files for the same test from one run. How does `RetryPlugin` group them, and which attempt does the generated report show?

level: middleimportance: must knowfreq 58%

basics

~20 s

RetryPlugin buckets every parsed result by the retry hash it derives, keeps the attempt with the newest start time as the visible one, hides the rest, and hangs them off the survivor as a retries block.

open as a page

In Allure 2, the previous report's `history/` directory is what gives the next report a trend. In which direction does that copy go, and why must it run before `allure generate` rather than after?

level: middleimportance: should knowfreq 44%

basics

~20 s

Out of the finished report and into the next run's results directory, before the generator runs. The trend is computed in that single pass, so anything placed later, or placed in the report directory, is never read.

open as a page

Allure's `StatusDetails` carries three booleans - `known`, `muted` and `flaky`. What different claim does each one make about a test result?

level: middleimportance: should knowfreq 46%

basics

~20 s

Each is an independent boolean on StatusDetails. flaky claims the case is unreliable, muted claims its result should not be held against the run, and known claims the failure is already accounted for. Setting one never sets another.

open as a page

When an Allure adapter leaves `historyId` unset on a `TestResult`, how does the `allure-java` writer compute a value before the result file is written, and what makes that computed value change?

level: middleimportance: should knowfreq 50%

basics

~20 s

When historyId is missing but testCaseId is set, the writer concatenates testCaseId with the run's parameter name and value pairs, sorted and minus any marked excluded, then hashes the result. A changed parameter value changes the key.

open as a page

In Allure 2, `RetryPlugin` flags every attempt but the last as hidden. What does that flag actually remove from the generated report, and what survives of those earlier attempts?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Hidden removes an attempt from the default result set most plugins read, so trees and widgets stop counting it. The attempt is still serialised to its own test-case data file in the report, reachable from the survivor by uid.

open as a page

In ReportPortal, a second attempt of a test item arrives while a launch is running. What does `retryOf` link, and what happens to the earlier item's place in that launch?

level: seniorimportance: should knowfreq 40%

basics

~20 s

retryOf points a superseded attempt at the item that replaced it. The newer attempt becomes the main item, keeping the tree position and the statistics; the older one has its path and launch link cleared and stops counting.

open as a page

Your pipeline copies the previous Allure 2 report's `history/` directory into each new run's results, yet the trend chart keeps restarting from one or two points. How do you diagnose that, and why can the lost trend not be rebuilt from the archived results directories?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Each report's history folder is built from the previous one plus the current run, so it is a chain. A single build that generates without it truncates everything earlier, permanently: a results directory holds only its own run.

open as a page

An Allure report shows a test as flaky, yet the `-result.json` the run wrote for it has `statusDetails.flaky` set to false. What are the two ways a result acquires that mark, and which one produced this one?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Two routes exist: the run writes the flag into statusDetails, or the report generator derives it from the case's stored series of earlier outcomes. A false flag in the raw file means the generator derived this one.

open as a page

A Java test method covered by an Allure report is renamed and moved into another class. What happens to that test's trend in the next generated report, and why is there no error telling you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Renaming or moving the method changes testCaseId, so historyId changes too and the result no longer matches its stored series. The report shows a new test with one point. Nothing errors: an unmatched key looks like a new test.

open as a page

Allure 2 backs its retry trend with a `retry-trend.json` file. What are the two numbers it stores for each build, and what is each one actually counting?

level: middleimportance: nice to knowfreq 28%

basics

~10 s

Each point in retry-trend.json holds a data map with two counters: retry, the number of attempts that were superseded in that build, and run, the number that stayed visible. Both count attempts, not tests.

open as a page

Allure 2's command line has no history option, so its trend depends on copying the previous report's `history/` directory into the next run's results. What does Allure 3 change about that, and does the carry-between-builds problem disappear?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Allure 3 makes history configuration explicit: historyPath, historyLimit and appendHistory are config keys and there is a history command, so the data no longer has to be smuggled in through the results directory. It still has to survive between builds.

open as a page

An Allure adaptor writes a result whose `StatusDetails` has `known`, `muted` and `flaky` all set to true. Which of the three survive into the generated report in Allure 2, and which in Allure 3?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only flaky survives in Allure 2, whose generator takes message, trace and flaky off StatusDetails and never reads the other two. Allure 3 keeps flaky and muted but drops known. Writing a flag does not mean a report honours it.

open as a page

An Allure report shows every case of a parameterised suite with a single history point on every build, and one of the recorded parameters carries a value that differs each run. How does that produce the symptom, and how would you fix it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

A per-run parameter value feeds the history key, so every build produces a fresh historyId and a series one point long. Fix it by flagging that parameter as excluded from the key, or by assigning historyId yourself.

open as a page