In an Allure 2 report, a results directory holds several `-result.json` files for the same test from one run. How does `RetryPlugin` group them, and which attempt does the generated report show?
answer
- several results, one visible row
- a grouping key, then a winner
- newest start time survives
- the rest are flagged hidden
basics
~20 sRetryPlugin buckets every parsed result by the retry hash it derives, keeps the attempt with the newest start time as the visible one, hides the rest, and hangs them off the survivor as a retries block.
solid answer
~50 s`RetryPlugin` is one of Allure 2's bundled extensions and it runs early, over every parsed result rather than the filtered set. It builds a map keyed by the value `getRetryHash()` returns, skipping any result whose hash is null, so every attempt of the same case with the same parameters lands in one bucket. Inside a bucket it picks the attempt with the newest start time as the survivor; status plays no part in that choice. Each of the other attempts is flagged `hidden` and `retry`. The survivor then receives an extra block named `retries`, holding one compact entry per earlier attempt, plus `setRetriesCount(...)` with how many there were and `setRetriesStatusChange(...)` with whether any of them ended in a different status. The report therefore shows one row per test, the last attempt, with the earlier ones reachable underneath it.
go deeper
Be able to say that Allure 2 collapses several results for the same test into a single row showing the last attempt, with the earlier ones tucked underneath it rather than listed beside it.
Explain the three moves: bucket by the retry hash, take the newest attempt by start time, then flag the others hidden and hang a retries block on the survivor. Say plainly that status never influences the choice.
Show you know when this runs relative to everything else, and what that ordering buys: the trees, widgets and trends built afterwards count survivors only, because the reconciliation already happened upstream of them.
Own the question of what a consolidated report must still expose about attempts — which fields have to reach the summary a reader actually opens, and what a reader can no longer reconstruct once the earlier attempts are collapsed.
## The problem the plugin exists to solve A runner that repeats a failing test writes **one result file per attempt**. Nothing in the Allure results format marks those files as related: each `-result.json` is a self-contained record with its own `uuid`, its own start and stop times, and its own `status`. If the generator rendered everything it parsed, a suite of two hundred tests with a handful of retries would produce more than two hundred rows, several of them the same test with contradictory outcomes, and every count in the report would be inflated by exactly the repeated work. `RetryPlugin`, in the `io.qameta.allure.retry` package, is the step that collapses that back down. It is registered among Allure 2's bundled extensions and it runs **early** — ahead of the plugins that build the trees, the widgets and the trends — so everything downstream sees a set of results that has already been reconciled. ## How the grouping actually works The plugin is an aggregator: it is handed every parsed launch and mutates the results in place. Its work is three moves. 1. **Bucket.** It flattens the results of every launch into one stream, keeps only those for which `getRetryHash()` returns a non-null value, and collects them into a map from that hash to a list of results. Two attempts of the same test with the same parameters produce the same hash and land in the same bucket. A result with **no** hash never enters the map at all and is left alone as an ordinary, visible test. 2. **Pick a survivor.** Within a bucket the plugin orders by start time, newest first, and takes the newest attempt that is not already hidden. Nothing about status enters this decision: a bucket whose newest attempt failed shows that failure, even if an earlier attempt passed. 3. **Demote the rest.** Every other attempt in the bucket is put through a small preparation step that sets two booleans on it — `hidden`, which takes it out of the result set most plugins consume, and `retry`, which marks it as a superseded attempt for anything that deliberately counts attempts. Notice what the grouping key does **not** depend on: the file the result came from, the order the files were read, or which results directory was passed on the command line. The plugin flattens across launches first, so attempts split over several results directories still reconcile into one bucket. ## What lands on the survivor Three things are attached to the attempt that stays visible. | what is set | shape | what it holds | |---|---|---| | the `retries` extra block | a list on the surviving result | one compact entry per earlier attempt | | `setRetriesCount(...)` | an `int` field | how many earlier attempts there were | | `setRetriesStatusChange(...)` | a `boolean` field | whether any of them ended in a different status | Each entry in the block is deliberately small. It carries the earlier attempt's `uid`, its status, its status message and its time — and nothing else. It is a **pointer plus a summary**, not a copy: the earlier attempt's own record still exists elsewhere in the report and the `uid` is how the surviving result's page reaches it. The count is the size of that list, which means it counts **every attempt except the survivor**. A test the runner executed three times shows a count of two. The status-change flag is computed by taking the statuses of the earlier attempts, discarding any that equal the survivor's, and asking whether anything is left; it is a plain boolean and it says only that the outcome moved, never in which direction. ## The things that catch people out - **The last attempt wins, not the best one.** There is no preference for a passing attempt. If a test passed and was then re-executed into a failure, the failure is what the report shows. - **Ordering is by start time only.** Allure 2 compares nothing else, so two attempts stamped with the same start have no defined tiebreak. - **A missing hash silently opts a result out.** If the derived retry hash comes back null, that result is never grouped, never hidden, and never gains a retries block — even with a sibling attempt sitting in the same directory. Two rows for one test in a report that otherwise collapses retries is almost always this. - **The plugin decides nothing about retrying.** It is purely a consolidation step over results that already exist; whether the runner retried anything, and how often, was settled long before the generator ran. ## The same idea in Allure 3 Allure 3 keeps the shape and moves the code. A `RetrySubstore` holds results in a map keyed by `retryHash`, re-sorts the bucket newest-first on every insert with the ingest order as a tiebreak when starts are equal, and then walks the bucket setting `isRetry` to true on everything but the element at index zero. The model exposes `retries` on a test result so the surviving attempt can offer the others. The vocabulary differs — `isRetry` rather than a hidden flag, an explicit tiebreak rather than none — but the reconciliation is recognisably the same one.
- What happens to a result whose retry hash comes back null?It never enters the grouping map: the plugin filters on a non-null hash before collecting. Such a result stays an ordinary, visible test with no retries block and a retry count of zero, even when another attempt of the same test sits in the same results directory.
- If two attempts carry identical start times, which one becomes the survivor?Allure 2 compares start time and nothing else, so identical stamps leave no defined tiebreak — whichever the reduction happens to keep, wins. Allure 3 closes that gap: its `RetrySubstore` falls back to the order results were ingested when the starts are equal.
saying these in an interview costs you the question
- Thinks the report shows the first attempt, not the last
- Assumes the plugin prefers a passing attempt over a later failure
- Believes the plugin decides whether a test gets retried
- Expects grouping by test name rather than by a derived hash