Allure 2's JUnit-XML reader decides a case is flaky by checking whether the `<testcase>` contains `<rerunError>` or `<rerunFailure>`, and never looks at `<flakyFailure>` or `<flakyError>`. What does that produce in the generated report, and what does it tell you about relying on this file?
answer
- the reader and the writer disagree
- only one of the two prefixes is inspected
- the wrong pair gets the mark
- hard red shows up as flaky
basics
~10 sThe report inverts the file's meaning. Cases labelled flaky are the ones Maven Surefire recorded as failing on every attempt, while the cases Surefire actually calls flaky get no mark at all.
solid answer
~40 sIt marks exactly the wrong cases. In Maven Surefire's file, `<rerunFailure>` and `<rerunError>` appear on a case that failed on *every* attempt, while `<flakyFailure>` and `<flakyError>` appear on a case that recovered. Allure 2 (2.47) checks only for the first pair, so an imported build shows hard-red cases labelled flaky and shows nothing for the cases that genuinely flaked - the two flaky element names appear nowhere in its source. Allure 3 (3.16.1) does not repeat that, but its JUnit-XML reader inspects none of the four, so the same file arrives there with no flake information at all. The lesson is that this format carries no enforcement: nothing validates that a reader's interpretation matches the writer's, so before trusting a consolidated flake number, find out what the consumer actually looks for.
code
java · 4 linesprivate boolean isFlaky(final XmlElement testCaseElement) {
return testCaseElement.contains(RERUN_ERROR_ELEMENT_NAME)
|| testCaseElement.contains(RERUN_FAILURE_ELEMENT_NAME);
}go deeper
Know that <rerunFailure> and <rerunError> belong to a case that never passed, so a tool keying its flaky label off those two elements is labelling the wrong cases.
Explain that this format carries no enforcement of meaning: a writer and a reader can attach opposite readings to the same element names and nothing in the file or the parse errors.
Describe how you would catch this in practice - a fixture holding both shapes, run through the consumer, with the resulting labels compared against what the writer recorded.
Own the call on whether the organisation keeps consolidating through a format nobody validates, or derives the verdict once at the producer and ships that as the contract.
It marks exactly the wrong cases, and nothing anywhere reports a problem. This is the cleanest live example of what the JUnit-XML interchange format does not do: it defines element names, and it defines nothing about what a reader must conclude from them. ## What the two prefixes mean to the writer Maven Surefire chooses the element name from the case's **final** outcome: - `<rerunFailure>` and `<rerunError>` appear on a case that **failed on every attempt**. The first attempt is the plain `<failure>` or `<error>`, and each later attempt gets a rerun element. The build is red. - `<flakyFailure>` and `<flakyError>` appear on a case that **failed and then passed**. Every failed attempt gets a flaky element, the final pass writes nothing, and the case is counted in the suite's `@flakes`. The build is green. ## What Allure 2's reader does with them Allure 2 (2.47) decides flakiness from the presence of the rerun pair only - it asks whether the `<testcase>` contains `<rerunError>` or `<rerunFailure>`, and returns true if either is there. The two flaky element names do not appear anywhere in its source. Compose the two halves and the result inverts: | the case, per Surefire | elements written | Allure 2's flaky mark | |---|---|---| | failed on every attempt (hard red) | `<failure>` + `<rerunFailure>` | **marked flaky** | | failed then passed (genuinely flaky) | `<flakyFailure>` | **not marked** | | passed first time | none | not marked | So the report labels as flaky the cases the producer says never passed, and says nothing about the cases the producer actually calls flaky. Both halves are wrong at once, which is why the symptom is confusing in practice: the flaky count is not merely inaccurate, it is anti-correlated with the truth. Allure 3 (3.16.1) does not repeat the mistake, but it does not fix the problem either - its JUnit-XML reader inspects **none** of the four repetition elements. The same file imported there arrives with no flake information at all, and nothing in the import reports the loss. ## Why nothing catches it Every layer that could have caught this is looking somewhere else: - **The schema cannot.** `surefire-test-report.xsd` declares element names, cardinalities and attributes. It has no way to express "a reader must treat this element as a recovered attempt", and the file in question validates perfectly. - **The build cannot.** Both tools parse the file without an error, so nothing fails, nothing warns, and the pipeline stays green. - **Spot-checking usually cannot.** A build with no re-runs contains none of the four elements, so the reader and the writer agree on every file the team looks at until the day re-runs are switched on somewhere. - **The vocabularies differ elsewhere too.** The same Allure 2 plugin accepts both `<testsuite>` and `<testsuites>` as a root, maps `<failure>` to `FAILED` and `<error>` to `BROKEN`, and treats a `notrun` status value as skipped. Two writers and one reader can hold three different vocabularies over one file. ## What to do about it The general lesson is worth more than the specific bug: **for this format, a consumer's interpretation is a fact you have to check, not a property you can assume.** Concretely: 1. **Derive the verdict yourself, at the producer.** Treat `<flakyFailure>` and `<flakyError>` as recovered attempts, and a `<failure>` or `<error>` with `<rerunFailure>` or `<rerunError>` as a case that never passed. Emit that verdict as data - a field, a tag, a separate summary - rather than hoping the consumer re-derives it correctly. 2. **Test the consumer against a fixture, not against your own builds.** One file containing a plain pass, a case that never passed, and a case that recovered, imported into whatever renders your reports, with the resulting labels compared against what the fixture says. That test is cheap to write and catches an entire class of interpretation drift. 3. **Do not treat a consolidated flake number as evidence until you know what produced it.** Before quoting a flake figure from a report, find out which element or field the tool looks for. If the answer is "the rerun elements", the number means the opposite of what its label says. 4. **Prefer a signal the consumer cannot misread.** If the reporting tool supports an explicit mark on a result, set it deliberately from a verdict you computed, rather than letting the tool infer one from element names. None of this is a criticism of a particular tool version. It is a property of an interchange format whose meaning lives in convention rather than in the schema, and it is the reason a consolidated result pipeline needs at least one test that crosses the producer/consumer boundary end to end.
- If you had to get a trustworthy flake signal out of these files, what would you do?Read the four elements yourself rather than relying on a consumer's interpretation: treat `<flakyFailure>` and `<flakyError>` as recovered attempts, and a `<failure>` with `<rerunFailure>` or `<rerunError>` as a case that never passed. Then feed the consolidator the derived verdict as data it cannot re-interpret, and test it against a fixture holding both shapes.
- Does importing the same file into Allure 3 fix the problem?No - it removes the wrong answer without supplying a right one. Allure 3 (3.16.1) inspects none of the four repetition elements when reading JUnit-XML, so a build that flaked and a build that failed on every attempt both arrive with no flake information, and nothing in the import reports the loss.
saying these in an interview costs you the question
- Assumes a report's flaky label was derived from the flaky elements
- Believes the XML schema forces readers to agree on meaning
- Trusts a consolidated flake count without checking the reader
- Thinks <rerunFailure> marks a case that eventually passed
- Expects a parse error when a reader ignores an element