An Allure report shows a test as flaky, yet the `-result.json` the run wrote for it has `statusDetails.flaky` set to false. What are the two ways a result acquires that mark, and which one produced this one?
answer
- the mark has more than one origin
- written by the run, or computed later
- the generator ORs the two together
- derivation needs a currently failing run
- grep the raw result file first
basics
~20 sTwo routes exist: the run writes the flag into statusDetails, or the report generator derives it from the case's stored series of earlier outcomes. A false flag in the raw file means the generator derived this one.
solid answer
~40 sA mark either comes from the run or is computed at generation time. In the first route the framework adaptor reads a marker declared on the test and calls `setFlaky(...)` before the `-result.json` is written, so the flag is a byte on disk. In the second the generator looks up the case's carried-forward series and adds its own verdict, which fires only when this run's status is `FAILED` or `BROKEN` and a `PASSED` appears before the last `FAILED` within a bounded window of recent entries. The two are combined with a logical OR, and the report carries no field saying which fired. Here the raw file says `false`, so the derived route produced it - which also tells you this run failed.
code
bash · 5 lines# what the run itself wrote, before any report was generated
grep -l '"flaky": *true' allure-results/*-result.json
# how many results carry the flag either way
grep -oh '"flaky": *[a-z]*' allure-results/*-result.json | sort | uniq -cgo deeper
Know that the flag in the report is not always the flag the run wrote, and that the raw result file is where you look to tell the two apart. That distinction alone is worth stating.
Explain both routes mechanically, including that they are combined with an OR so either alone is enough, and that the derived route needs a stored series to work from.
Demonstrate the diagnosis. Read the raw results before interpreting the report, and state the conditions the derived route implies about the current run and the recent series.
Take a position on which route your organisation should rely on, given that a derived mark silently disappears whenever the stored series does and a declared one never reflects today's run.
## The mark has two origins, and the report shows only their union A flaky flag on an Allure result can get there in exactly two ways, and they happen at different times, on different sides of the results directory. | | taken from the run | derived from the stored series | |---|---|---| | when | while the test executes | while the report is generated | | who | the framework adaptor | the report generator | | what it looks at | a marker declared on the test | the case's earlier outcomes | | visible in `-result.json` | yes, as `statusDetails.flaky` | no, the file is never rewritten | | needs history | no | yes | ## Route one: taken from the run The adaptor reads a marker declared on the test source -- an annotation on the method or class, a tag on a scenario -- and calls `setFlaky(...)` on the result's `StatusDetails` before writing the `-result.json`. Two things follow. First, the flag is a **declaration**, not an observation: it is true because somebody wrote it on the test, not because the runner watched the test misbehave. Second, it is durable and inspectable, because it is a byte in the results directory that you can read without generating anything. ## Route two: derived from the stored series When the report is generated, the generator looks up the case's carried-forward series of earlier outcomes and computes its own verdict, then merges it with whatever the run wrote. The rule it applies is narrow and worth memorising, because it is the same on both live Allure majors: 1. If there is no stored series for this case at all, the derived verdict is false. 2. If **this run's** status is not `FAILED` or `BROKEN`, the derived verdict is false. A passing case is never derived flaky, however mixed its past. 3. Otherwise, within a bounded window of the most recent prior entries, there must be a `PASSED` that occurs **before** the last `FAILED`. That ordering is the whole test: it is asking whether the case has switched back and forth recently, not merely whether it has ever failed. So the derived route only ever fires on a case that is red right now and was green not long ago. ## The report shows the union The two routes are combined with a logical OR. The generator sets the result's flag to `run-declared OR history-derived`, so either route alone is enough, and the derived route can never clear a flag the run set. Two consequences fall straight out of that: - **The report carries no field saying which route fired.** There is one boolean at the end, and it looks identical either way. - **The mark is not reproducible from the results directory alone.** Regenerate the same results with an empty history and the derived contribution disappears; regenerate with the history present and it comes back. Nothing about the run changed. ## Reading the evidence in your case The `-result.json` is the run's own testimony, written before any report existed and never rewritten by generation. So: - Raw file says `"flaky": true` and the report says flaky: the run declared it. Look at the test source for the marker. - Raw file says `"flaky": false` and the report says flaky: **the generator derived it.** That is the case in the scenario, and it tells you two further things for free -- this run's status must be `FAILED` or `BROKEN`, and the carried-forward series for this case must contain a recent switch from passing to failing. - Raw file says `"flaky": true` and the report says nothing: you are looking at the wrong result, or at a consumer that does not surface the flag. The practical move is to grep the raw results before you interpret the report, because the two answers disagree far more often than people expect, and only one of them is evidence of what the run actually did. ## Why this is worth knowing A team that treats the flag as a fact about the execution will draw the wrong conclusion in both directions. A case marked from a declaration tells you nothing about today's run. A case marked by derivation is telling you about the last few runs, not this one, and it will silently un-mark itself the moment the stored series is missing -- a fresh workspace, a first build on a new branch, a renamed case whose series key moved. Neither is wrong; they are simply answers to different questions, presented in the same box.
- The stored series for this case is present but the mark still does not appear. What would you check first?Whether this run actually failed. Derivation is skipped entirely unless the current status is `FAILED` or `BROKEN`, so a case that passed is never marked from history however mixed its past looks. Only after that would I check that the series really reached the generator rather than starting fresh.
- Does this derivation behave the same way in Allure 2 and Allure 3?Yes. Both majors apply the same rule and OR its verdict with whatever flag the run wrote: the current status must be `FAILED` or `BROKEN`, and within a bounded window of the most recent prior entries a passing outcome must occur before the last failing one. Neither major can clear a flag the run set.
It is like a fire alarm that can be set off either by someone pressing the manual call point or by a smoke sensor. The bell sounds identically both ways, and the panel records only that it rang, not which one rang it.
saying these in an interview costs you the question
- Assumes the flag always means an attempt was repeated
- Thinks a passing test can be derived flaky from history
- Believes the report records which route set the flag
- Says the derived verdict overwrites what the run wrote