skip to content

Marking a Flaky Outcome

The mark itself: which field carries it, the neighbouring claims it gets confused with, and the two ways a tool arrives at one - derived from a stored series, or taken from the run.

on this pageshow

explore

questions

4

In Allure's result model, `io.qameta.allure.model.Status` has exactly four values and none of them is flaky, so where does a result's flaky mark actually live and what writes it?

level: juniorimportance: must knowfreq 68%

answer

  1. outcome and claim live apart
  2. four status constants, none of them flaky
  3. a boolean hanging off the result
  4. StatusDetails, next to known and muted
  5. setFlaky(...) leaves status untouched

basics

~10 s

Flakiness is a boolean on StatusDetails, set with setFlaky(...), not a status value. The result keeps one of Status's four values - FAILED, BROKEN, PASSED or SKIPPED - and the flag rides alongside it.

solid answer

~40 s

Allure's writer model has no flaky outcome. `io.qameta.allure.model.Status` declares exactly `FAILED`, `BROKEN`, `PASSED` and `SKIPPED`, and every result carries one of them. The flaky claim is a separate boolean on the `StatusDetails` object hanging off the result, flipped with `setFlaky(...)`; `StatusDetails` also carries `known`, `muted`, `message`, `trace`, `actual` and `expected`. Because status and flag are orthogonal, a passing result can be marked flaky and so can a failing one, and the flag says nothing about which way the run went. In the serialised `-result.json` it shows up as `"flaky": true` inside `"statusDetails"`, next to an unchanged `"status"`. Anything that counts flaky cases reads that boolean, never the status.

code

json · 16 lines
json
{
  "uuid": "9f1c0d2a-6d3e-4a51-9c77-0b2e5a41e8c3",
  "historyId": "3b7e1f9c8a4d2e5061f7c9b0d3a4e5f6",
  "name": "checkout completes with a saved card",
  "fullName": "shop.CheckoutTest.completesWithSavedCard",
  "status": "passed",
  "stage": "finished",
  "statusDetails": {
    "known": false,
    "muted": false,
    "flaky": true,
    "message": "declared unreliable while the payment sandbox is shared"
  },
  "start": 1757000000000,
  "stop": 1757000004120
}

go deeper

for a junior

Be able to say that Allure records flakiness as a boolean on StatusDetails and that the four status values are FAILED, BROKEN, PASSED and SKIPPED, with no flaky member among them.

for a middle

Explain that the two are orthogonal: setFlaky(...) flips one boolean and leaves status, the message and the trace exactly as they were, so any combination of outcome and flag is legal.

for a senior

Show where this bites. A consumer that counts flaky cases must read the boolean, so a filter or gate written against status alone reports zero flaky tests however many carry the mark.

for a principal

Own the modelling argument: a fifth status forces every existing reader of the four to grow a branch and makes flaky and failed mutually exclusive, while a side flag stays ignorable by construction.

## Two separate things on one result Allure's writer model, the one `allure-java` serialises into a results directory, keeps a test's **outcome** and a test's **reputation** in two different places. The flaky mark belongs to the second. The outcome is the result's `status` field, typed by `io.qameta.allure.model.Status`. That enum declares exactly four constants: `FAILED`, `BROKEN`, `PASSED` and `SKIPPED`. There is no fifth constant for flakiness, no `FLAKY` member, and no way to smuggle one in, because the serialiser writes the enum's own lowercase wire value. A `-result.json` can therefore only ever say `"status": "failed"`, `"broken"`, `"passed"` or `"skipped"`. The reputation lives on `StatusDetails`, an optional object hanging off the result. `StatusDetails` declares seven fields, in this order: 1. `known` -- a boolean. 2. `muted` -- a boolean. 3. `flaky` -- a boolean. 4. `message` -- the short failure text. 5. `trace` -- the stack trace. 6. `actual` -- the observed value. 7. `expected` -- the value the assertion wanted. The flaky mark is item three: a plain boolean, flipped with `setFlaky(...)`. ## What setFlaky(...) actually changes Nothing except that one boolean. - It does **not** change `status`. A result marked flaky keeps whatever outcome the run gave it. - It does **not** turn a red run green, and it does not make a green run suspect to the runner. - It does **not** record anything about attempts. The flag is a claim about the case, not a log of what this execution did. - It does **not** fill `message` or `trace`; those still carry the failure text, if there was one. Because the two are orthogonal, all four combinations are legal and all four occur in real results directories: a `passed` result with `flaky` false (the ordinary case), a `passed` result with `flaky` true (a case somebody declared unreliable that happened to work), a `failed` result with `flaky` false, and a `failed` result with `flaky` true. ## What it looks like on disk ```json { "name": "checkout completes with a saved card", "status": "failed", "statusDetails": { "known": false, "muted": false, "flaky": true, "message": "expected the order id to be present" } } ``` `"status"` is untouched by the flag sitting beside it. Note also that all three booleans are written even though two of them are `false`. They are primitive booleans on the model class, and the writer's mapper omits only nulls, so a primitive can never be dropped. Absence of the whole `"statusDetails"` object is possible, and happens when nothing set any of its fields, but absence of one boolean while its siblings are present is not. ## Why a flag and not a status | | `status` | `statusDetails.flaky` | |---|---|---| | type | four-valued enum | boolean | | always present | yes | only when `statusDetails` exists | | what it records | how this execution ended | a claim about the case | | who must handle it | every reader | only readers that care | Modelling flakiness as a fifth status would force every consumer of the four to grow a branch for it, and would make the two facts mutually exclusive: a case could then be flaky or failed but not both, which is exactly the pair you most want to see together. A side flag is ignorable by construction. A tool that knows nothing about flakiness reads `status`, gets a well-formed answer, and never notices the boolean. ## What this means when you read results Anything that counts flaky cases has to read the boolean. A dashboard, a filter or a gate built by switching on `status` alone will report zero flaky tests no matter how many carry the mark, because the mark is not in the field it is looking at. Equally, a case that is genuinely marked is still counted in whatever bucket its `status` puts it in: marking a failing test flaky does not remove it from the failure count, and it does not move it out of `FAILED` into some softer category. The same shape holds for both live Allure majors, because both read the file `allure-java` writes. The flag is a boolean parsed off the result and the status stays one of a small closed set. Where the majors differ is in what they do with the flag afterwards, not in where it lives.

  • If a result's `status` is `passed` and its `statusDetails.flaky` is true, what has the writer actually told you?
    Only that something declared this case unreliable. The flag is a claim about the case, not a record of what this execution did, and the execution passed. Nothing in the flag says an attempt was repeated or how many attempts there were; those numbers live in entirely different fields.
  • What happens to `statusDetails` in the written `-result.json` when nothing sets any of its fields?
    The whole object is omitted. The writer's mapper serialises with non-null inclusion, so a null `statusDetails` disappears. But as soon as the object exists, all three booleans are written including `false`, because they are primitive booleans on the model class and can never be null.

saying these in an interview costs you the question

  • Says Allure has a FLAKY value in its status enum
  • Thinks setFlaky(...) changes the result's status
  • Assumes a flaky mark means the case passed on a repeat
  • Believes marking a result flaky turns a red run green
open as a page

Allure's `StatusDetails` carries three booleans - `known`, `muted` and `flaky`. What different claim does each one make about a test result?

level: middleimportance: should knowfreq 46%

basics

~20 s

Each is an independent boolean on StatusDetails. flaky claims the case is unreliable, muted claims its result should not be held against the run, and known claims the failure is already accounted for. Setting one never sets another.

open as a page

An Allure report shows a test as flaky, yet the `-result.json` the run wrote for it has `statusDetails.flaky` set to false. What are the two ways a result acquires that mark, and which one produced this one?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Two routes exist: the run writes the flag into statusDetails, or the report generator derives it from the case's stored series of earlier outcomes. A false flag in the raw file means the generator derived this one.

open as a page

An Allure adaptor writes a result whose `StatusDetails` has `known`, `muted` and `flaky` all set to true. Which of the three survive into the generated report in Allure 2, and which in Allure 3?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only flaky survives in Allure 2, whose generator takes message, trace and flaky off StatusDetails and never reads the other two. Allure 3 keeps flaky and muted but drops known. Writing a flag does not mean a report honours it.

open as a page