skip to content

Allure's `StatusDetails` carries three booleans - `known`, `muted` and `flaky`. What different claim does each one make about a test result?

level: middleimportance: should knowfreq 46%

answer

  1. three claims, not three degrees
  2. unreliable, silenced, already accounted for
  3. none of them touches status
  4. setting one never sets another
  5. primitive booleans, so false is written

basics

~20 s

Each is an independent boolean on StatusDetails. flaky claims the case is unreliable, muted claims its result should not be held against the run, and known claims the failure is already accounted for. Setting one never sets another.

solid answer

~40 s

They are three unrelated claims that happen to sit next to each other on the same object. `flaky` says the case cannot be trusted to give the same answer twice; `muted` says its outcome should not be held against the run; `known` says the failure is already accounted for. All three are primitive booleans, so once a `statusDetails` object exists all three are serialised, including the ones that are `false` - the only absence available is the absence of the whole object. None of them touches `status`, so a muted failure is still `FAILED` and still counted as one. And they are set by whatever writes the result, usually an adaptor copying a marker declared on the test itself.

go deeper

for a junior

Know the three names and that they sit together on StatusDetails alongside the message and trace. Being able to say they are separate booleans rather than one setting is enough at this level.

for a middle

Explain each claim in its own words and show they do not interact: no adaptor derives one from another, and none of them rewrites status or the failure text on the result.

for a senior

Bring up the direction problem. Report-side machinery reads these booleans as match conditions and never writes them, so a result carrying the mark got it from somewhere upstream of the report.

for a principal

Argue about which claims belong on a result at all. A boolean with no attached reason ages badly, so decide deliberately what evidence must accompany a claim before your teams are allowed to record it.

## One object, three independent booleans `StatusDetails` is the object that hangs off an Allure test result and carries everything about the result that is not its `status`. Three of its seven fields are booleans -- `known`, `muted` and `flaky` -- and the remaining four (`message`, `trace`, `actual`, `expected`) hold the failure text. The three booleans are not degrees of one idea. They are three unrelated claims that happen to live next to each other. | flag | the claim it makes | what it is not | |---|---|---| | `flaky` | this case cannot be trusted to give the same answer twice | not a record that a repeat attempt happened | | `muted` | this case's outcome should not be held against the run | not a change to the outcome itself | | `known` | this failure is already accounted for | not a link, and not an identifier | ## Where the values come from None of the three is computed by the model class. They are set by whatever writes the result, which in practice is the framework adaptor: it reads a marker declared on the test itself, such as an annotation on the method or class or a tag on a scenario, and copies it onto the result's `StatusDetails` before the file is written. That has a consequence people miss constantly. On the writer side these flags record **a declaration made in the source**, not **an observation made during the run**. A test declared unreliable is written with `flaky` true whether or not it misbehaved this time, and a test nobody declared is written with `flaky` false even if it failed and then succeeded. ## They do not interact Each boolean is set on its own and read on its own: - Setting `muted` does not set `flaky`. A silenced case is not thereby an unreliable one. - Setting `flaky` does not set `muted`. An unreliable case is still held against the run unless something else says otherwise. - Setting any of the three leaves `status` alone. A muted failure is still `FAILED`, and every count derived from `status` still counts it. - Nothing in the writer derives one from another, so the file can carry any combination of the eight. ## Absence and the value false ```json "statusDetails": { "known": false, "muted": false, "flaky": true, "message": "socket closed while draining the queue" } ``` The writer's mapper omits only nulls. `known`, `muted` and `flaky` are primitive booleans on the model class, so they can never be null and are always serialised once a `statusDetails` object exists, including the two that are `false` above. The only absence available is the absence of the whole object, which happens when nothing set any of its seven fields. There is no third state: a consumer that sees no `flaky` key has to read it as `false`, exactly like an explicit `false`. That matters when you are diffing two results directories written by different adaptors, because a difference in these keys is a difference in whether `statusDetails` exists at all, never a difference in the flags themselves. ## The direction of the flag The commonest confusion here is direction. Downstream machinery that mentions flakiness -- a report's grouping rules, a widget, a saved filter -- **reads** these booleans. It does not write them. A report's grouping rule can name the flag as one of its match conditions beside status, message and trace, which selects results whose own flag equals it, and a rule that names it still never marks anything. If you are working out why a result carries the mark, the answer is always upstream of the report's grouping, never inside it. ## What each one buys you `flaky` is the one worth setting deliberately, because it is the only one of the three that both live Allure majors carry into their report model. `muted` is worth setting when you need the run's readers to know that a result is deliberately not load-bearing, but you must check that your consumer honours it before you rely on it. `known` is the weakest of the three: it says a failure is not news, and it carries nothing that lets a reader find out why, so it is only useful next to something that does. None of the three is a decision. They are claims recorded on the result, and every consumer is free to honour, ignore or recompute them.

  • Does setting `muted` on a result change the run's pass and fail counts?
    Not by itself. `muted` is a claim carried on `StatusDetails`; the result keeps whatever `status` the run gave it, and the counts are computed from `status`. Whether the flag is honoured at all is up to the consumer reading the results, and the two live Allure majors do not agree about it.
  • Two adaptors write results for the same suite and one of them never sets these flags. What does a consumer see?
    Nothing distinguishable from a deliberate `false`. Once a `statusDetails` object exists all three booleans are serialised, and where no object exists the consumer sees absence, which every reader treats as not flaky, not muted, not known. There is no third state to represent 'not stated'.

saying these in an interview costs you the question

  • Treats muted and flaky as the same claim
  • Thinks marking a result muted makes it pass
  • Says a report's grouping rule can set the flaky flag
  • Assumes an unset boolean is missing from the file