skip to content

A colleague divides the tallest bar on an Allure 2 report's categories-trend tile by the `total` in `widgets/summary.json` and calls the result the share of tests hitting that failure class. What is each number's real denominator, and why does that ratio not mean what they think?

level: seniorimportance: should knowfreq 44%

answer

  1. numerator and denominator, different populations
  2. one tile counts failures, one counts everything
  3. this run versus a series of builds
  4. the bar's denominator is failed plus broken

basics

~20 s

The bar counts one build's failures; the total counts every result of the current run. Dividing them puts a failure numerator over an all-result denominator, and usually crosses two different builds as well, so the percentage means nothing.

solid answer

~40 s

In Allure 2 the two files describe different populations over different windows. `widgets/summary.json` is one run: `total` is the sum of `failed`, `broken`, `passed`, `skipped` and `unknown`, so it counts every result the report kept. `widgets/categories-trend.json` is a series of per-build points, newest first, and each point's `data` map counts how many results were assigned each category on **that** build — only categorised results contribute, and since every failed or broken result lands in some bucket, those values partition that build's failures. So the fraction puts a failure count over a test count, and unless the reader took the head of the array it also crosses two builds. The honest denominator for a bar is `failed` plus `broken` from the same build's summary file.

go deeper

for a junior

Be able to say that different overview panels count different things, and that a percentage built from two panels needs both numbers to come from the same run and the same population before it means anything at all.

for a middle

Explain each file's shape: one is a single object about this run with five status counters, the other an array of per-build points whose values count categorised results. Name the mismatch precisely rather than saying the number is roughly wrong.

for a senior

Show the production instinct — refuse the number, state the correct denominator, and check the head of the trend array against the report's own build identity before comparing anything. Expect to be handed such a ratio in a release summary.

for a principal

Own how numbers leave the report. Decide which comparisons your organisation may publish off a generated report at all, and where a figure needs a queryable store behind it rather than two tiles divided by hand in a meeting.

## Two tiles, two different populations The mistake is not arithmetic; it is that the numerator and the denominator are drawn from different populations, over different windows. Take the two files apart. **`widgets/summary.json`** describes exactly one run. Its `statistic` block holds `failed`, `broken`, `passed`, `skipped` and `unknown`, and `total` is the sum of those five. It counts every result the report kept, whatever its outcome. It is a **whole-run** number, and it counts results rather than distinct test cases. **`widgets/categories-trend.json`** is not about one run at all. It is an array of points, one per generated report, newest first. Each point carries `buildOrder`, `reportName`, `reportUrl` and a `data` map whose keys are category names and whose values are how many results were assigned that category on **that** build. Only results that were assigned a category contribute — and since every failed or broken result the report counted lands in some bucket, those values effectively partition that build's **failure population**. ## Why the ratio is meaningless Three separate errors are stacked into one division: 1. **Population mismatch.** A category count is a count of failures. `total` counts all results, overwhelmingly passes on a healthy suite. The quotient therefore falls when the suite grows, even if the failures are identical. 2. **Window mismatch.** `widgets/summary.json` is this run. The trend file is a series of past builds. Unless the reader took the point at the head of the array, they divided one build's failures by a different build's test count. 3. **Unit drift.** Both files count *results*, so the phrase "share of tests" is already wrong on both sides of the fraction before the division happens. ## What the numbers do line up with | the number | its honest denominator | |---|---| | a bar on the newest categories-trend point | `failed` plus `broken` in that build's `widgets/summary.json` | | `failed` in `widgets/summary.json` | `total` in the same file | | a `duration` metric on a duration-trend point | nothing — it is an absolute span, not a share | So the defensible sentence is: *on this build, X of its failed-and-broken results were classified as "Timeouts", out of F failed and B broken in total.* Both numbers then come from the same build, and the denominator stays inside the population the numerator was drawn from. ## Checking that two tiles describe the same run Because the trend file is newest-first and every point carries its build identity, lining tiles up is mechanical: - Take the **head** of the trend array. Everything below it belongs to an earlier report. - Compare its `buildOrder` and `reportName` with the report you actually have open. - Only then put its values next to the `widgets/summary.json` from the same generated report. If the head point's build identity does not match the report in front of you — because you are reading an archived report beside a fresh summary, or because the run supplied no executor block to label the point with — then the two tiles are not talking about the same thing and the comparison is void. That check takes seconds and is the difference between a number you can defend in a release review and one that merely looks precise. ## The habit worth carrying Before quoting any overview number as a percentage, write the fraction out in words: *what over what*. If the numerator and the denominator do not name the same population **and** the same run, the percentage is not merely imprecise, it is meaningless. That single check costs a sentence and catches almost every bad number that gets pasted into a release summary — and it is the whole reason it is worth knowing which file backs which tile in the first place.

  • How would you check the two numbers even describe the same build?
    By taking the head of the trend array. The points are written newest first and each carries its own `buildOrder`, `reportName` and `reportUrl`, so only the first entry lines up with the `widgets/summary.json` beside it in the same generated report. Anything further down belongs to an earlier one.
  • So which two numbers actually belong in that fraction?
    The newest categories-trend point's bar, over `failed` plus `broken` from the same build's `widgets/summary.json`. Every failed or broken result the report counted lands in some bucket, so those two counters are the bar's natural denominator, and both numbers then describe one run.

saying these in an interview costs you the question

  • Treats every widget total as the same denominator
  • Assumes passing results appear in the categories tile
  • Reads any trend point as belonging to the current run
  • Calls a count of results a count of distinct tests