skip to content

Allure 2 backs its retry trend with a `retry-trend.json` file. What are the two numbers it stores for each build, and what is each one actually counting?

level: middleimportance: nice to knowfreq 28%

answer

  1. one trend point per build
  2. two counters under one map
  3. attempts, not tests
  4. one counter is the superseded ones

basics

~10 s

Each point in retry-trend.json holds a data map with two counters: retry, the number of attempts that were superseded in that build, and run, the number that stayed visible. Both count attempts, not tests.

solid answer

~50 s

A point in `retry-trend.json` is a trend item: a `buildOrder`, optionally a `reportName` and `reportUrl`, and a `data` map. For the retry trend that map holds exactly two keys, `retry` and `run`. The plugin that fills them runs **after** `RetryPlugin`, so the retry flag on each result has already been set; it then walks the **unfiltered** results and increments `retry` for every result flagged as a superseded attempt and `run` for every one that is not. Both counters therefore measure attempts. Nothing in the file names a test, and a single badly behaved test retried many times is indistinguishable from many tests retried once. The file is written into the report's history data and again as widget data, and the current build's point is put in front of the points that came in from the previous report.

code

json · 18 lines
json
[
  {
    "buildOrder": 12,
    "reportName": "Nightly regression",
    "data": {
      "retry": 4,
      "run": 13
    }
  },
  {
    "buildOrder": 11,
    "reportName": "Nightly regression",
    "data": {
      "retry": 2,
      "run": 13
    }
  }
]

go deeper

for a junior

Recall that Allure 2's retry trend is backed by a file holding one point per build, and that each point carries two whole-number counters under a data map rather than a percentage.

for a middle

Explain that the two counters split on the retry flag the consolidation step has already set, and that the plugin has to take the unfiltered results because superseded attempts are invisible in the filtered set.

for a senior

Be ready to explain a spike without reaching for reliability. A larger suite or a raised retry allowance moves the line on its own, and the file names no tests, so it cannot answer which ones were involved.

for a principal

Own how much weight a two-counter-per-build artefact can carry, and say what a consolidated report has to keep beside it so a reader can get from a moved line back to the underlying results.

## The shape of the file `retry-trend.json` is a JSON array of trend points, newest first. Each point is the generic trend item shape Allure 2 uses for all its trend files: - `buildOrder` — the executor's build number for that report, when one was available. - `reportName` and `reportUrl` — where that point came from, so the rendered trend can link back. - `data` — a map of metric name to a whole-number count. What makes it the *retry* trend is the contents of `data`. The retry trend item initialises the map with exactly two keys and never adds a third: | key | what it counts | |---|---| | `retry` | attempts that were superseded by a later attempt in that build | | `run` | results that stayed visible in that build | ## How the two counters get their values This only makes sense in the light of what runs before it. The plugin that consolidates repeated attempts runs first: it groups the results, keeps the newest attempt of each group, and sets a boolean `retry` flag on every attempt it demotes. The trend plugin is registered immediately after it, so by the time the trend is computed that flag is already correct on every result. The trend plugin then does something the tree and summary plugins do not: it takes the **unfiltered** set of results, including the demoted ones, and folds each into the current point. The fold is a single branch — if the result carries the retry flag, increment `retry`; otherwise increment `run`. That is the whole computation. That choice of accessor is not incidental. Built on the filtered set, every result would look like a survivor and the `retry` counter would read zero on every build ever generated. Counting attempts is only possible from the population that still contains them. ## What the numbers do and do not mean Both counters are **counts of attempts**, and reading them as anything else is where this file gets misused. 1. **`retry` is not a count of tests.** One test retried several times contributes several. Many tests retried once contribute the same total. The file cannot tell those two runs apart, because nothing in it names a test. 2. **`run` is not the total number of executions.** It is the complement of `retry` within the same build, so the executions the runner performed are the two added together, not `run` alone. 3. **Neither counter looks at status.** A superseded attempt increments `retry` whether or not the outcome changed on the next attempt, and whether or not the visible attempt passed. 4. **A moving line has several innocent explanations.** The suite grew. The runner's retry allowance was raised. A slow environment pushed more tests into repetition for a day. The trend records how much repeated execution the build absorbed and nothing about why. Stated positively: this is a **workload** trend. It answers "how much of this build's execution was repetition?", which is a useful thing to watch precisely because it is cheap to compute and hard to argue with. It is not, and cannot be, a per-test signal. ## Where the older points come from The plugin does not recompute history. Older results are not present when a report is generated, so there is nothing to recount. Instead it reads the trend points that arrived with the parsed launch, computes one fresh point for the current build, puts that point at the front, and keeps the newest stretch of the list. Each generation therefore extends the series by exactly one entry, and every earlier entry is a value some previous generation computed and passed forward. One practical consequence: a report generated in a workspace that carried nothing forward produces a `retry-trend.json` with a single point in it. The line has not been reset by anything in the retry machinery — the plugin simply had no earlier points handed to it. The file is written twice per generation, once into the report's history data, which is the copy a later run can inherit, and once as widget data, which is the copy the rendered page reads. They hold the same list; they differ only in where the report keeps them.

  • Why can this trend jump sharply without any test having become less reliable?
    Because it counts attempts. A bigger suite, or a raised retry allowance on the runner, moves both counters on its own. The file also names no tests, so one test repeated many times and many tests repeated once produce the same point.
  • Where does the plugin get the older points from?
    From the trend data that arrived with the parsed launch. It computes one point for the current build, puts it at the front of whatever previous points came in, and keeps the newest stretch. It never recomputes history, because the earlier runs' results are not there to recount.

saying these in an interview costs you the question

  • Reads the retry counter as a count of flaky tests
  • Thinks the run counter covers every attempt the runner made
  • Assumes the trend is computed from the surviving results alone
  • Expects the file to name which tests were repeated