Allure 2's `widgets/summary.json` reports a run's time as both a `duration` and a `sumDuration`. What is each one computed from, and which of the two does the report's duration-trend tile plot?
answer
- one wall clock, one stopwatch total
- stop minus start, across the whole run
- sumDuration adds each result's own time
- the trend plots the wall clock
basics
~20 sIn Allure 2, duration is the run's wall-clock span, the latest stop minus the earliest start. sumDuration adds up each counted result's own duration. The duration-trend tile plots the wall-clock span, not the summed test time.
solid answer
~40 sBoth live in the group-time block of Allure 2's `widgets/summary.json`, and they measure different things. `start` and `stop` are the extremes across every counted result, and `duration` is simply `stop` minus `start` — the wall-clock span of the run. `sumDuration` is the sum of each result's own duration, with `minDuration` and `maxDuration` as the extremes of that same population. On a serial suite the two roughly agree; on a parallel one `sumDuration` can be many times `duration`, because overlapping tests all count toward the total but share the span. The duration trend, `widgets/duration-trend.json`, stores one metric per point keyed `duration`, and it is the wall-clock one — so doubling your runners makes that trend fall without a single test getting faster.
code
json · 10 lines{
"time": {
"start": 1757260800000,
"stop": 1757260980000,
"duration": 180000,
"minDuration": 12,
"maxDuration": 41230,
"sumDuration": 1412884
}
}go deeper
Know that the run's time block carries two different totals and that they are not interchangeable: one is the elapsed span of the run, the other adds up the individual tests. Say which you mean whenever you quote a duration.
Explain the arithmetic: start and stop are extremes over results, duration is their difference, and sumDuration is a running sum of per-result durations. Be able to predict how each one moves when the suite is sharded.
Demonstrate that you would not accept a falling duration trend as a performance win. Diagnose the ratio between the two numbers, spot a mis-stamped result stretching the span, and separate waiting time from test time.
Decide which of the two your organisation reports on, and defend it. Wall-clock span is what engineers wait for; summed test time is what the run costs to execute. Publishing one without the other invites optimising the cheaper number.
## Two different clocks in one block The `time` block of Allure 2's `widgets/summary.json` is a group-time model built by folding every counted result into one object. Six fields come out of that fold, and they are not measuring the same thing: | field | how it is produced | |---|---| | `start` | the earliest `start` of any counted result | | `stop` | the latest `stop` of any counted result | | `duration` | `stop` minus `start` | | `minDuration` | the shortest single result's own duration | | `maxDuration` | the longest single result's own duration | | `sumDuration` | every counted result's own duration, added up | So `duration` is a **span**: the wall-clock distance from the moment the earliest test started to the moment the latest one stopped. `sumDuration` is a **total of work**: the machine time actually spent inside tests. They answer different questions, and the file prints both without saying which is which. ## Why the two numbers diverge On a strictly serial suite the two are close. Every test runs after the last one finished, so the span is roughly the sum of the parts plus whatever gaps sat between them. The moment the suite runs in parallel they come apart: - With eight workers, eight results overlap in time. Their durations all still land in `sumDuration`, but they consume the same slice of the span, so `duration` grows by roughly an eighth of what `sumDuration` grows by. - **`sumDuration` can therefore be many times `duration`.** That is not double counting; it is the definition of the two fields. - They also come apart in the other direction. Because `start` and `stop` are extremes over results, any dead time inside the window — waiting on an environment, a long fixture between batches, a worker sitting idle — falls inside `duration` and inside none of the individual durations. A `duration` well above `sumDuration` on a serial suite says most of the wall clock is going on something that is not a test. - One result with a bad clock poisons only the span. A `stop` timestamp far in the future stretches `duration` for the whole run while `sumDuration` barely moves, which is a quick way to tell a genuinely slow run from a mis-stamped one. ## Which one the trend plots The duration trend is a separate file, `widgets/duration-trend.json`. Each point in it carries `buildOrder`, `reportName`, `reportUrl` and a `data` map of metric name to value — and for this trend the map holds exactly one entry, keyed `duration`, set from the same group-time `duration`. **The trend is a wall-clock trend.** That single fact changes how the chart can honestly be read: 1. **A falling trend is not evidence that tests got faster.** Add runners, shard the suite, or move to a bigger machine and the span drops while the work does not. 2. **A rising trend is not evidence that tests got slower.** Lose a runner, or serialise a stage that used to overlap, and the span rises with no change to any test. 3. **The machine-time bill is the other number.** If you care what the suite costs to run rather than how long you wait for it, `sumDuration` in `widgets/summary.json` is the field to compare across builds — and it is not the one the trend tile plots. ## Reading and quoting the pair The honest sentence names which clock it used. *"The suite takes eleven minutes of wall clock, on about ninety minutes of machine time across eight workers"* is a claim you can defend from these two fields. *"The suite takes ninety minutes"* is not, because it silently picks the number that suits the argument. A practical habit, whenever a duration is about to be quoted anywhere: pull both fields out of `widgets/summary.json` for the same build and look at the ratio. - A ratio near one says the suite is effectively serial, and the two numbers can be used almost interchangeably. - A large ratio says parallelism is doing the work, and the two must never be mixed in one sentence. - A ratio below one says the wall clock contains something that is not a test at all, and that is usually the more interesting finding. `minDuration` and `maxDuration` finish the picture, because they bound the tail of the same population `sumDuration` was summed over. A `maxDuration` close to `duration` means one test is holding the whole run open, and no amount of extra parallelism will shorten it — the critical path is a single test. A `maxDuration` far below `duration` means the span is being spent between the tests rather than inside them, which points at scheduling, queuing or setup rather than at any test you could optimise.
- The duration trend halved overnight and nobody touched a test. What do you check?Whether the run's parallelism changed. The trend plots `stop` minus `start` for the whole run, so more workers shorten it while `sumDuration` — the summed per-test time — stays flat. Compare both fields in `widgets/summary.json` across the two builds before telling anyone the suite got faster.
- What do `minDuration` and `maxDuration` tell you that `duration` cannot?They bound the tail. A `maxDuration` close to `duration` says one test is holding the whole run open, so extra parallelism will not help. A `maxDuration` far below it says most of the wall clock is going on scheduling, queuing or fixtures that sit outside any single test.
Think of a kitchen. The duration is the clock on the wall, from the first pan on to the last plate out; the sumDuration is what you get by adding up every cook's own stopwatch. Put four cooks in the kitchen and the wall clock drops while the stopwatch total does not move.
saying these in an interview costs you the question
- Reads duration as the total time tests spent running
- Adds each suite's duration to get the run's duration
- Says a falling duration trend proves tests got faster
- Assumes duration and sumDuration agree on any run