A defect arrival curve flattens two weeks before release. How do you tell stability from exhausted testing?
answer
- One line, several explanations
- What was the effort behind it
- Divide arrivals by something
- Try what the pack never runs
- Watch the age of open reports
basics
~20 sNormalise arrivals by test effort before believing the curve, then probe with something the suite has never exercised. A flat curve over falling effort, or over a suite that has already found everything it can, is not evidence of stability.
solid answer
~50 sA flat arrival curve has at least six explanations and only one of them means ship: genuine stability, effort that collapsed, an exhausted regression pack re-running the same paths, a blocking defect masking a whole region, reporters discouraged by rejections and deferrals, or a change in what gets counted. The two cheap discriminators are normalisation and novelty. Divide arrivals by test hours, cases executed or timeboxed sessions — a flat raw curve over falling effort is a rising rate. Then spend a short timebox on scenarios the pack has never run: a fresh exploratory charter, a different data shape, an unusual load profile. Read the closure curve and the age of the open backlog alongside it, because equal arrival and closure rates with a rising median age means old reports are ageing while new ones get cleared.
code
pseudocode · 8 linesfor week in release_weeks:
arrivals[week] = count(reports opened during week)
closures[week] = count(reports resolved during week)
backlog[week] = backlog[week - 1] + arrivals[week] - closures[week]
median_age[week] = median(week_end - opened_at for r in open_reports)
yield_rate[week] = arrivals[week] / test_hours[week]
report(yield_rate, backlog, median_age) // never arrivals alonego deeper
Know what the two curves are: reports opened per period and reports resolved per period, with the gap accumulating as the open backlog. Recall that a flat arrival line can mean nobody is testing.
Explain the normalisation. Be able to say that arrivals divided by test hours or cases executed is the reportable series, and that a flat raw curve over falling effort is actually a rising yield.
List the competing explanations before interpreting, and propose the discriminating checks: normalisation, a timeboxed novel probe, backlog ageing, the rejection rate. This is a release-readiness conversation, so show the judgement rather than the arithmetic.
Own what the organisation is allowed to conclude from these curves. Set the reporting standard — normalised yield, backlog age, and a novelty probe before any go decision — so that no release is signed on the shape of a single unnormalised line.
## The two curves and the third series **Arrival** is the count of new defect reports opened per period. **Closure** is the count moved to a terminal resolved state per period. The running difference is the **open backlog**. The series most teams forget is **backlog age** — the median age of open reports, or simply how many have been open longer than 30 days. The textbook release-readiness reading: arrivals rise, peak and decline; closures track behind; the backlog peaks and then drains; you ship when arrivals are low, the backlog is small, and it is not ageing. That reading silently assumes test effort is constant across the window. It almost never is, and that assumption is where the interview question lives. ## Six explanations for a flat arrival curve 1. **Genuine stability.** The remaining defects are getting harder to find because there are fewer of them. 2. **Effort collapsed.** Testers were reassigned, a holiday landed, the environment was down, or a broken build ate three days. 3. **Coverage exhaustion.** The same regression pack is being re-run. It has already found everything it is capable of finding, and re-running it produces a flat line forever. 4. **A blocking defect is masking a region.** One unfixed fault stops a whole flow being exercised, so nothing behind it can be reported. 5. **Reporting has been discouraged.** A run of rejections, deferrals or known-issue closures teaches people that filing is not worth the argument. 6. **The bookkeeping changed.** A new intake form, a merged project, or a quiet change in what counts as a defect. Only the first of those means ship. ## Cross-checks that separate them **Normalise by effort.** Arrivals per test hour, per case executed, or per timeboxed exploratory session. A flat raw curve over falling effort is a *rising* rate. **Probe with novelty.** Run something the pack has never exercised: a fresh exploratory charter, a different data shape, an unusual environment configuration, a load profile you have not used. If a short novel probe yields new defects at a healthy rate, the flat curve was exhaustion, not stability. **Read the closure curve and the backlog age together.** Equal arrival and closure rates with a rising median age means recent reports are being cleared while old ones sit — the totals look balanced and the backlog is quietly rotting. **Look at severity mix and location.** Late defects clustering in one area often means that area was only just reached, which is an argument against stability rather than for it. **Audit the intake path.** Count rejections and reclassifications over the same weeks. A rising rejection rate is a good explanation for a falling arrival rate that has nothing to do with the product. ## A worked example A subscription renewal job. Weekly arrivals across six weeks: 61, 78, 74, 52, 33, 31. The curve looks like a textbook decline and the release manager wants to ship. Executed test hours over the same weeks: 214, 220, 216, 141, 96, 88. Arrivals per test hour: 0.29, 0.35, 0.34, 0.37, 0.34, 0.35 — flat, not falling. The curve came down because two testers moved to another release, and the product's defect yield per hour of testing never changed. The team spent two days on a novel probe instead: running the renewal job at a 1,200-request-per-minute peak rather than the usual steady trickle. A currency-rounding drift surfaced — fractional cents accumulating across concurrently prorated renewals — plus two further reports behind it. The regression pack had never run that job at peak, so the shape of the curve had been describing the pack's reach, not the product's health. ## What to say when asked Say that the curve is a prompt, not a verdict. Name the confounders before offering an interpretation. Propose the normalisation and the novelty probe as the two cheap checks that discriminate between them. Bring backlog age in as the honest test of whether closure is real or cosmetic. And decline to sign a release decision on the shape of one line — the defensible statement is "arrivals per test hour have been flat for three weeks and a novel probe found nothing new", which is a very different claim from "the curve is going down".
- Arrivals and closures are running at the same rate, but the median age of the open backlog keeps rising. What is happening?Recent reports are being closed while older ones sit untouched, so the totals balance and the backlog quietly rots. It usually means triage is picking the cheap items, or that the old reports are blocked on a decision nobody is making — a design question, an environment, an owner who left. I would report the count over 30 days old rather than the median, name the blockers explicitly, and stop treating a stable backlog size as a healthy signal.
- How would you design a probe that distinguishes an exhausted suite from a genuinely stable product?Timebox it, and make it deliberately novel rather than more of the same: an exploratory charter over an area the pack does not cover, a fresh data set with different shapes and boundaries, an unusual environment or locale configuration, or a load profile the suite never applies. Fix the timebox in advance, record what was explored, and compare the yield per hour against the recent regression yield. A short probe that finds new defects at a healthy rate settles the question.
- What would make you argue that a falling arrival curve is bad news rather than good?If the fall coincides with something other than the product improving: test effort dropping, a spike in rejected or deferred reports, an environment outage, or a blocking defect that stops a major flow being exercised. The last one is the most dangerous, because everything behind the block is untested and invisible at once. In those cases the curve is describing our own capacity to look, and I would say so before anyone reads it as readiness.
saying these in an interview costs you the question
- Reads a flattening arrival curve straight through as readiness
- Never asks what the test effort was behind the curve
- Ignores the closure curve and the backlog age entirely
- Assumes a shrinking backlog means defects were actually fixed
- Never considers that a blocking defect is masking a region
- Treats the curve as a verdict rather than a prompt