An independent test group runs a 340-case regression pack after each two-week freeze, and defects reach authors nine days later. How would you change that?
answer
- Find out where the days actually go
- Segment the interval before moving people
- Split the pack by feedback budget
- Keep the group, change its remit
- Both sides measured on the same outcome
basics
~20 sMeasure where the nine days actually go before reorganising anything, then shorten the loop incrementally: move a fast subset of the pack in front of the merge, embed a tester or two, and keep the group for the work only it can do.
solid answer
~50 sDiagnose first: break the nine days into waiting for the freeze, waiting for an environment, running the pack, triage, and waiting for an author to be free. Each segment has a different fix, and reorganising before you know the split is guesswork. Then shorten the loop in steps. Split the 340 cases: the fast, high-value subset runs on every change before merge and is owned by the delivery team; the slow or cross-team remainder stays with the group. Embed a tester in the team so cases are shaped before code exists rather than executed after it. Give the team an environment it can reach without asking. Keep the group for what genuinely needs independence or specialised skill, and give it an enabling remit. Finally fix the incentives — while one side is measured on defects found and the other on features shipped, the wall regrows.
go deeper
Be ready to explain why a defect found nine days after the code was written costs more than one found the same hour, and why a long pack running only after a freeze delays that news.
Show the mechanics of shortening the loop: splitting a large regression pack into a fast pre-merge subset and a slower remainder, and giving the delivery team an environment it can reach on its own.
Demonstrate diagnosis before reorganisation. Segment the nine days, fix the segment that dominates, sequence the changes so embedding people comes after the build and environment constraints are gone, and prove it on one team first.
Own the incentive and capability argument: why a wall regrows whenever two functions are measured on opposing outcomes, what the independent group should be kept for, and how you protect deep testing skill once testers are scattered across teams.
## Read the symptom before naming the cure Nine days between writing a behaviour and hearing it is wrong is an **over-the-wall** pattern, and the instinct — dissolve the group, embed everyone — is the wrong first move. It discards concentrated skill you will need and answers a question you have not asked yet: *where do the nine days go?* Break the interval into segments and measure each for a few cycles: time from a change merging to the freeze; freeze to a testable build existing; build to an environment being available; environment to the 340 cases finishing; finish to a defect being triaged and assigned; assignment to the author having time to look. In practice these are wildly uneven. If seven of the nine days are waiting for the freeze window, embedding a tester changes almost nothing and cadence is the problem. If four days are environment provisioning, the fix is infrastructure. If the pack itself takes three days because much of it is executed by hand, the fix is in the pack. Reorganising people is expensive and slow; find out first whether people are the constraint. ## Why nine days costs more than nine days A defect that arrives nine days late lands on code the author has already changed, in a mental context they have lost, sometimes on top of two more changes to the same path. Take a quote engine for insurance policies where a retried submission wrote the applicant's quote record twice. Found within an hour, it is a small correction by the person who just wrote the retry path. Found nine days later, after the path has been touched twice more, it is archaeology: which change introduced it, does the later change already mask it, is anything downstream already counting duplicates. The same defect costs several times more to fix at the far end of the loop — the direction of that effect is uncontroversial, though the frequently quoted order-of-magnitude multipliers are contested and should not be recited as fact. ## Changes worth making, in order **Split the pack by feedback budget.** A 340-case pack is not one artefact. Identify the subset that is fast, deterministic and covers what actually breaks — typically a modest fraction — and make it run on every change before merge, owned and fixed by the delivery team. What remains is the slow, the cross-team and the genuinely manual, and it stays with the group on the longer cycle. This single move usually recovers most of the latency without touching the org chart. **Give the team the environment.** If the team cannot get a build into a realistic environment without filing a request, the queue is guaranteed regardless of who owns testing. This is often the cheapest large win. **Embed testers into the delivery teams.** Move one or two testers in permanently, so their work starts before code exists — shaping cases, arguing about the requirement, working alongside the author on the risky path — rather than executing after it. Do this after the previous two moves, not before, because embedding a tester into a team that still cannot build or deploy just relocates the waiting. **Keep the group, change its remit.** Do not dissolve it. It holds the deep exploratory skill, the cross-team scenarios that no single team can construct, and whatever independent pass a regulator or a customer contract requires. Give it an enabling remit as well: tooling and data design the teams cannot build alone, technique coaching, and a peer community for the now-scattered embedded testers, who otherwise become the only tester in the room with nobody to learn from. **Change what each side is measured on.** This is the structural root. While the group's success is defects found and the teams' success is features shipped, both are behaving rationally and the wall regrows however you seat people. Measure both on the delivered outcome — escaped defects and the interval from change to feedback — and the incentive to throw work over disappears. ## Sequencing and honesty about risk Do it incrementally and prove it on one team first. Running the reduced loop alongside the existing freeze for a couple of cycles tells you what the fast subset misses, and that evidence is what convinces the people whose job you are changing. Say the risks plainly: the fast subset will miss things at first, and you need a route for those escapes back into it; skill spread thin can mean nobody is expert, which is what the enabling remit is for; and the group's members are being asked to give up the part of the role they may most identify with, which is a change-management problem, not a logistics one. The measure of success is not the org chart. It is the interval from a change being written to someone competent judging it, and whether it is still falling three months later.
- How would you choose which cases move into the fast pre-merge subset?By what has actually caught defects and what covers paths whose failure is expensive, filtered by speed and determinism. Look at which cases have ever failed for a real reason, add the paths that carry money or data integrity, and exclude anything slow, manual or intermittently unstable. Keep the subset small enough to fit the feedback budget the team will tolerate, and let it grow only when an escape proves it should.
- The test group's lead argues that embedding testers will destroy their craft. Is that a fair objection?Partly, and it should be answered rather than dismissed. A tester who becomes the only one in a room does lose the peer group that developed their technique. That is precisely why the group survives with an enabling remit: a community of practice, shared tooling, technique sessions and the cross-team work that needs concentrated skill. Embedding without that support is the version of this change that genuinely does erode craft.
- After three months the interval has fallen but escaped defects have risen. What do you look at?Whether the fast subset covers what the long pack was actually catching. Take the escapes and ask, for each, whether a case existed in the 340 that would have caught it and did not move across, whether it needs a new case, or whether it was never covered at all. That analysis rebuilds the subset from evidence. Also check that the remainder is still being run rather than quietly abandoned.
saying these in an interview costs you the question
- Dissolves the test group before measuring anything
- Assumes automating all 340 cases fixes the latency
- Embeds testers while the team still cannot get an environment
- Leaves both sides measured on opposing outcomes
- Treats the nine days as one undifferentiated delay
- Promises the fast subset will miss nothing