Why narrow a defect to the shortest reliable reproduction before you report it?
answer
- Fewer conditions, faster diagnosis
- A claim, not a diary
- Delete one condition, then re-run
- Restore anything the failure needed
- Start state and hit rate count
basics
~20 sA short reproduction proves which conditions actually cause the failure. Every step you can delete without losing the failure is noise that slows diagnosis, invites a cannot-reproduce close, and hides the real trigger from whoever has to fix it.
solid answer
~50 sA reproduction is a claim about which conditions are sufficient to make the failure appear, not a diary of what you happened to do. So I reproduce it at least twice from a known start state, then delete one step or condition at a time and re-run; if the failure survives the deletion I keep the shorter version, if it disappears I put the condition back because it is part of the trigger. The start state counts as a condition - the build, the account's role, the seeded data, the configuration - so I record it rather than assume it. I also record how reliably it fires, for example nine attempts in ten, instead of writing 'always'. The result is evidence someone can act on in minutes rather than a walkthrough they have to re-derive.
code
pseudocode · 18 linessteps = recorded_walkthrough() # 41 recorded actions
start = known_state(build, account_role, seeded_data, settings)
for candidate in copy_of(steps):
trial = steps_without(candidate)
hits = 0
repeat 5 times:
reset_to(start)
if run(trial) shows_the_failure: hits = hits + 1
if hits >= 4:
steps = trial # condition was incidental, drop it
else:
keep(candidate) # condition is part of the trigger
report(start_state = start,
conditions = steps,
hit_rate = measure(steps, attempts = 20))go deeper
Be ready to say what a reproduction is for: proving which conditions cause the failure. Recall the loop - reset, reproduce twice, remove one condition, re-run, restore anything the failure needed - and remember to write down the build, the account and the data.
Explain the mechanics, especially the asymmetry: a deletion is only kept if the failure survives it. Be able to name the invisible conditions - identity, seeded data, configuration, leftover cache state, ordering - and to express reliability as a count of attempts rather than as 'sometimes'.
Show that you read the surviving conditions as evidence, not just as instructions, and that you know when minimising destroys the bug because the sequence itself is the trigger. Interviewers expect you to describe the cost side too: how long you spend narrowing before filing what you have.
Own the economics. Decide who is cheapest to do the isolation, what minimum evidence you require from everyone so that reports do not ping-pong, and how you keep that bar from quietly suppressing reports of the messy defects nobody can narrow.
### A reproduction is a claim, not a diary Most weak defect reports are transcripts: the tester writes down everything they did between logging in and noticing something wrong. That is a record of a session, not a reproduction. A reproduction is a much stronger claim - *these conditions, from this starting point, are enough to make the failure appear* - and the work of narrowing is the work of turning the first thing into the second. The claim matters because the reader is not you. They did not see the screen, they do not know which of your forty-one actions was load-bearing, and they will spend their first twenty minutes re-deriving what you already know. Every incidental step you leave in is a step they have to perform, doubt, and eventually eliminate themselves. ### The minimisation loop The loop is mechanical and worth doing by hand at least once until it becomes instinct: 1. **Get a known start state.** A fresh account or a reset data set, a named build, a recorded configuration. Without this, nothing after it is repeatable. 2. **Reproduce it twice.** One observation is an anecdote. Two from the same start state is a reproduction. 3. **Delete one condition and re-run.** A step, a field value, a role, a piece of seeded data, an option. 4. **Keep the deletion only if the failure survives it.** If the failure disappears, restore the condition - you have just learned it is part of the trigger, which is a finding in itself. 5. **Stop when every remaining condition is load-bearing**, then measure how often the shortened version fires. The important asymmetry is in step 4. Deleting steps until the failure stops appearing and then filing the shortened version anyway is the single most common way a reproduction becomes untrustworthy - it produces steps that genuinely do not reproduce anything, and the report comes back closed as not reproducible. ### Conditions that are not in the step list Juniors usually minimise the clicks and forget everything around them. The conditions that most often turn out to be essential are invisible in a step list: - **Identity and permissions** - which role, which account, which team the account belongs to. - **Data** - not 'some trips' but the specific shape: a record with an empty optional field, a very large history, a duplicate. - **Build and configuration** - the exact build identifier, feature switches, regional or unit settings. - **Leftover state** - a cache warmed by the previous test, a background job that ran between your two attempts. - **Timing and ordering** - two actions inside the same second, a page left open while something else changed underneath it. If you cannot say which of these you controlled, you have not finished narrowing; you have only shortened the visible part. ### Reliability belongs in the reproduction 'Always' and 'sometimes' are both worthless to the person fixing it. A count is not: *fires on 9 of 10 attempts on this build, with this account and this data*. That single line tells the reader whether to expect it on their first try, and it tells you honestly whether you have finished isolating. A hit rate well below one is a signal that a condition you have not identified is still varying between runs. ### When a long reproduction is the right answer Sometimes the length is the bug. Accumulated state, a specific ordering, a slow leak, a scheduled job - these need a sequence, and cutting it destroys the evidence. The discipline still applies: minimise everything you can, then state plainly that the remaining sequence is essential and say what you tried to remove and could not. A reader trusts 'I removed twelve steps and these six are required' far more than a bare six-step list with no history. ### A worked example A tester on a ride-hailing dispatcher records a forty-one-action walkthrough that ends with a completed trip showing the wrong fare. Narrowing it takes about twenty minutes: the map interactions, the driver chat, the two cancelled requests and the profile visit all come out with the failure intact. What does not come out is the seeded trip whose promotion code was applied after the trip started, the dispatcher account being in the regional supervisor role, and accepting the trip while a second request is already open. Six conditions remain, from a reset account, firing on 9 of 10 attempts. The report is now a twelve-line note that a developer can reproduce before their coffee cools - and the three surviving conditions are themselves the strongest hint about where the defect lives.
- How many times do you reproduce a defect before you call the reproduction reliable?At least twice from a reset start state before I believe it at all, and then a fixed batch - twenty attempts is a reasonable default - to get a hit rate. I report the count, not an adverb: 'fires on 18 of 20 attempts' or 'fires on 3 of 47'. A hit rate far below one is my own signal that some condition is still varying between runs and the narrowing is not finished.
- When is a long reproduction sequence the right thing to report?When the length is the defect: accumulated data, a required ordering, a background job that has to run between two actions, or a resource that only exhausts after many cycles. I still minimise everything that can go, then say explicitly that the remaining sequence is essential and list what I removed. That tells the reader the length was tested rather than merely copied out of my notes.
- You cut a reproduction from forty-one steps to six. What do the six surviving conditions tell you beyond how to trigger it?They are the best cheap hypothesis about where the defect lives. If the survivors are a particular role, a promotion applied mid-trip and a second request open at the same time, the defect is almost certainly in how those three interact, not in the map or the chat. I put that reading in the report as a hypothesis, clearly labelled as one, so the developer can accept or discard it without having to re-run my narrowing.
Narrowing a reproduction is like isolating an ingredient after a bad meal: you do not re-eat the whole menu, you drop one dish at a time and see whether you still get sick.
saying these in an interview costs you the question
- Pastes the whole recorded walkthrough unchanged
- Writes 'it does not work' with no start state
- Claims it always fails after a single observation
- Deletes steps without re-running to confirm
- Keeps deleting until the failure disappears, then files that
- Treats build, role and seeded data as outside the reproduction