As a lead, how much investigation do you require before a defect is filed?
answer
- Investigation cost moves, never disappears
- Ask who can produce it cheapest
- Cheap floor, timeboxed ceiling
- Silent wrong results override the budget
- Watch for defects quietly not filed
basics
~20 sSet a cheap, fixed evidence floor plus a timebox rather than an open-ended demand: capture build, start state, artefacts and a hit rate, isolate for an agreed period, then file with what you have and say where you stopped.
solid answer
~50 sI treat it as a cost-allocation decision, not a standards question. Investigation is real work that lands on whoever does it, so the question is who can produce the missing evidence most cheaply. A tester can always supply the floor - build identifier, start state, artefacts from failing runs, and a hit rate - cheap enough to require universally. Deep isolation is different: the person with the code and traces is usually faster, so I timebox the tester's attempt rather than demand a root cause. Two carve-outs override the timebox: anything suggesting a silent wrong result gets investigated immediately whatever the budget, and anything near a deadline gets filed early with partial evidence, because on a three-week release train a thin report on day 4 beats a thorough one on day 18. Then I watch the signals: round trips per report, cannot-reproduce closes, and whether marginal defects have quietly stopped being filed.
go deeper
Know what you are always expected to supply: the build, the start state, the artefacts from the failing attempt, what you expected, and how many attempts failed. If you cannot isolate further, file with what you have and say where you stopped.
Explain why the floor and the isolation timebox are separate policies, and what each buys - the floor removes round trips, the timebox bounds an open-ended cost. Be ready to say what you write when the timebox expires with the trigger still unknown.
Show that you decide case by case who can produce the missing evidence most cheaply, and that you override the budget for suspected silent wrong results or a near deadline. Interviewers expect you to justify handing an isolation over rather than grinding on it.
Own the whole incentive system: the cost allocation between testers and developers, the measurable signals including reports that quietly stop being filed, and the tooling investment - capture, resettable data, identifiable builds - that makes a cheap floor enforceable at all.
### The real question is who pays Every defect needs a certain amount of investigation before it can be fixed. That work does not disappear when you set a policy - it moves. A high bar before filing moves it onto testers, who have the failing environment and the observations but not the code. A low bar moves it onto developers, who have the code and the traces but must first re-create conditions someone else already had. Neither is free, and a lead who frames this as 'testers should file better reports' has skipped the only interesting part of the decision. ### The floor and the ceiling The policy that survives contact with reality has two separate parts, and conflating them is the usual mistake. **A cheap, non-negotiable floor.** Evidence that costs the reporter almost nothing because they already had it: the build identifier, the start state (account, role, data, configuration), the artefacts from the failing attempt, what was expected and why, and a hit rate as a count of attempts. This is a minutes-long ask, it eliminates most round trips on its own, and requiring it universally is fair because everyone can meet it. **A timeboxed ceiling on isolation.** Narrowing conditions and chasing a hidden variable is open-ended, and its cost is unpredictable in a way the floor's is not. So it gets a clock - twenty to forty minutes is a common shape - after which the reporter files what they have, states where they stopped, and lists the hypotheses they eliminated. What you must not write is 'investigate until you find the cause', which silently converts an unbounded cost into an unbounded delay. ### Overrides that beat the budget Two situations justify ignoring the ceiling. **Suspected silent wrong results.** When a symptom suggests a value may be persisted, exported or aggregated wrongly with nothing in the product ever comparing it, mildness of the symptom is not evidence of small damage - it is evidence of poor detection. That gets investigated immediately, because every day it stays unfiled adds records that may be unrepairable. **Deadline proximity.** On a three-week release train, a thin report on day 4 can still be fixed and re-tested; a thorough one on day 18 usually cannot. Near a cut-off the correct instruction is to file early and continue investigating on the open report, which is exactly the opposite of the instinct to polish first. ### Signals worth watching Any policy here can be measured, and a lead who does not measure it is guessing: - **Round trips per report** - how often a report bounces for missing information. Rising means the floor is too low or unclear. - **Cannot-reproduce closes** - usually an environment, data or artefact problem rather than a diligence problem. - **Time from filing to first diagnosis** - what the floor is actually buying. - **The marginal-defect count** - the one everybody forgets. If odd, hard-to-narrow observations have stopped being filed, the bar is suppressing the reports you most need, and this shows up as a quiet absence rather than a complaint. ### Perverse incentives to design against A strict bar makes filing expensive, so reporters stop filing anything they cannot make tidy - and the untidy ones are disproportionately the serious ones. A bar tied to individual metrics is worse: it teaches people to file only what they can already explain. Conversely, no bar at all floods the queue and trains developers to close reports unread. The stable design is a floor that is cheap and universal, a ceiling that is bounded and explicit, an easy path for 'here is something odd I could not pin down' that is welcomed rather than penalised, and pairing on hard isolations so the knowledge moves instead of the ticket. ### Make the environment carry some of the load Much of what is argued about as diligence is really tooling. If failing runs do not leave artefacts, if test data is not resettable, if builds are not identifiable, then even a diligent reporter cannot meet a reasonable floor, and the policy becomes a demand for something the system does not permit. Cheap continuous capture, reset-on-demand data and per-attempt build identifiers usually buy more report quality than any exhortation, and they also make the floor genuinely cheap - which is the condition that lets you enforce it without argument.
- How would you notice that your evidence bar has started suppressing reports?By its silence, which is why it needs deliberate checking. I compare the rate of odd, hard-to-narrow observations being filed before and after, ask directly in debriefs what people saw but did not file, and watch whether defects are surfacing late in a cycle that someone clearly noticed early. A bar is doing damage when the reports that vanish are the messy ones rather than the trivial ones.
- When is a developer better placed to do the isolation than the tester who found it?When the remaining question is about internal state rather than external conditions - which branch was taken, what the code did between two operations, why an ordering matters. The tester's advantage is the reproducing environment and the observations; the developer's is the traces and the code. Once the cheap floor is met and the tester's timebox is spent, moving it is usually cheaper than continuing, and pairing on the handover keeps both the context and the learning.
- How does a short release train change what you ask reporters to do?It inverts the polish-first instinct. On a three-week train, work that lands after roughly day 16 will not be fixed and re-tested in that cycle, so near the cut-off I ask for filing first and investigating afterwards on the open report. Earlier in the train the timebox stands, because there is room for isolation to pay for itself in reduced round trips.
It is triage staffing at an emergency door: you insist on the two-minute observations from everyone, cap how long the first responder investigates, and keep the door welcoming enough that people still come in with the vague symptoms.
saying these in an interview costs you the question
- Demands a root cause before any report is filed
- Says testers should simply write better reports
- Sets an unbounded 'investigate until you find it' rule
- Ignores that the bar suppresses messy reports
- Measures reporters individually on report quality
- Blames diligence when artefacts are impossible to capture