skip to content

When does a written, scripted test procedure still beat exploratory testing?

level: seniorimportance: should knowfreq 58%

answer

  1. One property decides it: sameness
  2. Re-run, prove, hand off, reproduce
  3. External reviewer was not in the room
  4. Comparable results across many configurations
  5. Scripts re-find only what their author knew

basics

~20 s

A written procedure wins whenever sameness matters more than discovery: re-running the identical check every build, producing an evidence trail someone outside the team can audit, handing work to a person without product knowledge, and confirming a known defect the same way twice.

solid answer

~50 s

Scripts buy repeatability, and repeatability is what exploration cannot give. Four situations turn that into a decision. First, regression: a check that must produce the same verdict on every build needs to be written down, or it drifts with whoever runs it. Second, evidence: where an external reviewer must see what was executed and what the outcome was, the procedure and its result are the artefact, and "a skilled tester explored for an afternoon" does not substitute. Third, hand-off: a written procedure lets someone without product knowledge run it, and lets many people run the same thing across many configurations and compare results. Fourth, confirmation and reproduction: once a defect is understood, the shortest reliable steps are worth writing so the fix can be verified identically. Everywhere else, a script mostly re-finds what its author already knew.

go deeper

for a junior

Be ready to name two situations where writing the steps down matters: checking the same thing on every build, and letting someone else reproduce a defect exactly as you did.

for a middle

Explain the mechanism rather than listing situations. A written procedure guarantees two runs are the same run, which is what makes results comparable across builds, people and configurations.

for a senior

Show you decide per piece of work. Say which parts of a release need sameness and which need discovery, and volunteer the costs of scripts — narrowed attention, maintenance, and diminishing information from a suite that has passed for years.

for a principal

Own the policy question: which contexts in your organisation genuinely require a documented evidence trail, and how you stop that requirement from quietly expanding until every hour is spent executing procedures nobody learns from.

## The one thing a script has A scripted procedure is a set of steps and an expected outcome, fixed in advance. Everything it is good at follows from one property: two runs are the same run. Exploration cannot offer that, because the whole mechanism is that each result changes the next move — two afternoons by two testers on the same mission produce different paths and different findings, which is a feature when you want discovery and a liability when you want comparison. So the question is never "which is better". It is "does this piece of work need sameness or discovery?" ## Four places sameness wins **Re-execution against change.** A check whose job is to say the same thing on every build has to be written down. Undocumented, it drifts: the person who runs it this month interprets a step differently, skips a precondition that used to be obvious, or checks a slightly different thing, and the signal stops meaning anything. Note that this is an argument for *writing the check down*, not for who or what executes it — whether that written check should be automated is a separate decision with its own criteria. **An external evidence trail.** In a regulated or contractual setting, someone outside the team must later be able to see what was executed, against which build, with what result and by whom. A procedure plus its recorded outcome is that artefact. Exploration produces notes and a hand-back, which are honest but not comparable across time, and cannot be reviewed by someone who was not there. If a customer or a regulator has to be shown a specific behaviour was verified before release, write it down. **Hand-off across people.** A procedure encodes product knowledge so that someone who lacks it can still execute. That matters when work moves to a different team or time zone, when a build must be sanity-checked by whoever is around, or when the same behaviour must be checked across many configurations — devices, locales, deployment shapes — and the results must be comparable. Twenty people exploring twenty configurations produce twenty stories; twenty people running the same procedure produce a comparison. **Confirmation and reproduction.** Once a defect is understood, the shortest reliable steps become a written artefact so the fix can be verified the same way and so the check can rejoin the regular set. Exploration finds it; the script is how you prove it stays fixed. ## A worked case On a grant-application review queue, exploration found that interrupting a bulk decision produced a partial-failure rollback: 11 of 37 applications came back marked withdrawn while their detail view still said pending. Nothing scripted would have found it — the test only existed because the tester had just learned how submission wrote to three places at once. But the moment it was understood, the value flipped. The team wrote down the exact steps: seed 37 applications in a known state, start the bulk action, sever the connection at the second write, then assert the counts in the list, the detail view and the audit record. That written check is now what confirms the fix, what runs against every later build, and what a reviewer can be shown. Discovery produced it; sameness preserves it. ## What the script costs Be honest about the other side, because interviewers listen for it. A script only finds what its author thought of when they wrote it, and it keeps finding exactly that — so a suite that has passed for many releases is producing less and less new information, a phenomenon usually described as the pesticide paradox. Scripts also carry maintenance: every product change either updates them or leaves them lying about what they check. And a written pass count is a seductive metric, because it measures how much was executed rather than how much was learned. A hundred green procedures on a feature nobody has explored is confidence without evidence. There is a subtler cost too. A detailed procedure tells a tester what to look at, which reliably stops them from looking at anything else — the failure sitting one screen over is not in the steps, so it is not observed. That effect is one of the strongest practical arguments for keeping unscripted time in the plan even where scripts are mandatory. ## Under peak conditions and other special cases Some work needs both. Verifying that the review queue holds up at a 1,200-request-per-minute peak needs a fixed, repeatable procedure to make two runs comparable, and needs a person watching the system during the run to notice the thing nobody thought to assert — a lag that only appears after several minutes, an error that is logged but not surfaced. Treat "scripted or exploratory" as a property of each piece of work rather than of the team or the release. ## Answering this in an interview Name the property first — sameness — then the four cases it buys, then volunteer the cost. A senior answer ends with the pairing rather than a winner: exploration to find what nobody specified, a written procedure to hold what you found, and a deliberate decision about which each piece of work needs.

  • A team says its written suite is comprehensive, so unscripted time is unnecessary. What is your response?
    The suite is comprehensive with respect to what its authors imagined, which is the only thing it can be. It cannot contain the tests that could only be designed after seeing how the product actually behaves, and a detailed procedure actively narrows attention to its own steps. I would keep a bounded block of unscripted work on each risky area and judge it by what it finds that the suite structurally could not.
  • You wrote a procedure for a defect found while exploring. How detailed should the steps be?
    Detailed enough that someone without the product knowledge reproduces it reliably, and no more. Fix the preconditions and the data that matter, state the interruption or timing precisely, and assert every place the state is visible rather than just the one that looked wrong. Extra decoration makes the procedure brittle against product changes without making the verdict more trustworthy.
  • Does a written procedure lose its value once a check is executed by a machine instead of a person?
    No — writing it down is what made it repeatable and reviewable in the first place, and the encoded intent survives whoever or whatever runs it. What changes is the cost profile and how often it can run. Whether a given written check deserves to be executed automatically is a separate decision with its own criteria, and answering it does not change why the check was worth documenting.

saying these in an interview costs you the question

  • Claiming exploration can replace all documented tests
  • Claiming scripts find more defects than exploration generally
  • Treating a green pass count as evidence of quality
  • Ignoring that scripts only find what their author imagined
  • Forgetting audit and hand-off needs entirely
  • Choosing one approach for the whole team instead of per work item

context