skip to content

How do you budget unscripted exploration against scripted testing for a release?

level: principalimportance: should knowfreq 44%

answer

  1. No single ratio survives two areas
  2. Discover something or confirm something
  3. Novelty, thin specification, consequence, recent surprises
  4. Exploration is cut first under pressure
  5. Escaped-defect audit tells you the split

basics

~20 s

Allocate by what each area needs, not by a fixed ratio. Spend written, repeatable effort where a verdict must be reproduced or shown to someone outside the team, and unscripted time where the risk is real but nobody yet knows what could go wrong.

solid answer

~50 s

I split by information need per area rather than by a company-wide percentage. Areas that are new, redesigned, badly specified, or carry serious consequences get unscripted blocks, because there the useful tests cannot be written in advance. Areas that are stable, contractual or externally auditable get written procedures, because there the need is a reproducible verdict rather than discovery. Then I protect the exploration budget explicitly — it is the first thing cut under schedule pressure, precisely when the change is largest and discovery matters most. To defend it I report exploration in areas covered, risks addressed and what remains unexplored, and I audit escaped defects for whether a written check could ever have caught them. A recurring answer of no is the evidence that the split is wrong, and it is more persuasive than any ratio.

go deeper

for a junior

Be ready to say the mix depends on the area rather than on a company rule, and to give one example each of an area that needs discovery and an area that needs a repeatable check.

for a middle

Explain the signals that move an area one way or the other — novelty, thin specification, consequence, stability, an external evidence requirement — and why a single ratio under-tests one area while wasting effort on another.

for a senior

Show you make and defend the call per release. Talk about protecting the unscripted block when the schedule compresses, and about reporting what remains unexplored rather than only what was found.

for a principal

Own the evidence loop and the staffing angle: an escaped-defect audit that tells you whether the gap is discovery or regression discipline, and a deliberate choice about who explores which area, since the return varies enormously with product knowledge.

## Why a fixed ratio is the wrong instrument Every team eventually proposes a number — seventy-thirty, or a fixed exploration day per sprint. The number is comforting and it is not a decision. Two areas in the same release can need opposite things: a re-skinned settings page that has not changed behaviour needs a repeatable check that nothing broke, while a newly built decision engine needs a person who is allowed to follow surprises. A single ratio applied across both under-tests one and wastes effort on the other. The useful question per area is: **do we need to discover something, or to confirm something?** Discovery is where the tests worth running cannot be written in advance, because they depend on facts nobody knows yet. Confirmation is where the expected behaviour is already agreed and the job is to show it still holds. ## The allocation heuristics A few signals reliably move an area toward unscripted time: - **Novelty.** The area is new, rebuilt, or integrated with something for the first time. Nobody's imagination of the failure modes is trustworthy yet. - **Thin or contested specification.** If what it should do is being decided as it is built, tests derived from the document will encode the document's gaps. - **Consequence.** Where being wrong is expensive — money, safety, legal standing, irreversible actions — you want someone actively hunting, not only confirming. - **Recent surprises.** An area that produced unexpected defects last release is telling you the team's model of it is wrong. And toward written, repeatable effort: - **An external evidence requirement** — a reviewer outside the team must see what was verified. - **Stability** — behaviour is settled and the risk is regression rather than the unknown. - **Hand-off and comparability** — many people, many configurations, results that must line up. - **Confirmed defects** — once understood, the reproduction is worth fixing in writing. ## A worked allocation A release of a grant-application review queue changes 14 areas. Nine are cosmetic or settled behaviour with an existing written pack: they need re-execution, not attention. Three carry a rebuilt bulk-decision path with a new write ordering — thin specification, irreversible consequences, and this is exactly where a partial-failure rollback was found last time, leaving 11 of 37 applications in disagreeing states across three views. Those three get most of the unscripted budget. The remaining two touch a queue that must hold up at a 1,200-request-per-minute peak; that needs a repeatable procedure so two runs are comparable, plus a person watching during the run, because the interesting failures there are the ones nobody thought to assert. Notice what the allocation is made of: a judgement per area, defensible in a sentence each, and re-made every release rather than fixed as a policy. ## Protecting the budget Unscripted time has a structural weakness in a plan: it has no visible deliverable until it produces one. When the schedule compresses, the written pack survives — it is countable and it is required — and exploration is quietly dropped, which is exactly backwards, because compression usually means the change was bigger or later than planned. Three defences work. First, make it a named, sized block of work in the plan, with an area and a risk attached, so cutting it is a visible decision rather than a silent omission. Second, report it in terms leadership can weigh — areas covered, risks addressed, and specifically what is still unexplored, which is the sentence that gets attention. Third, and most persuasive, run an escaped-defect audit: for each defect that reached production, ask whether any reasonable written check would have caught it. When the answer keeps coming back no, you have direct evidence that the missing capability is discovery, not more procedures. Conversely, if escapes are mostly things a written check would have caught, the honest conclusion is that your regression discipline is the gap and you should say so. ## Who explores, not just how much Budget is also a staffing question. The output of unscripted work scales with product knowledge and with the tester's stock of ideas about how things fail, so the same hour spent by different people returns wildly different value. That argues for putting your most product-literate people on the highest-consequence areas, for pairing a newcomer with them rather than handing a newcomer the risky area alone, and for pulling in people with different models of the system — support staff know what users actually do, developers know where the wiring is fragile. It also argues against treating exploration as a job title's private activity. Where a team is short of testers, deliberately allocated exploration by developers on each other's work buys real information, provided it is given a mission and a hand-back like any other block. ## Answering this in an interview Refuse the ratio, explicitly and politely, then give the per-area rule and two or three signals on each side. Finish with how you would know the split was wrong — the escaped-defect audit — because that is what separates a principal answer from a confident opinion.

  • Leadership asks for one number: what percentage of testing should be unscripted?
    I would give a starting number if pressed, say it out loud as a starting number, and immediately attach the rule that moves it — more unscripted time where the area is new, thinly specified or expensive to get wrong, less where behaviour is settled or an external reviewer needs an evidence trail. Then I would commit to reporting back with the escaped-defect audit, so the next conversation is about evidence rather than about the number.
  • How would you tell whether the current split is producing value?
    Two checks. For each defect that escaped to production, ask whether any reasonable written check could plausibly have caught it — a run of no means the gap is discovery. And look at what unscripted blocks actually returned: findings that the written pack could not structurally have produced, versus re-discoveries of what the pack already covers. Cheap findings from exploration that duplicate the pack usually mean the missions were aimed at settled areas.
  • A schedule slips and two days of unscripted time are the obvious cut. How do you argue?
    I point out that the slip usually means the change landed larger or later than planned, which raises the amount nobody has learned about it — the condition under which discovery is worth most. Then I make the cut explicit rather than fighting it outright: name the areas that will ship unexplored and the risks attached, and let the decision be made with that stated. Most of the time the block survives; when it does not, the record is honest.

saying these in an interview costs you the question

  • Applying one fixed ratio across every area
  • Treating unscripted time as slack rather than planned work
  • Cutting exploration first whenever the schedule slips
  • Judging exploration by defect counts alone
  • Assuming any hour of exploration returns the same value
  • Never checking whether escaped defects were script-findable

context