skip to content

Exploratory & Manual Testing

Testing where a person investigates the product instead of replaying a script, choosing what to try next and deciding what counts as wrong. Interviewers use it to see how you think without a spec.

on this pageshow

questions

25

Why narrow a defect to the shortest reliable reproduction before you report it?

level: juniorimportance: must knowfreq 68%

answer

  1. Fewer conditions, faster diagnosis
  2. A claim, not a diary
  3. Delete one condition, then re-run
  4. Restore anything the failure needed
  5. Start state and hit rate count

basics

~20 s

A short reproduction proves which conditions actually cause the failure. Every step you can delete without losing the failure is noise that slows diagnosis, invites a cannot-reproduce close, and hides the real trigger from whoever has to fix it.

solid answer

~50 s

A reproduction is a claim about which conditions are sufficient to make the failure appear, not a diary of what you happened to do. So I reproduce it at least twice from a known start state, then delete one step or condition at a time and re-run; if the failure survives the deletion I keep the shorter version, if it disappears I put the condition back because it is part of the trigger. The start state counts as a condition - the build, the account's role, the seeded data, the configuration - so I record it rather than assume it. I also record how reliably it fires, for example nine attempts in ten, instead of writing 'always'. The result is evidence someone can act on in minutes rather than a walkthrough they have to re-derive.

code

pseudocode · 18 lines
pseudocode
steps = recorded_walkthrough()            # 41 recorded actions
start = known_state(build, account_role, seeded_data, settings)

for candidate in copy_of(steps):
    trial = steps_without(candidate)
    hits  = 0
    repeat 5 times:
        reset_to(start)
        if run(trial) shows_the_failure: hits = hits + 1

    if hits >= 4:
        steps = trial                     # condition was incidental, drop it
    else:
        keep(candidate)                   # condition is part of the trigger

report(start_state = start,
       conditions  = steps,
       hit_rate    = measure(steps, attempts = 20))

go deeper

for a junior

Be ready to say what a reproduction is for: proving which conditions cause the failure. Recall the loop - reset, reproduce twice, remove one condition, re-run, restore anything the failure needed - and remember to write down the build, the account and the data.

for a middle

Explain the mechanics, especially the asymmetry: a deletion is only kept if the failure survives it. Be able to name the invisible conditions - identity, seeded data, configuration, leftover cache state, ordering - and to express reliability as a count of attempts rather than as 'sometimes'.

for a senior

Show that you read the surviving conditions as evidence, not just as instructions, and that you know when minimising destroys the bug because the sequence itself is the trigger. Interviewers expect you to describe the cost side too: how long you spend narrowing before filing what you have.

for a principal

Own the economics. Decide who is cheapest to do the isolation, what minimum evidence you require from everyone so that reports do not ping-pong, and how you keep that bar from quietly suppressing reports of the messy defects nobody can narrow.

### A reproduction is a claim, not a diary Most weak defect reports are transcripts: the tester writes down everything they did between logging in and noticing something wrong. That is a record of a session, not a reproduction. A reproduction is a much stronger claim - *these conditions, from this starting point, are enough to make the failure appear* - and the work of narrowing is the work of turning the first thing into the second. The claim matters because the reader is not you. They did not see the screen, they do not know which of your forty-one actions was load-bearing, and they will spend their first twenty minutes re-deriving what you already know. Every incidental step you leave in is a step they have to perform, doubt, and eventually eliminate themselves. ### The minimisation loop The loop is mechanical and worth doing by hand at least once until it becomes instinct: 1. **Get a known start state.** A fresh account or a reset data set, a named build, a recorded configuration. Without this, nothing after it is repeatable. 2. **Reproduce it twice.** One observation is an anecdote. Two from the same start state is a reproduction. 3. **Delete one condition and re-run.** A step, a field value, a role, a piece of seeded data, an option. 4. **Keep the deletion only if the failure survives it.** If the failure disappears, restore the condition - you have just learned it is part of the trigger, which is a finding in itself. 5. **Stop when every remaining condition is load-bearing**, then measure how often the shortened version fires. The important asymmetry is in step 4. Deleting steps until the failure stops appearing and then filing the shortened version anyway is the single most common way a reproduction becomes untrustworthy - it produces steps that genuinely do not reproduce anything, and the report comes back closed as not reproducible. ### Conditions that are not in the step list Juniors usually minimise the clicks and forget everything around them. The conditions that most often turn out to be essential are invisible in a step list: - **Identity and permissions** - which role, which account, which team the account belongs to. - **Data** - not 'some trips' but the specific shape: a record with an empty optional field, a very large history, a duplicate. - **Build and configuration** - the exact build identifier, feature switches, regional or unit settings. - **Leftover state** - a cache warmed by the previous test, a background job that ran between your two attempts. - **Timing and ordering** - two actions inside the same second, a page left open while something else changed underneath it. If you cannot say which of these you controlled, you have not finished narrowing; you have only shortened the visible part. ### Reliability belongs in the reproduction 'Always' and 'sometimes' are both worthless to the person fixing it. A count is not: *fires on 9 of 10 attempts on this build, with this account and this data*. That single line tells the reader whether to expect it on their first try, and it tells you honestly whether you have finished isolating. A hit rate well below one is a signal that a condition you have not identified is still varying between runs. ### When a long reproduction is the right answer Sometimes the length is the bug. Accumulated state, a specific ordering, a slow leak, a scheduled job - these need a sequence, and cutting it destroys the evidence. The discipline still applies: minimise everything you can, then state plainly that the remaining sequence is essential and say what you tried to remove and could not. A reader trusts 'I removed twelve steps and these six are required' far more than a bare six-step list with no history. ### A worked example A tester on a ride-hailing dispatcher records a forty-one-action walkthrough that ends with a completed trip showing the wrong fare. Narrowing it takes about twenty minutes: the map interactions, the driver chat, the two cancelled requests and the profile visit all come out with the failure intact. What does not come out is the seeded trip whose promotion code was applied after the trip started, the dispatcher account being in the regional supervisor role, and accepting the trip while a second request is already open. Six conditions remain, from a reset account, firing on 9 of 10 attempts. The report is now a twelve-line note that a developer can reproduce before their coffee cools - and the three surviving conditions are themselves the strongest hint about where the defect lives.

  • How many times do you reproduce a defect before you call the reproduction reliable?
    At least twice from a reset start state before I believe it at all, and then a fixed batch - twenty attempts is a reasonable default - to get a hit rate. I report the count, not an adverb: 'fires on 18 of 20 attempts' or 'fires on 3 of 47'. A hit rate far below one is my own signal that some condition is still varying between runs and the narrowing is not finished.
  • When is a long reproduction sequence the right thing to report?
    When the length is the defect: accumulated data, a required ordering, a background job that has to run between two actions, or a resource that only exhausts after many cycles. I still minimise everything that can go, then say explicitly that the remaining sequence is essential and list what I removed. That tells the reader the length was tested rather than merely copied out of my notes.
  • You cut a reproduction from forty-one steps to six. What do the six surviving conditions tell you beyond how to trigger it?
    They are the best cheap hypothesis about where the defect lives. If the survivors are a particular role, a promotion applied mid-trip and a second request open at the same time, the defect is almost certainly in how those three interact, not in the map or the chat. I put that reading in the report as a hypothesis, clearly labelled as one, so the developer can accept or discard it without having to re-run my narrowing.

Narrowing a reproduction is like isolating an ingredient after a bad meal: you do not re-eat the whole menu, you drop one dish at a time and see whether you still get sick.

saying these in an interview costs you the question

  • Pastes the whole recorded walkthrough unchanged
  • Writes 'it does not work' with no start state
  • Claims it always fails after a single observation
  • Deletes steps without re-running to confirm
  • Keeps deleting until the failure disappears, then files that
  • Treats build, role and seeded data as outside the reproduction

context

open as a page

What is a bug bash and who takes part in one?

level: juniorimportance: must knowfreq 56%

basics

~20 s

A bug bash is a short, time-boxed event where many people — developers, support, product and testers — hunt defects in one shared build at once, each covering an assigned area and filing into one channel.

open as a page

What is a test charter in exploratory testing, and what does its three-part shape name?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A test charter is a one- or two-sentence mission for a single exploratory session. The common shape names three things: what to explore, what to explore it with, and what information the session should discover.

open as a page

What is a test oracle, and what makes one fallible?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A test oracle is whatever you use to decide that an observed behaviour is wrong: a written specification, a standard, a comparable product, an earlier version, or your own expectation. Each is an imperfect model, so it can be wrong too.

open as a page

What is a timeboxed test session in session-based test management?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A session is one uninterrupted block of testing, commonly 60, 90 or 120 minutes, spent on a single charter and written up as one reviewable record. The session, not the test case, is the unit of work you plan and count.

open as a page

What is exploratory testing, and how does it differ from ad-hoc testing?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Exploratory testing learns the product, designs a test and runs it as one loop, with each result steering the next test. Ad-hoc testing is the same poking without a mission, a time bound, or any record of what was covered.

open as a page

What does a session sheet record after an exploratory test session?

level: middleimportance: must knowfreq 51%

basics

~20 s

A session sheet records the charter, the areas covered, the tester and start time, running test notes, the defects found, the obstacles hit, and a rough split of where the time went. It is the session's only durable output.

open as a page

How do you turn an intermittent failure into a reliable reproduction?

level: middleimportance: should knowfreq 58%

basics

~20 s

Treat intermittence as a hidden variable, not randomness. Capture artefacts from a failing run, diff it against a passing one, then flip one candidate condition at a time and re-run a fixed batch until the hit rate moves.

open as a page

Before a bug bash, how do you prepare the build, the data and the area split?

level: middleimportance: should knowfreq 46%

basics

~10 s

Freeze one deployed build and block deploys for the window, seed an account and interesting data per role, carve the product into named assigned areas, publish known issues, and agree one capture format.

open as a page

How do you tell an exploratory test charter is too broad, and how do you split it?

level: middleimportance: should knowfreq 57%

basics

~20 s

A charter is too broad when one uninterrupted session cannot finish it and the tester cannot say afterwards what was covered. Split it along a real dimension — sub-area, data condition, user role, quality attribute or recent change.

open as a page

Which oracles can a tester use when no written specification exists?

level: middleimportance: should knowfreq 54%

basics

~20 s

Consistency oracles: with history, with the product's own claims, with comparable products, with itself across features, with users' expectations, with its stated purpose, and with applicable standards - plus whether the behaviour is explainable at all.

open as a page

How do you probe a mild symptom for the worse failure hiding behind it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Treat the visible symptom as one sample of a defect, not its full extent. Vary data, timing, role and configuration around it, then follow the bad value downstream - into storage, exports and totals - to see whether something worse happens unseen.

open as a page

How do you stop a bug-bash capture channel from filling with duplicate reports?

level: seniorimportance: should knowfreq 39%

basics

~10 s

Post known issues before the start, fix one short title shape with an area tag, keep the growing list visible so people scan before filing, and put one person on duty merging duplicates live.

open as a page

How do you proceed when a result looks wrong but no oracle can settle it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Name it as an oracle problem: rank candidate oracles by strength and independence, prove a contradiction with a partial check such as cross-view consistency, spend expert time on one narrow question, and report the ambiguity if nothing resolves it.

open as a page

How do you report exploratory testing progress when there is no test-case count?

level: seniorimportance: should knowfreq 49%

basics

~20 s

Report in sessions and areas: sessions completed against sessions planned, which product areas have had sessions and which have not, roughly how much of the time stayed on charter, and the obstacles currently slowing sessions down.

open as a page

When does a written, scripted test procedure still beat exploratory testing?

level: seniorimportance: should knowfreq 58%

basics

~20 s

A written procedure wins whenever sameness matters more than discovery: re-running the identical check every build, producing an evidence trail someone outside the team can audit, handing work to a person without product knowledge, and confirming a known defect the same way twice.

open as a page

As a lead, how much investigation do you require before a defect is filed?

level: principalimportance: should knowfreq 41%

basics

~20 s

Set a cheap, fixed evidence floor plus a timebox rather than an open-ended demand: capture build, start state, artefacts and a hit rate, isolate for an agreed period, then file with what you have and say where you stopped.

open as a page

When is paying an external crowd for device and locale breadth worth it?

level: principalimportance: should knowfreq 31%

basics

~10 s

Buy a crowd when the real-device, network and locale breadth you need exceeds what you can staff, the build is safe to expose, and you have reviewer capacity to filter a low acceptance ratio.

open as a page

Sessions completed per week is proposed as a tester productivity target — how do you respond?

level: principalimportance: should knowfreq 36%

basics

~20 s

Push back, and offer something better. A session count is a good planning input and a terrible individual target: it is trivially inflated by shorter, shallower, easier sessions, and the moment it drives appraisal the session sheets stop being honest.

open as a page

How do you budget unscripted exploration against scripted testing for a release?

level: principalimportance: should knowfreq 44%

basics

~20 s

Allocate by what each area needs, not by a fixed ratio. Spend written, repeatable effort where a verdict must be reproduced or shown to someone outside the team, and unscripted time where the risk is real but nobody yet knows what could go wrong.

open as a page

What happens in a session debrief, and who takes part in it?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

A debrief is a short conversation, usually a few minutes, in which the tester walks one other person through the session sheet: what happened, what was found, what got in the way, and what the area needs next.

open as a page

What is a tour in exploratory testing, and how do a feature tour and a data tour differ?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

A tour is a pass through the product with one lens held fixed, used to generate test ideas. A feature tour walks every function to learn what exists; a data tour follows one piece of data through every place that creates, changes, displays or deletes it.

open as a page

How do you raise a defect's perceived impact honestly instead of inflating it?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

Raise impact with evidence, never adjectives: replace a contrived path with the most realistic one that triggers the same defect, count who is exposed, follow the consequence to money, data or safety, and state the conditions that limit it.

open as a page

Where do new exploratory test charters come from, and how do you keep a charter backlog useful?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Charters come from recent changes, unusual data shapes, user roles, quality attributes, field and support signal, and above all from findings in earlier sessions. Keeping the backlog useful is mostly subtraction: retire stale entries, merge duplicates, keep each session-sized.

open as a page

How do you decide whether a stronger, costlier oracle is worth building?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Weigh what a wrong answer costs and how silently it fails against the oracle's build, run and maintenance cost, its false-alarm rate, and above all its independence from the implementation. Buy strength where failures are expensive and quiet.

open as a page