skip to content

Test Planning & Governance

How a team decides what to verify, how deeply, when to start and when a build is good enough to ship. Interviewers probe it because owning a release takes more than writing tests.

on this pageshow

explore

questions

page 1 of 2

What are entry criteria for a test cycle, and what do they prevent?

level: juniorimportance: must knowfreq 62%

answer

  1. Conditions before the cycle may start
  2. Build, environment, data, agreement
  3. Answerable yes or no by an outsider
  4. The smoke subset green on this build
  5. The agreed way to say not yet

basics

~20 s

Entry criteria are the conditions that must hold before a test cycle starts: a deployed build with a known change list, a healthy environment, seeded data, and agreed acceptance criteria. They prevent a cycle from burning its budget on setup failures.

solid answer

~50 s

Entry criteria are the agreed conditions that must be true before a test cycle is allowed to start — often called a **definition of ready for test**. The usual set: a build deployed to the intended environment with a published change list, the environment healthy and its dependencies reachable or stubbed, test data seeded to a known state, acceptance criteria agreed and testable, and the smoke subset green on that exact build. The point is not ceremony. A cycle started on a half-ready build spends its time diagnosing infrastructure, and every failure it reports has to be re-triaged once the setup is fixed. Entry criteria also give whoever owns the cycle a defensible way to say *not yet* without it sounding like an opinion. Keep them few, binary and observable — each one answerable yes or no by somebody who was not in the room when it was written.

code

pseudocode · 14 lines
pseudocode
ready_for_test(cycle):
    checks = [
        ("build deployed to test env", deployed_version(env) == cycle.build),
        ("change list published",      cycle.change_list is not empty),
        ("dependencies reachable",     all(ping(d) for d in env.dependencies)),
        ("data seeded",                seeded_personas(env) == expected_personas),
        ("smoke subset green",         smoke_result(cycle.build) == "pass"),
    ]
    unmet = [name for (name, ok) in checks if not ok]
    if unmet is empty:
        return START
    if waiver_recorded_for(unmet, approved_by=cycle.owner):
        return START_WITH_WAIVER
    return HOLD

go deeper

for a junior

Be ready to list the usual entry conditions in your own words — build deployed and identified, environment up, data seeded, acceptance criteria agreed, smoke subset green — and to say why starting without them wastes the cycle.

for a middle

An interviewer expects you to explain the wording discipline: each criterion binary and observable, checked by someone uninvolved, few enough to verify in minutes. Show that you know 'stable enough' is not a criterion.

for a senior

Demonstrate that you have used them under pressure: a recorded waiver instead of a silent start, re-triage of failures contaminated by a bad environment, and the habit of attributing lost time to the unmet condition rather than to slow testing.

for a principal

Own the tradeoff between a set so loose it is decorative and one so strict nothing starts. Be able to say how you would tune the criteria after a quarter of evidence, and how waivers get reviewed rather than accumulated.

## What an entry criterion is An **entry criterion** is a condition the team has agreed must hold before a test cycle is allowed to begin. It is a *process* condition, and that is what separates it from the other thing called a criterion in the same conversation: a feature's **acceptance criteria** describe what the software must do to be considered correct, while an entry criterion describes whether the situation is fit for anyone to start checking anything at all. Teams that work in short iterations often carry the same idea under the name **definition of ready for test**. ## The usual set Most teams' entry criteria fall into five families. 1. **Build identity.** A build is deployed to the intended environment and you can name exactly what is in it — a version string, and a change list of what merged since the last cycle. Without this, a failure cannot be attributed and a fix cannot be confirmed. 2. **Environment health.** The environment is up, its collaborators are reachable or replaced by agreed stand-ins, and background jobs and schedulers are in the state the cases assume. 3. **Data readiness.** Test data is seeded to a known state, including the awkward accounts — the expired one, the one at a limit, the one in a locale that formats dates differently. 4. **Agreement.** Acceptance criteria for the scope are written, reviewed and testable; anything ambiguous has been raised before execution rather than discovered as a disputed defect. 5. **A live check.** The smoke subset passes on that exact build. This is the one criterion that is not a promise but a measurement, and it is usually the most valuable of the five. ## Few, binary, observable The quality bar for an entry criterion is whether a person outside the conversation can answer it yes or no. "The build is stable enough to test" fails that bar: stable by whose judgement, measured how? "The smoke subset passed on build 4.19.2 in the shared test environment" passes it. Vague criteria do not merely fail to help — they are worse than nothing, because they create the appearance of a gate while leaving the decision exactly where it was, in the loudest person's hands. ## A worked cycle An 11-person team building a document e-signing flow agrees five entry criteria for each cycle. On one cycle the build lands at 09:40 with a change list of 14 merges, but the signature-callback dependency is returning errors and the data seed produces only 3 of the 9 signer personas the suite assumes. The criteria are not met; the cycle starts anyway, because the date is close. By mid-afternoon, 22 of the first 31 failures trace to the missing personas and the unreachable callback. All 22 have to be re-run once the environment is repaired, and two of them were filed as defects and had to be withdrawn — which cost developer time as well as test time. The one genuine defect of the day, an off-by-one at the signer-limit boundary where an envelope capped at 12 signers quietly accepted a 13th, is not found until the following morning. The lesson is not the defect. It is that a day of execution produced results nobody could trust, and the team paid for the same cases twice. ## The two failure modes Entry criteria go wrong in both directions. **Too loose**, and they are decoration: nobody checks them, cycles start on anything, and the criteria are quoted only in the post-mortem. **Too strict**, and no cycle ever legitimately starts — "zero open defects in the component" is the classic overreach, since a component always has open defects and the team simply learns to ignore the list. A useful set is short enough to check in a few minutes at the start of the cycle. When a criterion cannot be met and the team decides to proceed anyway, that should be a recorded **waiver**, not a silence: which criterion was unmet, who decided to proceed, and what the expected cost is. The waiver is what makes the later conversation about wasted effort a fact rather than a memory. ## Who owns them The criteria are agreed by the people affected — whoever produces the build, whoever owns the environment, and whoever executes — and written before the cycle rather than argued during it. Ownership matters because the value of an entry criterion is almost entirely in the moment of pressure: it exists so that *not yet* is a previously agreed position rather than an individual's resistance. ## The symmetry with the exit side Entry criteria protect the cycle's budget; exit criteria protect the release decision. Both work by the same discipline — written in advance, phrased so they can be checked, and either met, waived on the record, or not met. A team that gets the entry side right usually finds the exit conversation easier, because it is already in the habit of stating conditions instead of impressions.

  • Who decides that an entry criterion has been met?
    Whoever owns the cycle checks it, but the criterion has to be phrased so that the check is not a judgement call — a version match, a green smoke result, a seeded-data count. If deciding whether a criterion is met needs a debate, the criterion is written wrong. Agreement on the wording belongs to everyone affected: the people who produce the build, own the environment and execute the cases.
  • The date is fixed and an entry criterion is unmet. What do you do?
    You can still start — the criteria are not a veto — but you record a waiver naming the unmet criterion, who decided to proceed and the expected cost, and you plan for re-execution of anything the gap contaminates. What you must not do is start silently, because then the wasted re-runs look like slow testing rather than a known consequence of a known decision.
  • How is an entry criterion different from a feature's acceptance criteria?
    An acceptance criterion is about the software: it states an observable behaviour the feature must exhibit. An entry criterion is about the situation: it states whether the build, environment, data and agreements are in a state where checking that behaviour is worth doing. Confusingly, a written and agreed set of acceptance criteria is itself often one of the entry criteria — but the two answer different questions.

A pre-flight checklist: none of the items make the flight succeed, they only establish that starting is not a waste of fuel.

saying these in an interview costs you the question

  • Treats entry criteria as paperwork with no power to delay a cycle
  • Writes criteria like 'the build is stable' that nobody can check
  • Confuses entry criteria with a feature's acceptance criteria
  • Demands zero open defects before any cycle may start
  • Starts on an unready build without recording the decision
  • Assumes readiness is the environment owner's problem alone

context

open as a page

What is the difference between verification and validation in a software lifecycle?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Verification asks whether the product was built according to its specification. Validation asks whether that specification described the right thing to build. One checks conformance to a written statement, the other checks fitness for a real need.

open as a page

In a test cycle status report, why is a blocked case counted separately from a failed one?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A failed case ran and gave the wrong result, so it is evidence about the product. A blocked case never ran, because something stopped it. Merging the two hides untested scope and blames the product for an environment problem.

open as a page

What is the difference between confirmation testing and regression testing after a defect fix?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Confirmation testing re-runs the case that exposed the defect, to prove the fix works. Regression testing re-runs other cases around the change, to prove the fix broke nothing that previously passed. Both are needed after a fix.

open as a page

What is risk-based test selection, and how do likelihood and impact rank what gets tested first?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Risk-based test selection ranks each area by how likely it is to fail and how badly a failure would hurt, then spends the deepest testing on the highest-ranked areas and the least on the lowest.

open as a page

What does shift-left mean in testing, and what can a tester do before any code exists?

level: juniorimportance: must knowfreq 71%

basics

~20 s

Shift-left means moving verification earlier in delivery, toward requirements and design, instead of leaving it as a phase after the build. Before code exists a tester reviews requirements for ambiguity and testability and agrees concrete examples of expected behaviour.

open as a page

What is the difference between a test strategy document and a test plan?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A test strategy is standing and general: which test levels run, the stance on environments and data, the automation stance, tooling policy, who owns quality. A test plan is per-release: scope, schedule, staffing, deliverables, risks, sign-off.

open as a page

Why is a test case count a weak basis for sizing a test effort?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Because cases vary hugely in cost: a one-step check and a case needing days of seeded data both count as one. Size instead from scope and the depth each risk band needs, then convert with measured throughput.

open as a page

What does reading a requirement-to-case traceability link forward tell you that reading it backward does not?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Forward reading starts at a requirement and asks which test cases verify it, answering whether anything checks it. Backward reading starts at a test case and asks which requirement justifies it, answering whether that case still earns its keep.

open as a page

In a traceability record linking requirements to test cases and defects, which gap and orphan findings appear, and what does each call for?

level: juniorimportance: must knowfreq 52%

basics

~20 s

Four findings recur: a requirement with no linked test case, a test case linked to no requirement, a defect linked to no test case, and a linked case that has never been executed. Each calls for a different action.

open as a page

What columns does a requirements traceability record need, and what does each answer?

level: juniorimportance: must knowfreq 46%

basics

~20 s

One row per requirement unit, carrying its identifier and text, the cases that verify it, the latest execution verdict, the build it ran against, and any linked defects. Every column answers one question about that statement's verification.

open as a page

What is the difference between code-execution coverage and requirement coverage?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Code-execution coverage measures how much of the written code ran during testing; requirement coverage measures how many agreed requirements have a check confirming them. The denominators differ, so one can be high while the other is low.

open as a page

In the V-model, how does each specification stage pair with a verification level?

level: middleimportance: must knowfreq 64%

basics

~20 s

The V-model bends the sequence into a V. Each specification stage on the descending arm has a partner verification level at the same height on the ascending arm, and that stage's document is the source for the checks performed at its partner level.

open as a page

What belongs in a smoke suite that gates a build, and what disqualifies a case from it?

level: middleimportance: must knowfreq 66%

basics

~20 s

A build-gating smoke suite is a short, broad set of checks that decides whether a build is worth testing further. Include one fast, deterministic case per critical path. Exclude slow, data-hungry, flaky or deep edge-case checks.

open as a page

When a requirement changes, how far do you walk its traceability links to build the impact set?

level: middleimportance: must knowfreq 58%

basics

~20 s

Walk outward hop by hop — the requirement's acceptance criteria, the test cases linked to them, then the defects raised in that area before — and stop at the hop that stops adding anything you would re-verify.

open as a page

What can a maintained traceability record prove that links left as a by-product of ordinary work cannot?

level: middleimportance: must knowfreq 52%

basics

~20 s

A maintained traceability record proves, as of a stated moment, that every requirement had a named verification against a named version. By-product links only show what someone happened to reference, and fall silent wherever nobody referenced anything.

open as a page

How should blocked and not-run cases count against a pass-rate exit criterion?

level: middleimportance: should knowfreq 51%

basics

~20 s

Blocked and not-run cases are unknowns, not passes. Compute the pass rate against the planned case count rather than the executed count, and report blocked and not-run alongside it — otherwise the criterion rewards running less.

open as a page

Your test execution burn-down has been flat for three days — how do you diagnose why?

level: middleimportance: should knowfreq 44%

basics

~20 s

Split the single line into its parts. If blocked is rising, execution is obstructed; if the planned total is rising, the cycle is moving but scope grew; if neither moved, throughput itself collapsed. The chart names the symptom, the state series names the cause.

open as a page

What is the difference between a product risk and a project risk in test planning?

level: middleimportance: should knowfreq 46%

basics

~20 s

A product risk is a way the delivered system could fail in use, so the response is test depth and design. A project risk threatens the delivery effort itself, so the response is planning, contingency and escalation.

open as a page

How do you review a requirement for testability, and what makes one untestable?

level: middleimportance: should knowfreq 57%

basics

~20 s

Check two properties: controllability, whether you can put the system into the state the rule describes, and observability, whether you can see the promised outcome. A requirement is untestable when it names no measurable outcome, no reachable precondition, or no way to observe the result.

open as a page

What belongs in a one-page test strategy, and what should be left out?

level: middleimportance: should knowfreq 52%

basics

~20 s

A one-page test strategy states standing decisions: which test levels run, the stance on environments and test data, the automation stance, sanctioned tooling classes, and who owns quality. Leave out release scope, dates, staffing and anything already recorded elsewhere.

open as a page

When do you estimate test effort from historical throughput rather than expert three-point ranges?

level: middleimportance: should knowfreq 46%

basics

~20 s

Use historical throughput when the coming work resembles work you have already measured on the same product and team. Use expert three-point ranges and consensus rounds when the work is new, the process changed, or no comparable history exists.

open as a page

Which traceability findings are honest noise rather than plan defects, and how do you scope the analysis so the list stays actionable?

level: middleimportance: should knowfreq 34%

basics

~20 s

Cases written for build health or conventions, statements verified by review, and requirements not yet in scope are noise rather than defects. Scope it to the current release and record each exception with a reason and a review date.

open as a page

What must a verification evidence record carry for a reviewer outside the team to rely on it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Four things at minimum: what was verified, the exact version it ran against, who or what performed it and when, and the outcome with the stored output it was read from — kept readable for as long as required.

open as a page

Requirements are all confirmed but much of the code never runs — what does a requirement-confirmed figure miss?

level: middleimportance: should knowfreq 45%

basics

~20 s

It misses everything not on the agreed list. Failure paths, defensive branches, compatibility code and residue from removed features are in no requirement, so no value of that figure reports them. Only the code-side denominator contains them.

open as a page

A build slips and your test window halves — how do you re-plan depth and scope mid-cycle?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Re-forecast from measured throughput to learn how much fits, then cut depth before breadth: one pass over every planned area, variations and repeat permutations dropped. Record exactly what was cut so the shortfall is visible rather than absorbed silently.

open as a page

How do you choose the regression subset to run for a change when the full pack no longer fits the feedback budget?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Select by change impact: the cases traceable to the changed components or requirements, cases asserting on what the change writes, cases that failed recently in that area, and the new case for the change. Keep the full pack running on a slower schedule as the safety net.

open as a page

A marketplace bidding engine has produced silent data corruption in an area your risk register ranked low. How do you re-rank mid-release?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Treat the incident as evidence about the ranking method, not just one row. Raise that item on both axes, re-score every item resting on the same assumption, and fund the new depth by cooling an area the evidence has cleared.

open as a page

Why do long test plan documents go stale faster than the code they describe?

level: seniorimportance: should knowfreq 41%

basics

~10 s

Nothing forces a document to change when reality does. Wrong code fails a build; a wrong sentence fails nothing, so a long plan's duplicated facts drift silently until nobody trusts the document.

open as a page

showing 1–30 of 55