skip to content

questions

4

What is risk-based test selection, and how do likelihood and impact rank what gets tested first?

level: juniorimportance: must knowfreq 68%

answer

  1. Two axes, not one
  2. How probable, and how costly
  3. Ranked list drives depth
  4. Bands get declared techniques
  5. Ordinal scores multiply badly

basics

~20 s

Risk-based test selection ranks each area by how likely it is to fail and how badly a failure would hurt, then spends the deepest testing on the highest-ranked areas and the least on the lowest.

solid answer

~50 s

Exhaustive testing is impossible, so the real question is never *what could we test* but *what do we test first and how deeply*. Risk-based selection answers it by scoring each area on two axes: **likelihood** — how probable a failure is, driven by complexity, recent change, unfamiliar technology and past defect density — and **impact** — how much a failure costs in money, safety, legality, data integrity or affected users. The two scores place the area in a band, and each band gets a declared depth: top-band areas get several techniques plus exploratory sessions, middle bands get the main flows and key negative cases, bottom-band areas get a shallow check or nothing at all. The output is a ranked risk register, not a feeling, so the plan can be argued with, reviewed by developers and business people, and revised when evidence changes it.

code

pseudocode · 15 lines
pseudocode
item = {
    area: "settlement calculation",
    likelihood: 4,      # heavy recent change, complex branching
    impact: 5,          # money moved, loss hard to reverse
    detectable: false   # wrong values look plausible
}

band(item):
    if item.impact >= 4 and item.likelihood >= 3: return "TOP"
    if item.impact >= 4 or item.likelihood >= 4:  return "MIDDLE"
    return "BOTTOM"

depth("TOP")    = [boundary, negative_paths, combinatorial, exploratory_session, design_review]
depth("MIDDLE") = [main_flows, key_failure_paths]
depth("BOTTOM") = [shallow_confirmation]   # or: none, recorded as accepted

go deeper

for a junior

Be ready to name the two axes and give a concrete driver for each — recent change for likelihood, money or data loss for impact — and to say that the ranking decides how much depth an area gets, not merely the order of execution.

for a middle

An interviewer expects the mechanics: small ordinal scales with written descriptors, bands mapped to declared techniques, and who supplies each axis. Be able to explain why multiplying two ordinal scores can mislead.

for a senior

Show the ranking changing behaviour under pressure: which area you cut when a build slips, what evidence moves an item between bands mid-release, and how you keep the register from becoming a document written once and never opened.

for a principal

Own the framing across teams — one scale definition so two teams' registers can be compared, and a rule for what escalates. Be ready to argue that the value of the method is the deliberate decision not to test, not the ranked list itself.

## The problem it solves Any non-trivial system has more possible inputs, states and paths than could be exercised in any budget. That is not a resourcing complaint, it is a structural fact, and it means every test plan is already a selection. Risk-based selection makes that selection **explicit, ranked and defensible** instead of implicit and driven by whoever spoke last. ## The two axes A *risk item* is a way the delivered product could fail — an area, a feature, a quality attribute, an integration. Each item is scored on two axes. **Likelihood** — how probable is a failure here? It is estimated from evidence about the code and the team, not from intuition alone: - structural complexity and the number of interacting paths; - recent or heavy change (churn), and brand-new code with no field exposure; - unfamiliar technology, or a team new to the area; - historic defect density in that component; - number of external integrations and asynchronous steps; - ambiguous or contested requirements — a specification nobody can restate the same way twice tends to be built wrong. **Impact** — if it does fail, how bad is it? Estimated from consequence, not from engineering effort: - money moved or lost, and whether the loss is reversible; - safety and legal or regulatory exposure; - data integrity, especially corruption that spreads before anyone notices; - how many users are hit, and whether the failure is visible to them or silent; - whether the failure is detectable and recoverable in production, or discovered only by a customer. Detectability deserves emphasis because it is easy to miss. A loud failure that stops the flow is often cheaper than a quiet one that writes wrong values for a week: the loud one is self-reporting, the quiet one accumulates. ## From score to depth The ranking is only useful if it changes behaviour. The usual mechanism is **bands**, with a depth declared per band before execution starts: - **top band** — multiple techniques combined: boundary and equivalence analysis, negative and error paths, combinatorial coverage of the parameters that interact, exploratory sessions with a charter, and a second pair of eyes on the design; - **middle band** — the main success paths plus the handful of failure paths that matter, largely scripted, automated where the case is stable; - **bottom band** — a shallow confirmation that the area exists and responds, or a documented decision to do nothing. Writing depth per band rather than per case is what keeps the plan from silently drifting back to uniform effort as deadlines press. ## Scoring in practice Scales are usually small ordinals — three or five points per axis — with each point given a written descriptor so two people score the same item the same way. Scoring is done as a group, because the two axes have different owners: developers and architects have the best information about likelihood, while business, support and legal have the best information about impact. A scoring workshop where those groups disagree loudly is doing its job; the disagreement is the finding. One caution about the arithmetic. Multiplying two ordinal scores is convenient but not strictly meaningful: a 5x1 and a 1x5 both come out as 5, yet a rare catastrophe and a frequent nuisance deserve different treatment. Many teams therefore sort by impact first and use likelihood to break ties inside an impact tier, or plot items on a two-axis grid and treat the top-right corner as the band rather than trusting the product of the scores. ## The artefact and its life The ranking lives in a **risk register** — a list of items with owner, likelihood, impact, band, planned mitigation and current status — often visualised as a grid with likelihood on one axis and impact on the other. The register is a living document. Every executed test, every defect found, and every piece of real usage evidence is information about likelihood, and the ranking is expected to move during the release rather than being fixed once at planning time. ## Why interviewers ask it Because it reveals whether a candidate can *choose*. Anyone can list techniques; the skill is in saying which areas get them and, harder, which areas deliberately get nothing. Risk-based selection is the vocabulary that lets that second sentence be said out loud to a business stakeholder without it sounding like negligence.

  • Two items score the same after multiplying likelihood by impact — one rare and catastrophic, one frequent and trivial. How do you order them?
    The equal product is an artefact of multiplying ordinal scores, not a real tie. Sort by impact first and let likelihood break ties inside the tier, because a rare catastrophe usually needs verification and a containment plan while a frequent nuisance can often be handled by a fast fix path. Say which rule you are using in the register so the ordering is reproducible rather than a judgement call re-made each week.
  • What evidence would you use to estimate likelihood for an area you have never tested?
    Change history and code churn, the defect record of the surrounding component, the number of integrations and asynchronous steps involved, how new the technology is to the team, and how confidently the developers can restate the requirement. Where none of that exists, treat the unknown itself as elevated likelihood and buy information cheaply first — a short exploratory session that tells you whether the area deserves a band, before committing depth to it.
  • Who should be in the room when likelihood and impact are scored?
    Both sides of the estimate. Developers and architects carry the best information about likelihood — complexity, churn, fragility — while product, support, legal or operations carry the best information about impact. Testers facilitate and hold the scale definitions steady. Scoring alone and circulating the result for comment produces a register nobody argues with and therefore nobody owns.

Insurance underwriting works the same way: nobody inspects every house on the street, they rank by how likely a claim is and how large it would be, then spend the inspector's day on the top of that list.

saying these in an interview costs you the question

  • Claims every area must be tested equally
  • Scores only impact and ignores likelihood
  • Treats the ranking as fixed after planning
  • Confuses a high-risk area with a difficult one
  • Ranks by how easy an area is to test
  • Produces a register that never changes any depth

context

open as a page

What is the difference between a product risk and a project risk in test planning?

level: middleimportance: should knowfreq 46%

basics

~20 s

A product risk is a way the delivered system could fail in use, so the response is test depth and design. A project risk threatens the delivery effort itself, so the response is planning, contingency and escalation.

open as a page

A marketplace bidding engine has produced silent data corruption in an area your risk register ranked low. How do you re-rank mid-release?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Treat the incident as evidence about the ranking method, not just one row. Raise that item on both axes, re-score every item resting on the same assumption, and fund the new depth by cooling an area the evidence has cleared.

open as a page

As the owner of a risk register, how do you defend a deliberate decision to leave an area unverified?

level: principalimportance: should knowfreq 40%

basics

~20 s

Make it a decision, not an omission: name the failure being accepted, score it on the same scale as what you did verify, offer a cheaper containment, and have the consequence-owner accept it with a review date.

open as a page