skip to content

How do you decide which journeys deserve end-to-end coverage, and keep that set from growing?

level: principalimportance: should knowfreq 44%

answer

  1. Treat the set as a portfolio
  2. Consequence and irreversibility of failure
  3. Could a cheaper test see this
  4. Publish a budget, enforce admission
  5. After each failure, push the check down

basics

~20 s

Select by consequence and uniqueness: journeys whose silent failure is unacceptable and whose outcome no cheaper test can observe. Then govern the set with a published budget, an admission rule for adding one, and a push-down habit after every failure.

solid answer

~40 s

Treat the set of journeys as a **portfolio under a budget**, not a backlog. Selection has two filters: **consequence** - what the business loses if this journey silently breaks, and whether that loss is reversible - and **uniqueness** - could a cheaper test observe the same outcome? A journey that passes both earns a slot; everything else belongs a level down. Then govern it: publish a budget in wall-clock and in expected diagnosis hours, make adding a journey require naming what it uniquely catches, and after every failure ask which cheaper test would have caught the same defect. Review the set on a cadence against evidence: journeys that have never failed for a real defect are candidates for retirement, and defects that escaped to users are candidates for a new one.

code

pseudocode · 8 lines
pseudocode
admitJourney(candidate):
    if not candidate.failureIsCostlyOrIrreversible():
        return REJECT("low consequence - cover it lower down")
    if candidate.outcomeObservableBy(cheaperLevel):
        return REJECT("not unique - write it at " + cheaperLevel)
    if layer.wallClock() + candidate.cost() > BUDGET:
        return REQUIRE_RETIREMENT(layer.lowestValueJourney())
    return ADMIT

go deeper

for a junior

Be ready to say that not every feature deserves a journey, and to give the first filter: what actually goes wrong for the business if this path silently breaks?

for a middle

Explain the uniqueness filter concretely - if a cheaper test could observe the same outcome, write it there - and be able to walk one feature and say which single path you would cover at this level and why.

for a senior

Show the counter-pressure you would install: an admission rule that names what a new journey uniquely catches, and the habit of writing a cheaper test after each failure so the expensive one may eventually retire.

for a principal

Own the portfolio framing end to end: a published budget in wall-clock and diagnosis hours, ownership placed where the cost is paid, an escaped-defect review that lets the set grow deliberately, and honesty that no universal number exists.

### Why this is a portfolio problem Every end-to-end journey is a standing liability: it consumes pipeline wall-clock, it needs an environment healthy enough to run, it breaks when the interface it drives is redesigned, and each ambiguous failure costs engineer hours to diagnose. It is also an asset: it is the only artefact that catches wiring, configuration and cross-component sequencing defects on a path real users walk. A portfolio holds assets under a constraint, and the constraint here is not moral - it is a budget nobody usually writes down. So write it down. Two numbers work: a **wall-clock ceiling** for the layer in the pipeline (what delay the team will tolerate before a release decision), and an **expected diagnosis load** (how many ambiguous failures per week the team can absorb). Both are choices, and different products choose differently - a payments platform and an internal reporting tool should not land on the same numbers. ### The two selection filters **Consequence.** Ask what is lost if this journey silently breaks for a day, and whether the loss is reversible. Money movement, safety, regulatory obligations, anything a user cannot undo, and anything that would be discovered by a customer rather than by the team all raise the score. For an airline seat-map service, *passenger reserves and confirms a seat on a paid booking* scores high: a silent failure is lost revenue plus a support queue. *Passenger changes the cabin-view colour theme* does not. **Uniqueness.** Ask which cheaper test could observe the same outcome. If a narrow test can, the journey adds cost without adding information. The journeys that survive this filter are the ones whose value lies in the crossing itself - several separately deployed components, real configuration, real serialization between them, real sequencing. A useful third input, where it is available, is **actual usage**: which journeys real traffic exercises most. It is evidence rather than opinion, and it usually reorders a team's intuitive list. ### Governance, because selection decays Selection is a one-off; the set grows continuously. Three mechanisms hold it: 1. **An admission rule.** Adding a journey requires stating in the change what it uniquely catches that no cheaper test can. Where the budget is already spent, adding one means retiring one. This converts a vague preference for fewer tests into a decision someone has to make explicitly. 2. **A push-down habit.** After every genuine end-to-end failure, ask which cheaper test would have caught the same defect, and write it. Do this consistently and the layer stops being the place defects are found and becomes the place they are prevented from escaping - and some journeys become retirable. 3. **A periodic review against evidence.** For each journey ask: has it caught a real defect in the last year, how many hours has it cost in diagnosis, and does the risk it covers still exist? A journey that has never failed for a real reason is not proof of a healthy system; it may simply be watching a path nothing changes any more. ### The escape-driven counterweight Governance that only ever removes is as wrong as growth that only ever adds. The counterweight is the **escaped-defect review**: when a defect reaches users, ask what would have caught it. Sometimes the honest answer is a journey that does not exist, and then the set should grow - deliberately, with the budget revisited rather than quietly exceeded. ### The organisational dimension The hardest part of this is rarely technical. Ownership decides everything: a suite owned by a separate group tends to grow, because the people adding journeys do not pay the diagnosis cost. Putting the journeys with the teams who own the components they cross aligns cost with benefit, and makes retirement conversations possible. Similarly, a shared suite that everyone depends on and nobody owns becomes the place where nobody may delete anything - which is how a thin layer stops being thin. Be honest, too, about what is contested. There is no evidence-backed universal number of journeys, no defensible ratio that applies across products, and the common claims about defect-cost multipliers by phase are widely repeated and weakly supported. What is defensible is the *shape of the reasoning*: risk times reversibility, uniqueness against cheaper levels, an explicit budget, and evidence-driven review. Assert the method, not a number. ### How to answer this in an interview Name the filters, name the budget in units you would actually publish, and give the two governance mechanisms that create the counter-pressure - the admission rule and the push-down habit. Then say what you would measure to know the portfolio is working: defects caught only here, escaped defects that a journey should have caught, and hours spent diagnosing. A candidate who answers with a number of tests has missed the question; a candidate who answers with a decision procedure and its feedback loop has owned it.

  • A journey has run for a year without ever failing. Is that a reason to keep it or to retire it?
    Neither on its own - it is a prompt to ask why. If it guards a path that still carries high consequence and still changes, a long green streak is the value being delivered. If the path has been frozen for a year, the journey is watching something nothing touches, and its budget slot is better spent elsewhere. Decide on the risk and the change rate, not on the failure count alone.
  • How would you measure whether this layer is earning its cost?
    Three quantities: defects caught here that no cheaper test could have caught, defects that escaped to users which a journey should have caught, and the engineer-hours spent diagnosing failures at this level. Rising escapes argue for growing the set; a high diagnosis cost with few unique catches argues for shrinking it or investing in reporting. Track them per journey where you can, because the portfolio decision is per journey.
  • A team asks to add nine journeys for a new feature. How do you respond?
    Ask which single journey carries the consequence, and what the other eight would catch that a narrower test could not. Usually one earns a slot and the rest describe behaviours that belong a level down, where they will run faster and diagnose better. If several genuinely cross different real boundaries, revisit the budget explicitly rather than letting the layer grow by default.

It is like deciding which fire drills a building actually runs: you rehearse the evacuations whose failure would be catastrophic and unrecoverable, not every door in the floor plan, and you retire a drill when the risk it rehearses is gone.

saying these in an interview costs you the question

  • Prescribing a universal number or ratio of journeys
  • Adding a journey for every new feature by default
  • Never retiring a journey once it exists
  • Quoting defect-cost multipliers as settled fact
  • Letting a group that pays no diagnosis cost own the suite
  • Judging the layer only on pipeline minutes

context