skip to content

Your team can sustain only a handful of new automated cases per release. How do you decide which behaviours get them?

level: principalimportance: should knowfreq 38%

answer

  1. It is portfolio choice, not enthusiasm
  2. Fund the recurring bill first
  3. Expected loss avoided per hour spent
  4. Defects cluster in the same areas
  5. Name what you refused, and why

basics

~20 s

Rank candidates by expected loss avoided per engineer-hour: severity and defect history, run frequency, manual cost, and expected upkeep. Reserve capacity for upkeep of existing cases before funding new ones, and record what stays manual and why.

solid answer

~50 s

I treat it as a portfolio under a fixed budget, and the budget that binds is **upkeep capacity**, not authoring time. First I reserve the hours the existing set consumes in repair and triage; whatever is left funds new cases. Then I rank candidates by expected loss avoided per hour spent: severity of the failure, likelihood inferred from defect history in that area, how often the check would run, and the expected upkeep rate given how fast that behaviour and its interface change. High severity plus stable behaviour goes first, because it is cheap to hold and expensive to lose. Low severity on a churning interface goes last or never. I write down what we chose not to automate and why, so the decision can be revisited, and I refuse a bare percentage target because it drives people to author cases with weak oracles.

code

pseudocode · 14 lines
pseudocode
reserved_upkeep_hours = 16          # measured last quarter
budget_hours          = 40 - reserved_upkeep_hours

for c in candidates:
    c.score = (c.severity * c.history_likelihood * c.runs_per_year) / (c.authoring_hours + c.upkeep_hours_per_year)

funded = []
spent  = 0
for c in sorted(candidates, key=score, descending=True):
    if spent + c.authoring_hours <= budget_hours:
        funded.append(c)
        spent = spent + c.authoring_hours
    else:
        record_manual(c, reason='below the line this quarter')

go deeper

for a junior

You will not own this call, but understand the shape: there is a fixed amount of time, cases cost time forever once written, and the ones guarding severe and stable behaviour are funded first.

for a middle

Be able to argue a single candidate's place in a queue using severity, run frequency and expected upkeep, and to say why the easiest case to write is often not the most valuable one.

for a senior

Show that you bring numbers to the conversation — measured repair hours, defect clustering by area — and that you can propose checking a severe rule through a steadier channel when its usual surface is in redesign.

for a principal

Own the budget and the refusals. Be ready to state the team's sustainable upkeep capacity, to defend a written list of what stays manual and why, and to replace a handed-down percentage target with a measure tied to defects that actually escaped.

### The constraint is upkeep, not authoring The question sounds like prioritisation of new work. It is really capacity planning, and the resource that runs out first is not the hours available to write cases — it is the hours available to keep them working. Every case added draws on that budget for as long as it exists. A team that funds authoring from the whole budget and upkeep from whatever is left ends up with a large set nobody trusts, because repairs queue and failures stop being read. So the first move is to size the recurring bill. A 4-person team on a marketplace bidding engine measured a quarter: 31 repairs across the existing set, averaging 24 minutes each, plus roughly 3.5 hours of triage across the quarter — about 16 hours, or 5% of one engineer's quarter. That is the number that goes on the wall, and it is the number to grow deliberately rather than by accident. ### Ranking the candidates With the remaining hours, I rank by *expected loss avoided per hour spent*, which needs four inputs per candidate: **Severity.** What the failure costs if it escapes. On the bidding engine, an account acting outside its role — cancelling an auction it does not own — is a different order of loss from a mis-formatted countdown. Severity here is technical and business impact, judged with whoever owns the area, not an analyser's rule level. **Likelihood, from defect history.** Defects cluster. If an area produced a permission-escalation defect two releases ago, the prior that it produces another is much higher than for an area that has been quiet for two years. Reading the last two or three quarters of reported defects and grouping them by area is the cheapest prioritisation input available, and most teams already have the data. **Repetition.** How many times the check would actually run before the behaviour is replaced. This is what amortises the authoring cost. **Expected upkeep.** How fast the rule, its data and the surface it is driven through change. This is the term that turns an attractive candidate into a standing bill, and it is why two candidates with identical severity can rank very differently. The ordering that falls out is stable across teams: severe and stable behaviours first; severe but churning behaviours next, often checked through a more stable channel than the one in redesign; low-severity stable behaviours after that; low-severity churning behaviours never. Notice that the second row is where judgement earns its keep — the temptation is to skip a high-severity behaviour because its surface is unstable, when the better answer is usually to guard the same rule somewhere steadier. ### What you deliberately do not fund A portfolio decision is only real if something is refused. Naming the refusals is what separates a lead from a list-maker: the subjective judgements that have no mechanical oracle, the one-offs, the behaviours whose expected outcome nobody has agreed, and the areas about to be rebuilt. Those go on a written list with the reason and a trigger for revisiting — the interface ships, the promotion becomes permanent, the oracle is agreed. Without the reason recorded, a later reader cannot tell whether the entry is still true, and it silently becomes permanent. ### Resisting the percentage target Somewhere above this decision, someone will ask for a number: automate 80% of the cases, or reach a coverage figure by a date. The honest response is arithmetic rather than resistance in principle. A percentage target ignores that cases are not interchangeable — the last 20% is disproportionately the material with weak oracles and unstable surfaces, which is precisely the material that produces cases failing on unchanged behaviour. And a target denominated in case count rewards authoring cheap cases with weak assertions. I offer a better measure instead: of the defects that escaped to production last quarter, how many would a proposed case have caught, and what does the set cost per quarter to hold? Those two numbers direct effort where a percentage cannot. ### Reviewing the decision The ranking is not a one-off. Defect history moves, interfaces settle, features retire. I revisit at a regular planning point with three questions: what escaped that we could have caught cheaply, what did we pay to repair and was it worth holding, and which of the manual-for-now entries have had their reason expire. That review is also where a case that has stopped earning its place gets raised — though the mechanics of retiring and quarantining cases belong to whoever owns the set's ongoing hygiene, not to this decision. ### What interviewers listen for They want to hear that you count the recurring cost before the new work, that your ranking has inputs rather than instincts, that you can name what you refused and why, and that you would push back on a number handed down without a cost model behind it. Certainty is the wrong register here: the good answer is a defensible process, stated as a process, with the places it can be wrong named out loud.

  • A high-severity behaviour sits behind an interface being rebuilt this quarter. Do you fund it?
    Usually yes, but not through the surface in redesign. The rule is what matters, so I look for a steadier channel to exercise the same rule and check it there; the case then survives the rebuild. If no steadier channel exists, I keep it manual with a named owner and a revisit trigger tied to the rebuild landing, rather than paying to author a case that will be rewritten within weeks.
  • Leadership sets a target percentage of cases to be automated. How do you respond?
    With the cost model rather than a refusal. I show what the existing set costs per quarter to hold, and point out that the remaining candidates are disproportionately the ones with weak oracles and unstable surfaces — so hitting the number produces cases that fail on unchanged behaviour and erode trust. Then I offer a measure that actually tracks the goal: what escaped last quarter, and which proposed cases would have caught it.
  • How do you use defect history without simply chasing the last incident?
    Group reported defects by area over two or three quarters rather than reacting to the newest one, and look for clustering. One incident is an anecdote; a cluster is a prior worth spending on. I also weight by severity, because a run of low-impact defects in one area can outnumber a single severe one while mattering less. The output is a ranking input, not a mandate to automate whatever failed most recently.

saying these in an interview costs you the question

  • Funds new cases without reserving upkeep capacity
  • Accepts a percentage target with no cost model
  • Ranks by ease of automating rather than value
  • Chases the most recent incident as the priority
  • Cannot name anything the team chose not to automate
  • Measures the effort by number of automated cases

context