Which properties make a test case a good candidate for automation rather than manual running?
answer
- Automation is an investment, not a virtue
- Count the runs over a horizon
- Is the behaviour still being redesigned?
- Can a machine judge the outcome?
- Severe defect history lowers the bar
basics
~20 sA case pays back when it repeats often, the behaviour it checks is stable, its expected result can be judged mechanically, and running it by hand is slow or error-prone. High risk and a defect history raise its value further.
solid answer
~40 sAutomation is an investment, so I look for the properties that make it repay. **Repetition** first: a check re-run every release earns its cost many times, a one-off never does. **Stability of the behaviour** second: if the rules the check encodes are still being redesigned, I will pay to rewrite the case before it has run enough times. **A mechanical oracle** third: something must be able to say the behaviour is wrong without a human judging taste or wording. **Manual cost** fourth: a check needing eleven minutes of careful setup by hand is worth more automated than one that takes twenty seconds. Finally **risk and defect history** — an area that has produced severe defects before repays automation at a much lower repetition count, because the loss avoided is larger.
go deeper
Be ready to name the properties plainly: it repeats, the behaviour is stable, the expected result can be checked mechanically, and doing it by hand is slow or error-prone. Give one example of a case you would leave manual.
An interviewer expects you to explain the mechanics of payback — authoring, upkeep, run time and triage on the cost side, manual effort saved and defects caught earlier on the return side — and to weigh two candidate cases against each other out loud.
Show production judgement: use defect history and severity to argue that a high-risk check repays at a much lower repetition count, and point out where a team is paying upkeep on cases that guard behaviour nobody is changing any more.
Own the framing that automation is a funded investment with a recurring cost, not a target percentage. Be ready to say which classes of check your organisation deliberately never automates, and how that decision is written down and revisited.
### Automation is an investment, not a virtue The question behind "what should we automate?" is never "what *can* be automated?" — with enough effort almost anything can. It is "which cases return more than they cost?" A candidate has four cost components, and a weak answer only ever counts the first: 1. **Authoring** — designing the case, building its data and preconditions, making it run in a pipeline, and getting it reviewed. 2. **Upkeep** — repairing the case every time the behaviour, the data or the interface it drives changes. 3. **Run cost** — the minutes it adds to every run of the set it belongs to, multiplied by how often that set runs. 4. **Triage** — the human time spent reading each failure to decide whether it is a real defect. Against that sits the return: the manual effort it removes, plus the value of catching a regression earlier and more reliably than a person would. ### The five properties that predict payback **Repetition over a horizon.** Count the runs the case will actually get before the behaviour it checks is replaced. A check re-run on every change to a long-lived rule may run several hundred times; a check for a launch-day migration runs once. The horizon matters as much as the rate: a case that repays after nine months of running, guarding a feature already scheduled for replacement in four, is a loss no matter how often it runs in between. **Stability of the behaviour.** The relevant churn is in the *rule being checked*, not the code around it. A pricing rule that changes twice a quarter forces a rewrite of every case that encodes it. A rule that has held for two years is cheap to guard. Interfaces count here too: a screen being redesigned this quarter will invalidate the way the case drives it, even when the underlying rule is untouched. **A mechanical oracle.** The oracle is the thing that says the behaviour is wrong. Automation needs one that a machine can evaluate — a status, a stored record, a computed total, a rejected request. Where the judgement is "does this read well", "is this layout attractive", "would a first-time user find this", no mechanical oracle exists, and no amount of repetition rescues the case. **Manual run cost.** The comparator is not the abstract idea of a manual test but the real cost of one honest manual run: setup, careful observation, and the error rate of a bored person doing it for the fortieth time. Checks that are tedious and easy to get wrong by hand are exactly the ones a machine does better. **Risk and defect history.** Expected loss avoided, not run count alone, is the return. An area that has already produced severe defects is likely to produce more, and the cost of each is high — so it repays automation at a far lower repetition count than a cosmetic check that repeats twice as often. ### A worked example On a marketplace bidding engine, a 4-person team weighs two candidates. The first: a bidder who does not hold the seller role must not be able to cancel a live auction. The rule has not changed in 21 months, the outcome is a rejected request with a specific error, and this area produced a permission-escalation defect two releases ago that let 63 auctions be cancelled by the wrong account. Authoring is estimated at 3 hours 40 minutes; the team ships 14 times a year and would run it on every change, roughly 210 times a year in the pipeline. It is an easy yes — the repetition is high, the oracle is exact, the rule is stable, and the loss avoided is severe. The second: confirming that the new bid-confirmation screen "feels responsive and reassuring" during a promotion running for eleven days. It runs a handful of times, the interface is being redesigned mid-promotion, and the oracle is human judgement. It stays manual, and the team spends the saved hours on the first case. ### Partial automation The decision is rarely all-or-nothing. Even where the judgement must stay human, the expensive arrangement around it — creating accounts, seeding auctions, putting the system in the right state, capturing evidence — is repetitive, mechanical and worth automating. A good answer says which *part* of a check is being automated, not just which check.
- Two candidate cases repeat equally often. What would make you automate one and not the other?The tie is broken by risk and by expected upkeep. If one guards a rule whose failure would be severe — money moving wrongly, a permission being escalated — and the other guards a cosmetic detail, the first returns far more per hour spent. Upkeep is the other half: a case whose behaviour or interface is about to be redesigned will need rewriting before it has run enough times to repay, so equal repetition does not mean equal payback.
- A check has no mechanical oracle. Is there anything left worth automating about it?Usually yes — the arrangement, not the judgement. Creating the accounts, seeding the data, driving the system into the state where the judgement is made, and capturing artefacts for a person to review are all repetitive and mechanical. Automating those turns a fifteen-minute manual chore into a two-minute one and leaves the human doing only the part that needs a human.
Automating a check is like buying a machine for a workshop rather than doing the job by hand: it is worth it when the same job comes back often, the job is not about to change shape, and you can tell mechanically whether the machine got it right.
saying these in an interview costs you the question
- Says everything should eventually be automated
- Counts authoring cost but never upkeep or triage
- Automates a one-off check because it was tedious once
- Ignores that the behaviour is still being redesigned
- Treats automated case count as the goal
- Assumes a human judgement can be encoded as a mechanical check