skip to content

Which checks do you deliberately leave manual instead of automating, and why?

level: middleimportance: must knowfreq 66%

answer

  1. Four families, each with its own reason
  2. Nothing to amortise a one-off against
  3. Taste and wording have no mechanical oracle
  4. Nobody agrees the expected outcome
  5. Manual for now, with a revisit date

basics

~20 s

Checks that run once, that need a person to judge look, wording or feel, that have no dependable expected result, and that drive an interface still being redesigned stay manual. Each of those costs more to automate than it returns.

solid answer

~50 s

Four families stay manual. **One-offs** — a migration verified before a single launch has no repetition to amortise authoring against. **Subjective judgements** — whether wording reassures, whether a layout reads well, whether a first-time user can find a control; there is no mechanical oracle, only a person. **Checks with a poor oracle** — where nobody can state what the right outcome is, or the outcome legitimately varies run to run; encoding a guess produces a case that either never fails or fails randomly. **Unstable interfaces** — where the screen or contract being driven is mid-redesign, so the case will be rewritten before it has run enough times to repay. I say "for now" rather than "never": when the interface settles or the check starts repeating weekly, I revisit it. And I still automate the arrangement around a manual judgement even when the judgement stays human.

go deeper

for a junior

Learn the four families by heart — one-off, subjective, no dependable expected result, unstable interface — and be able to give one concrete example of each rather than reciting the labels.

for a middle

Explain the mechanics: why a one-off has nothing to amortise against, what an oracle is and how it fails, and why an interface mid-redesign forces rewrites before a case has repaid. Expect to be pushed on a borderline example.

for a senior

Show that you make the decision explicit and reversible — a recorded reason and a revisit trigger per manual check — and that you split arrangement from judgement so the tedious setup gets automated even when the verdict stays human.

for a principal

Own the policy: which classes of check the organisation deliberately never automates, how that is communicated so it does not read as laziness, and how you resist a percentage target that pushes teams to automate checks with no oracle.

### "Not automated" is a decision, not a failure Interviewers ask this question because the reflex answer — "automate everything" — is expensive and common. A strong answer names the families of check that stay manual and gives the reason for each, because the reason is what generalises to a case the interviewer has not mentioned. ### One-offs and near-one-offs Authoring cost is amortised across runs. A check that runs once has nothing to amortise against. The classic instance is a data migration verified the day it happens: careful, high-stakes, and never repeated in that form. Near-one-offs are the trap — a check expected to run "a few times" during an eleven-day promotion looks repetitive but will not accumulate enough runs before the promotion ends. The test is the count over the *horizon*, not the presence of repetition in principle. The exception worth naming: a throwaway harness written to *find* a defect rather than to guard against one. Driving several hundred generated inputs at a rule to hunt for a bad case is automation whose payback lands inside a single sitting, and it is fine to delete afterwards. That is a different economic decision from adding a case to a set that runs forever. ### Subjective judgements Some outcomes are only meaningful to a person: whether copy reassures a nervous buyer, whether a layout is pleasant, whether an error message helps rather than blames, whether the flow feels fast. There is no oracle a machine can evaluate here — the best it can do is detect that pixels changed, which is not the same claim and generates failures for every intentional redesign. Usability and aesthetic judgement stay with people, and pushing them into automation produces cases that either assert nothing or assert the wrong thing. ### A poor or missing oracle An oracle is the thing that says the behaviour is wrong. It can fail in two ways. First, nobody can state the expected outcome: the rule is undocumented, contested between two people, or genuinely "whatever the system does today". Encoding today's behaviour then locks in whatever defect is currently present, and every future correction shows up as a red case. Second, the outcome legitimately varies — an ordering that is not specified, a value that depends on the moment, a recommendation that is allowed to differ. A case built over a varying outcome either asserts something so loose it can never fail or fails on unchanged behaviour. Get the oracle agreed first; then decide about automating. ### Unstable interfaces When the surface a case drives is mid-redesign, the case is a liability: it will be rewritten before it has run enough times to repay. The right move is to wait for the surface to settle, or — where possible — to drive the same rule through a more stable channel and check it there. "Wait" is a real answer, not a cop-out, provided somebody has said when to look again. ### A worked example A 4-person team on a marketplace bidding engine sorted a backlog of 47 candidate checks and left 12 manual. Two were launch-day migration verifications. Five were judgements about the redesigned bid-confirmation screen — clarity of the countdown, tone of the outbid message. Three were checks over a recommendation ranking whose expected order nobody would commit to. Two drove a seller console being rebuilt over the next quarter. Against those, the team automated the permission rules around cancelling an auction immediately: exact oracle, stable rule for 21 months, and a permission-escalation defect in that area two releases earlier. Notice what the split produced. It did not produce twelve untested behaviours — it produced twelve behaviours checked by a person, deliberately, with the reason recorded next to each. The record matters, because the reasons expire: an interface settles, a promotion becomes permanent, an oracle finally gets agreed. A list of manual-for-now checks with no revisit date silently becomes a list of things nobody looks at. ### Automate the arrangement, keep the judgement Even inside a manual check, most of the minutes are mechanical: creating accounts with the right roles, seeding auctions in the right state, moving the clock, capturing evidence. Automating that arrangement can turn a 14-minute manual check into a 2-minute one and reduces the human error rate of the setup. The judgement stays with a person; the tedium does not. Candidates who answer only "automated or manual" miss this, and it is often the highest-return move available.

  • How do you stop a manual-for-now list from becoming a list nobody ever revisits?
    Record the reason next to each entry, not just the decision, and give it a trigger for review — the interface ships, the promotion becomes permanent, the oracle gets agreed. The reasons are what expire, so a list of reasons can be re-checked cheaply at a planning point. Without the reason, a later reader cannot tell whether the entry is still true and will leave it alone forever.
  • Someone proposes automating a check whose expected outcome nobody can state. What do you do?
    Stop and get the oracle agreed before writing anything. Encoding current behaviour as the expectation locks in whatever defect exists today and turns every future correction into a red case. The productive move is to take the disagreement to whoever owns the rule; often the argument itself surfaces the defect, which is more valuable than the case would have been.
  • Is a throwaway script that drives hundreds of inputs at a rule automation worth writing?
    Yes, but it is a different decision. That script is a tool for finding a defect now, and its payback lands inside a single sitting, so it does not need to repay upkeep over years. It is fine to delete it afterwards. Confusing it with a case added to a set that runs forever is what leads teams to carry hundreds of cases nobody would author today.

saying these in an interview costs you the question

  • Answers that everything should be automated eventually
  • Calls a check automated when it asserts nothing that can fail
  • Encodes today's behaviour as the expected outcome
  • Treats manual as a permanent verdict with no revisit
  • Automates against a screen being redesigned this quarter
  • Claims a machine can judge wording or layout quality

context