How do you estimate whether an automated case will repay the cost of writing and maintaining it?
answer
- Turn the opinion into an estimate
- Cost side has a recurring term
- The saving per run is a difference
- Runs before the behaviour is replaced
- Upkeep capacity is the real ceiling
basics
~20 sCompare authoring plus recurring upkeep and triage against the manual cost per run multiplied by the runs expected before the behaviour changes. The break-even run count, and whether the horizon reaches it, decides. Most estimates omit upkeep.
solid answer
~50 sI build a crude model out loud rather than argue from principle. Cost = authoring (design, data, making it run in the pipeline, review) + upkeep per period + triage per failure. Return = honest manual cost of one run x runs expected over the horizon, plus the value of catching the regression earlier. Break-even run count is roughly authoring divided by the saving per run, and the decisive question is whether the behaviour survives long enough to reach it. The two terms people forget are **upkeep** — which never stops and scales with how often the behaviour changes — and **triage**, the human minutes spent on every failure including the false ones. I present the numbers as an estimate with a range, and I say explicitly that the figures often quoted for automation return on investment are contested rather than measured.
code
pseudocode · 9 linesauthoring_minutes = 220
manual_minutes = 11.0
triage_per_run = 0.4
saving_per_run = manual_minutes - triage_per_run
break_even_runs = authoring_minutes / saving_per_run
runs_over_horizon = 190
print(break_even_runs) # about 20.8 runs
print(runs_over_horizon > break_even_runs) # true -> the case repaysgo deeper
Know that an automated case costs more than the hours spent writing it, and that it has to be run many times to be worth it. Being able to say why a one-off does not repay is enough at this level.
Be able to write the comparison down: authoring plus upkeep plus triage against manual cost times runs. Explain why the per-run saving is a difference rather than the full manual time.
Demonstrate that you use real numbers from your own set — measured repair rates, measured triage minutes — rather than invented ones, and that you can say which assumption the conclusion is most sensitive to.
Own the aggregate: the team's sustainable upkeep capacity is the real constraint, and every case added draws on it forever. Be ready to use that number to push back on a coverage percentage handed down as a target.
### Why a number beats an opinion "We should automate this" and "that is not worth automating" are equally unfalsifiable. A rough model turns the argument into one about inputs, which is a much better argument to have: someone can dispute your estimate of upkeep, and disputing it produces information. ### The model For a single candidate case, over a chosen horizon: ``` cost = authoring + (upkeep_per_period * periods) + (triage_per_failure * expected_failures) return = (manual_minutes_per_run * runs_over_horizon) + value_of_earlier_detection ``` And the break-even run count, ignoring upkeep for a moment: ``` break_even_runs = authoring / (manual_minutes_per_run - automated_triage_per_run) ``` Four things about this model matter more than its arithmetic. **The denominator is a difference, not the manual cost.** An automated run is not free of human time: someone reads the result when it goes red. If a case fails spuriously often, the triage term can eat most of the saving and push break-even out beyond the horizon. **Upkeep is a recurring subscription.** It is the term candidates leave out, and it is the one that decides most real cases. Upkeep scales with how often the behaviour, its data or its interface changes, not with how often the case runs. A case guarding a rule that changes twice a quarter is a standing bill; a case guarding a rule untouched for two years is nearly free after authoring. **Runs over the horizon, not runs in principle.** The horizon ends when the behaviour is replaced, the feature is retired, or the interface is rebuilt — whichever comes first. A case that repays after 180 runs, on a behaviour that will be redesigned after 60, loses money however elegant it is. **Earlier detection is real but not in the arithmetic.** Catching a permission regression in eight minutes rather than during a manual pass three weeks later is worth something, and for severe classes of defect it is worth more than the whole labour saving. Add it as an explicit risk term rather than pretending the labour arithmetic alone justifies the case. Be careful with the numbers you attach: the popular multipliers for how much more a late defect costs are contested and their original basis is disputed, so cite them as an argument for direction, not as a measured constant. ### A worked example A 4-person team on a marketplace bidding engine weighs one candidate: a check that an account without the seller role cannot cancel a live auction — the area that produced a permission-escalation defect two releases ago. - Authoring, including seeding two accounts and a live auction and getting it running in the pipeline: **3 h 40 m**. - Manual cost of one honest run, including setup: **11 minutes**, and the team admits that in practice it gets skipped under time pressure. - Triage per automated run, averaged over green and red: **0.4 minutes**. - Expected upkeep: the rule has been stable 21 months, so **about 25 minutes per quarter**. - Runs over the horizon: the pipeline runs it on every change, roughly **190 times** in the first year. Break-even on labour alone: 220 minutes / (11 - 0.4) is about **21 runs** — reached in the first fortnight. Upkeep of 25 minutes a quarter is trivial against a saving of about 10.6 minutes per avoided manual run. Add the severity of the regression it guards, and it is not a close call. Now the same team's second candidate: a check on the redesigned bid-confirmation screen, authoring 2 h 10 m, manual cost 3 minutes, running perhaps 9 times before the promotion ends, on a screen being rebuilt. Break-even is around 43 runs against a horizon of 9. The arithmetic says no, and now the team has something better than a taste argument to say so. ### Arguing against automating everything The strongest version of the argument is arithmetic, not principle. Suppose a 4-person team can sustain, say, six hours a week of automation upkeep in total. Every case added draws on that fixed budget forever. Once the budget is exceeded, the marginal case is not free — it is paid for out of the repair time of existing cases, which then start failing for reasons nobody chases. That is how a set of cases stops being trusted. So the real constraint on "automate everything" is not authoring capacity but *upkeep capacity*, and it is the number to put in front of whoever set the percentage target. ### Presenting the estimate Give a range, name your assumptions, and say which input the answer is most sensitive to — usually upkeep or the horizon. An estimate that says "this repays in about three weeks unless the interface is rebuilt before then, in which case never" is far more useful to a decision-maker than a single confident number.
- Which input in that estimate is usually the least reliable, and how do you handle it?Upkeep. It depends on how often the behaviour, its data and its interface change, and nobody knows that in advance. I handle it by taking a rate from the cases already in the set — how many needed repair last quarter and how long each took — instead of guessing, and by presenting the result as a range with the upkeep assumption stated. If the decision flips within that range, the honest answer is that it is a marginal case.
- How does the value of catching a defect earlier enter the estimate?As an explicit risk term, separate from the labour arithmetic: expected frequency of the regression times what an escape would cost. For a severe class such as a permission escalation, that term dominates and the labour saving becomes almost irrelevant. I keep it separate and I avoid quoting the popular late-defect cost multipliers as fact — their basis is contested, so they support the direction of the argument, not a specific number.
- Your estimate says a case does not repay, but the team wants it anyway. What now?Ask which input they disagree with. Usually it is the horizon — they expect the behaviour to outlive my assumption — or the risk term, which the arithmetic underweights. Both are legitimate and both are checkable. The estimate's job is to move the argument from taste to inputs; if a stated assumption is wrong, I update it and the answer changes honestly.
Buying a machine for a workshop has a sticker price and a service contract. Teams argue about the sticker price and forget the service contract, which is the payment that never stops and decides whether the machine was worth it.
saying these in an interview costs you the question
- Counts authoring cost and ignores recurring upkeep
- Uses manual run time as the whole per-run saving
- Assumes the behaviour will be checked forever
- Quotes a late-defect cost multiplier as measured fact
- Treats a single point estimate as certain
- Ignores triage time on failures, including false ones