Your team has a backlog of recurring manual operations tasks and limited engineering time. How do you decide which ones to automate, which to leave manual, and which to remove the need for entirely?
answer
- eliminate, automate, or keep manual
- frequency times time, minus maintenance forever
- risk can outweigh the arithmetic
- rare and destructive stays human-triggered
- self-service before a closed loop
basics
~20 sRank by payoff against cost: how often the task occurs, how long it takes, and what it risks, weighed against building and maintaining the automation forever. Always ask first whether the task can be eliminated — fixing the cause beats automating the symptom.
solid answer
~50 sI start with three columns, not two. Eliminate, automate, or keep manual. Elimination is checked first and is usually underused: many recurring tasks exist because of a defect, a missing default, or a design decision, and removing the cause beats automating the workaround permanently. For what remains, the payoff is roughly frequency times handling time over the automation's lifetime, weighted by risk — a task done monthly but capable of causing an outage can justify automation that pure time arithmetic would not. Against that sits the real cost: build, plus maintenance forever, plus the on-call burden of a new production actor. What I deliberately leave manual is the rare, high-blast-radius action. Automation nobody has run in eight months is stale and untrusted at 3am. For those I build the tool but keep a human on the trigger, and rehearse it. And I prefer self-service over closed loops when the requester can safely act themselves.
go deeper
Be able to say that automation is worth it when a task is frequent and repetitive, and that some tasks are better removed at the source than automated.
Work the arithmetic out loud — frequency times handling time against build plus maintenance — and note that maintenance is an ongoing cost most estimates leave out.
Show where you would refuse to automate: rare destructive actions, judgement-heavy decisions, and procedures the team does not yet fully understand, and explain that unexercised automation is untrusted automation.
Treat it as a portfolio: every automation has an owner, a maintenance tax and a decommission path, self-service usually beats a closed loop, and repeated manual work in one category is a signal to fix the platform rather than write another script.
## Three outcomes, not two The framing that separates a strong answer from a weak one is that "automate it" is one of three options. **Eliminate.** A surprising share of recurring operational work exists because something upstream is wrong: a service that needs a weekly restart, a certificate nobody automated the renewal of, a config that must be hand-copied because two systems disagree on a default, a capacity ceiling that requires manual intervention every launch. Automating these makes them cheap and permanent. Removing the cause makes them zero. Elimination should be checked first for every item on the list, because it is the only option with no ongoing cost. **Automate.** The task must genuinely recur, the decision inside it must be expressible, and the payoff must beat the total cost of ownership. **Keep manual, but make it good.** Documented, rehearsed, and supported by tooling that does the dangerous parts — while a human retains the decision. ## The arithmetic, and its limits The naive model is: `frequency x handling time x horizon` versus `build cost + maintenance`. A task consuming two hours a week is roughly a hundred hours a year, which comfortably justifies a week of engineering. A task consuming two hours a quarter does not. Two corrections matter. First, **maintenance is not free** and is routinely omitted: every automation is a system with dependencies, credentials and assumptions that rot. Budget for it as an ongoing tax, not a one-off. Second, **risk is a multiplier**. A rare task that can cause an outage when performed badly — a failover, a schema change, a certificate rotation — may deserve tooling out of proportion to the time it consumes, because the value is consistency and blast-radius control rather than hours saved. ## Why some tasks should stay manual The senior half of this answer is knowing where to stop. - **Rare plus destructive.** Automation exercised twice a year is stale automation. Its dependencies have moved, its credentials have expired, and the person who wrote it has left. At 3am nobody trusts it and someone does it by hand anyway — badly, under pressure, having also lost the practice. Prefer a rehearsed procedure plus a tool that performs the fiddly steps under human control. - **Judgement-heavy.** If the hard part is deciding *whether* to act rather than *how*, automate the how and keep the whether. Most bad auto-remediation stories start with automating a decision that needed context. - **Irreversible.** Anything that destroys data or capacity in a way you cannot undo deserves a human confirmation for a long time, regardless of frequency. - **Not yet understood.** Automating a procedure you do not fully understand encodes the misunderstanding and makes it fast. ## Self-service is usually the best value Before reaching for a closed loop, ask who is asking. If the work arrives as requests from other teams, the highest-leverage move is not to make your execution faster but to remove yourself from it: expose the action with its safety checks built in so the requester performs it. This is far cheaper than a control loop, it scales with demand rather than with your headcount, and it forces you to build exactly the guardrails a closed loop would later need. Beware the anti-pattern of moving toil rather than removing it — automation that makes your team's work disappear by pushing manual steps onto another team is not a win, it is an accounting trick. ## Keeping rarely used automation honest If you do automate something rare, it must be exercised on a schedule; otherwise its next run is its first test, during an incident. Run it periodically against a non-production target, or fold it into a rehearsal, and treat a failed exercise with the same seriousness as a failed deploy. Automation you cannot demonstrate working this quarter should be assumed broken. ## The organisational view At scale this becomes a portfolio question. Every automation has an owner, a maintenance cost and a decommission path; unowned automation with production credentials is a liability, and periodically retiring automation whose task no longer exists is part of the job. It is also worth measuring: if the same category of task keeps producing new manual work, the answer is upstream in the platform, not in one more script. ## How to answer in an interview Give the three-column framing, do the arithmetic out loud with a concrete example, then spend the rest of the answer on where you would deliberately not automate and why. Ending on "eliminate first, self-service second, close the loop only where frequency and confidence justify it" shows you have owned this portfolio rather than read about it.
- How do you account for the maintenance cost when you are pitching an automation project?Estimate it as a recurring tax rather than a one-off — a rough rule is that the automation will need attention whenever its dependencies, credentials or the underlying procedure change, which for most operations tooling is several times a year. State it explicitly in the proposal alongside the saving. Automation pitched with only a build cost is how teams accumulate a fleet of scripts nobody has budget to maintain.
- What do you do with automation whose underlying task no longer exists?Decommission it deliberately. Unowned automation holding production credentials is a standing risk with no offsetting value, and it will eventually fire on a condition nobody remembers. Reviewing the inventory periodically — what does this do, who owns it, when did it last run successfully — and retiring the dead entries is part of running the portfolio.
- A task is frequent but each run requires a judgement call. How do you handle it?Split it. Automate the mechanical parts — gathering the data, checking preconditions, executing the steps once chosen — and present the human with the decision and the evidence. That removes most of the elapsed time and nearly all of the execution risk while keeping the judgement where it belongs. Automating the decision itself is where most bad remediation stories begin.
saying these in an interview costs you the question
- If it repeats, automate it, always
- Maintenance of the automation is free after launch
- Rare dangerous tasks are the best automation candidates
- Moving the manual steps to another team counts as removing toil
- Never revisit whether an automation is still needed