A recurring production fix currently runs as a script an on-call engineer executes by hand after being paged. What do you gain, and what do you take on, by promoting it to closed-loop automation that runs with no human in the loop?
answer
- each rung trades effort for autonomy
- speed and consistency, bought unsupervised
- you inherit a new production actor
- rate limit, audit trail, named owner
- close the loop last, not first
basics
~20 sClosing the loop buys speed and consistency: the fix runs in seconds, identically, without waking anyone. In exchange you take on a new production actor that changes systems unsupervised, so it needs guardrails, its own telemetry, and a named owner.
solid answer
~50 sYou gain three things: latency (the fix happens in seconds rather than after someone wakes up and reads a document), consistency (the same steps every time, including at 3am), and a reclaimed on-call night. What you take on is a system that changes production without supervision, and it deserves the same rigour as the service it repairs — tests, a staged rollout, an audit trail of every action, a rate limit so one bad signal cannot become a thousand actions, and an owner who gets paged when it misbehaves. The subtler cost is that you removed the human who used to notice the fix was needed unusually often. So the automation has to report on itself: a metric or ticket per run, and an alert when the run rate rises. I would normally climb one rung at a time — script it, make it self-service, watch it run under supervision, then close the loop.
go deeper
Be able to name the rungs in order — manual steps, a written procedure, a script, a self-service tool, then automatic remediation — and say that each one hands more of the work to the machine.
Explain what each rung buys and what it costs to build, and be specific that unattended automation needs its own tests, logging and limits because no human is watching it act.
Show the judgment that the top rung is optional. Talk about validating a remediator in observe-only mode, bounding its blast radius, and instrumenting it so a rising action rate becomes a page.
Own the argument that self-service is usually the highest-leverage rung, and be ready to defend where your organisation stops climbing: which actions stay human-triggered, who owns each automation, and how you prevent unowned robots accumulating production access.
## The ladder Operations work is usually described as climbing rungs, each one moving a little more of the decision from a person to a machine: 1. **Undocumented manual steps.** One or two people know how to fix it. Everything depends on who is awake. 2. **A documented procedure.** Anyone on the rotation can execute it, but every step is still a human action, and the document decays between uses. 3. **A script a human runs.** The steps are now executable and consistent; the human still decides *whether* and *when*. 4. **A self-service tool.** The action is exposed so the people who need it — including other teams — can trigger it themselves, without filing a ticket to you. The decision stays human; the execution and the safety checks are machine-owned. 5. **Closed-loop auto-remediation.** A control loop observes a condition and acts on it with no human involved at all. Each rung costs more to build and more to trust than the one below it. The interview point is that the top rung is not automatically the goal. ## What each rung actually buys Going from a document to a script removes *variance*: the tired engineer no longer skips step 4. Going from a script to a self-service tool removes *you* from the critical path — this is often the highest-leverage rung and the most under-built one, because it converts your team from an execution bottleneck into a platform. Going from self-service to closed loop removes *human latency*, which only matters when the condition recurs often, when minutes of exposure are expensive, or when it recurs at 3am. If the condition happens twice a year, the closed loop buys you almost nothing and costs you a piece of unattended machinery that will be stale the next time it fires. ## What closing the loop costs When the loop closes, the automation becomes a production system in its own right: - **It needs to be operable.** Someone must be able to answer, at 3am and half asleep, "is the automation doing this, and how do I stop it?" That means an off switch that is documented and has actually been used. - **It needs an audit trail.** Every action, with a timestamp and the input that triggered it. Without that, a future incident starts with the unanswerable question "did a human do that, or did the robot?" - **It needs limits.** A rate limit (at most N actions per window) and a blast-radius cap (never act on more than X% of the fleet) convert a runaway from an outage into a page. - **It needs to be trustworthy under a broken input.** The remediator's decision is only as good as the signal it reads. Ask what it does when the signal itself fails — the correct answer is usually "stop and escalate", not "act on everything". - **It needs an owner.** Automation with no owner is abandoned code with production credentials. ## The two failure modes you inherit **Masking.** The human who ran the script by hand implicitly tracked how often it was needed. The machine does not, unless you make it. A remediator that quietly repairs a worsening fault buys time and spends it invisibly, until the fault outgrows the remediation. **Amplification.** A human executing a fix on a bad signal does it once, notices it did not help, and stops. A loop does it continuously. Whatever bounded the damage before was a person's judgement, and closing the loop deletes it; the rate limit and the health precondition are what you put back in its place. ## Safeguards that make the top rung defensible - A **dry-run** path used in staging and in review, so the action's scope is visible before it is real. - **Idempotency** — repeating the action converges instead of compounding, because retries and overlapping runs are normal. - A **rate limit and a circuit breaker on the automation itself**: after N actions in a window, or N failed attempts, the loop stops acting and pages instead. - **Self-reporting**: a counter per action, and an alert on the *rate* of remediation, not only on the unremediated symptom. - **A staged rollout**: run in observe-only mode first, comparing what it would have done against what humans actually did. ## How to answer in an interview Name the ladder, then show that you treat the top rung as a decision rather than a destination. The strongest version is: "I would make it self-service first, keep it human-triggered while we watch the trigger rate, and close the loop only once we know the signal is reliable and the action is bounded." That answer demonstrates that you have owned automation after it shipped, not just written it.
- Which rung do teams most often skip, and why does that hurt them?The self-service rung. Teams jump from "we run the script for you" straight to a closed loop, or stay stuck answering tickets. Exposing the action so requesters can run it themselves removes your team from the critical path immediately, at a fraction of the cost of a control loop, and it surfaces the safety checks you will need later anyway.
- How would you validate a new remediator before letting it act on production?Run it in observe-only mode for a couple of weeks: it evaluates its trigger and logs the action it would have taken, but does not take it. Then compare its decisions against what on-call actually did. That measures its false-positive rate on real traffic, which no staging environment will tell you, and it gives you a defensible threshold to launch at.
- When is a closed loop the wrong answer even though the task is clearly automatable?When the action is irreversible or has a huge blast radius — deleting data, failing over a region, terminating a stateful cluster — and when the condition is rare. Rare automation rots and nobody trusts it under pressure. For those, build the tool but leave a human on the trigger, and rehearse it.
saying these in an interview costs you the question
- Automate everything the moment it repeats twice
- Closed-loop automation does not need monitoring of its own
- The script is finished once it works by hand once
- Nobody has to own the automation after it ships
- Removing the page is the only goal that matters