You are building automation that terminates and replaces any instance whose health check is failing. A monitoring bug briefly reports every instance in the fleet as unhealthy. What safeguards stop that from destroying the fleet?
answer
- the signal can be wrong too
- everything unhealthy means the check broke
- bound the rate and the blast radius
- a breaker that pages instead of continuing
- observe-only before it gets authority
basics
~20 sBound the automation before it acts: refuse to act when an implausible share of the fleet looks unhealthy, cap actions per time window, cap how much capacity may be missing at once, and trip a breaker that pages a human instead of continuing.
solid answer
~50 sThe core insight is that a remediator is only as trustworthy as the signal it reads, so it must treat an implausible signal as a signal failure rather than as a fleet failure. The first safeguard is a global health precondition: if more than some share of the fleet — say 20% — looks unhealthy at once, do nothing and page, because one instance being sick is likely, all of them being sick simultaneously is almost always a broken check. Then bound the action itself: a rate limit of at most N replacements per hour, a concurrency cap so capacity never drops below what serves peak traffic, and a circuit breaker that trips after repeated failures or repeated actions and escalates instead of continuing. Add an audit log of every decision and an off switch that on-call has actually used. What a human did implicitly — notice it was not helping and stop — has to be encoded.
code
bash · 27 lines#!/usr/bin/env bash
# Replace one unhealthy node - only if the fleet-wide signal is believable.
set -euo pipefail
inventory=/var/lib/fleet/nodes # one line per node, state in the last field
actions=/var/lib/fleet/actions
node=$1
total=$(wc -l < "$inventory")
bad=$(grep -c ' unhealthy$' "$inventory" || true)
recent=$(grep -c "^$(date +%Y%m%d%H) " "$actions" 2>/dev/null || echo 0)
# 1. Implausible signal: if a fifth of the fleet is "sick", suspect the check.
if [ "$(( bad * 100 / total ))" -ge 20 ]; then
logger -p daemon.err "remediator: $bad/$total unhealthy - refusing to act"
exit 1
fi
# 2. Rate limit: at most two replacements per hour, then escalate.
if [ "$recent" -ge 2 ]; then
logger -p daemon.err "remediator: hourly action limit reached - escalating"
exit 1
fi
# 3. Audit trail first, action second.
echo "$(date +%Y%m%d%H) replace $node" >> "$actions"
logger -p daemon.warning "remediator: replacing $node"go deeper
Know that automation acting on a wrong input can do far more damage than a person would, and that limiting how many actions it may take is the basic protection.
Explain the specific guards — a plausibility check on the fraction of the fleet reporting unhealthy, a per-window rate limit, and a persistence requirement before the loop acts at all.
Show that you would make the loop doubt its own input, cap the capacity it may remove against the fleet's redundancy headroom, trip a breaker when its actions are not improving things, and roll it out in observe-only mode to measure false positives.
Set the bar for the whole estate: no unattended actor gets destructive authority without a plausibility precondition, a bounded blast radius, an audit trail and an exercised off switch, and be ready to say who reviews new remediators against that bar.
## The failure mode being defended against A human executing a replacement on a bad signal replaces one instance, sees no improvement, and stops to think. A control loop replaces one instance, observes that the fleet is still unhealthy, and replaces another — forever, at machine speed. The judgement that bounded the damage was never written down anywhere, and closing the loop deleted it. Every safeguard below is an attempt to write that judgement back into code. ## Safeguard 1: distrust an implausible signal The most valuable guard is a precondition on the *shape* of the input rather than on any single reading. One instance failing its check is an ordinary event. Ninety percent of instances failing simultaneously is not a fleet event at all — it is overwhelmingly likely to be a broken check, an expired credential the health endpoint depends on, a DNS failure in the checker, or a bad deploy of the check itself. So the loop's first decision is: *is this signal believable?* If the unhealthy fraction exceeds a threshold, the correct action is to take no action and escalate. This inverts the naive design, where more evidence of failure means more aggressive remediation; here, too much evidence of failure means stop. ```bash if [ "$(( bad * 100 / total ))" -ge 20 ]; then logger -p daemon.err "remediator: $bad/$total unhealthy - signal implausible, refusing to act" exit 1 fi ``` ## Safeguard 2: rate limit the actor Even with a believable signal, cap how fast the loop may act: at most N replacements per hour, or per rolling window. The limit is derived from how quickly replacements can genuinely be needed in normal operation — if the fleet has never lost more than two instances an hour, two per hour is a generous ceiling and anything beyond it deserves a human. A rate limit converts a runaway from an outage into a page. ## Safeguard 3: cap the blast radius Separately from the rate, cap the *simultaneous* impact: never take out more capacity than the fleet can lose while still serving peak demand. If the service needs six instances at peak and runs eight, the loop may have at most one instance out of service at a time and still keep a spare. This is the same reasoning as redundancy headroom — running N+1 means one unit may be absent, N+2 means two — and the remediator must be aware of the headroom it is consuming, not just of the instance it is fixing. ## Safeguard 4: a circuit breaker on the automation itself The loop should watch its own effectiveness. If it has replaced three instances and the unhealthy count has not improved, its model of the world is wrong; continuing is destructive. Tripping a breaker — stop acting, hold the state, page — is the mechanical version of the human who stopped to think. A breaker that only resets on human acknowledgement is stronger than one that resets on a timer, because a timed reset resumes the runaway at 3am with nobody watching. ## Safeguard 5: verification between actions A loop that fires on a stale reading acts on a world that no longer exists. Re-read the health signal after each action, wait for the replacement to become healthy before starting another, and require the condition to persist for a duration before acting at all. A required persistence window is cheap and eliminates the entire class of transient blips — the monitoring bug in this scenario might not even survive it. ## Safeguard 6: an audit trail and an off switch Every action, with a timestamp, the instance affected, and the reading that triggered it, written somewhere a human can read during an incident. Without it, the incident begins with an unanswerable question about who or what terminated the instances. And the off switch must be documented in the same place on-call already looks, and must have been exercised — an untested kill switch is not a control, it is a belief. ## Safeguard 7: earn the authority gradually Run the loop in observe-only mode first, logging the action it would have taken. Compare against reality for a few weeks. That measures the false-positive rate on live signals, which is the number that decides whether this automation is safe at all — and it is a number no staging environment can produce. ## How to answer in an interview Lead with the signal-plausibility check, because it is the guard that specifically defeats the scenario as posed, and many candidates only reach for rate limiting. Then layer rate limit, blast-radius cap, breaker, audit trail and off switch. Finish with observe-only rollout. The through-line to state explicitly: the remediator is only as good as its input, so it must be built to doubt its input.
- Should the circuit breaker reset itself automatically after a cooldown?Prefer a reset that requires human acknowledgement. A timed reset means the loop resumes the same runaway an hour later, typically overnight, with nobody watching — you have delayed the damage rather than stopped it. If you must auto-reset, make the cooldown long, log the reset loudly, and keep a hard daily action ceiling underneath it that no reset clears.
- How does requiring the unhealthy condition to persist before acting change the safety profile?It removes the entire class of transient blips at almost no cost. A check that must fail continuously for, say, two minutes will not fire on a single scrape failure, a deploy restart, or a brief network partition in the checker. The trade is exactly that persistence window of extra exposure when the failure is real, which is usually a very good deal.
- What do you log so that a later incident review can reconstruct what the automation did?For each decision: the timestamp, the reading that triggered it, the guard checks that passed, the action taken, and the outcome — including the decisions where it chose not to act. The non-actions matter as much as the actions, because "why did the remediator sit still" is a question incident reviews ask just as often as "why did it fire".
saying these in an interview costs you the question
- More failures detected means remediate harder
- A rate limit alone is enough protection
- The health signal is assumed to be correct
- A kill switch nobody has ever tested counts as a control
- Auto-reset the breaker after a short cooldown