skip to content

A watchdog restarts your API process whenever its health check fails. It has been quietly restarting the process about three times a night for two months; last night the restarts could not keep up and the service was down for 40 minutes. What went wrong in the way that automation was built?

level: seniorimportance: must knowfreq 55%

answer

  1. the fix hid its own evidence
  2. borrowed time, spent silently
  3. count every automated action
  4. alert on remediation rate, not the fault
  5. let the loop give up and escalate

basics

~20 s

The watchdog suppressed the symptom without reporting it, so a slowly worsening fault stayed invisible until it outran the remediation. Auto-remediation must count every action, escalate when the action rate climbs, and refuse to keep repairing indefinitely.

solid answer

~50 s

The automation worked exactly as written, and that is the problem: it treated a recurring failure as a routine event instead of as evidence. Self-healing does not repair anything — it buys time, and here the team spent two months of borrowed time without knowing it. The fix is to make the remediator report on itself: emit a counter for every restart, open a ticket or raise a low-priority alert the first time the daily count exceeds baseline, and page when the rate crosses a threshold — three restarts an hour, say, rather than three a night. It should also have a ceiling: after N restarts in a window it stops restarting and escalates, because at that point the loop is no longer mitigating, it is hiding an outage in progress. And those restarts almost certainly dropped in-flight requests, so users were seeing errors the whole time — that should have shown up in the service's own error rate.

code

bash · 20 lines
bash
#!/usr/bin/env bash
# Restart the API, but count every restart and escalate rather than hide it.
set -euo pipefail

state=/var/lib/watchdog/restarts
mkdir -p "$(dirname "$state")"
hour=$(date +%Y%m%d%H)
count=$(grep -c "^$hour " "$state" 2>/dev/null || echo 0)

if [ "$count" -ge 3 ]; then
  logger -p daemon.err "watchdog: $count restarts this hour - refusing to restart, escalating"
  exit 1
fi

# Capture evidence before the restart destroys it.
logger -p daemon.warning "watchdog: restart $((count + 1)) this hour"
ps -o pid,rss,nlwp -p "$(pgrep -f api-server | head -n1)" >> /var/log/watchdog-evidence.log || true

echo "$hour restart" >> "$state"
systemctl restart api

go deeper

for a junior

Know that restarting a failing process hides the failure rather than fixing it, and that every automatic restart should be logged and counted so someone can see it happened.

for a middle

Explain why the alert has to move from the fault to the rate of remediation once a fix is automated, and describe the two tiers — a ticket when the count exceeds baseline, a page when the rate climbs.

for a senior

Demonstrate that you would give the loop a ceiling so it escalates instead of grinding, capture diagnostics before each remediation, and check whether the remediated errors are visible in the service's reliability measurement at all.

for a principal

Own the standard: any remediator your organisation ships must emit an action counter, an escalation threshold and an owner before it is allowed to run unattended, so hidden risk cannot accumulate quietly across dozens of services.

## Self-healing defers a problem; it does not solve one A watchdog that restarts a failing process is a legitimate and valuable control. What it is not is a repair. It converts a long outage into a short blip, and it converts an obvious failure into an invisible one. That trade is worth making — as long as the invisibility is deliberately undone by instrumentation. When it is not, the automation becomes a mechanism for accumulating hidden risk at a steady rate. The pattern in this scenario is the textbook one. Something in the process degrades over hours — a leak, a file-descriptor exhaustion, a connection pool that never drains. Restarting resets it. The reset is cheap enough that nobody sees it, so the underlying defect is never diagnosed, and it gets worse: the interval between restarts shrinks from eight hours to four to twenty minutes, and eventually the process cannot stay up long enough to be useful. ## What the automation should have emitted A remediator has three outputs, not one: 1. **The action itself** — restart the process. 2. **A record of the action** — a counter, incremented every time, with a label saying which condition fired. This is the difference between "three restarts a night" being a fact somebody could have noticed and being nothing at all. 3. **An escalation when the record looks wrong** — this is the part teams skip. The escalation has two tiers. A *ticket* tier: the first time a night's restart count exceeds the normal baseline, file work for the owning team; this is not urgent, but it must not be silent. A *page* tier: when the rate crosses a threshold that means the remediation is losing — restarts an hour rather than a night, or a failure to come back healthy after the restart. ## Alert on the remediation, not only on the fault The key inversion is that once you automate a fix, the fault stops being a good signal — the automation is designed to make it go away. The signal that carries information is now the *rate of remediation*. A team that only alerts on "the API is down" will never fire an alert, because the API is only down for four seconds at a time. A team that alerts on "the watchdog restarted the API more than three times in an hour" would have been paged in week one, at leisure, with a diagnosable process still running. ## Give the loop a ceiling A remediator should be allowed to give up. Encoding a limit — after N actions in a window, stop acting and escalate — does two things. It prevents the automation from grinding indefinitely against a fault it cannot fix, and it forces a human decision at exactly the moment when human judgement is worth more than another restart. In this incident, a ceiling would have converted a 40-minute unexplained outage into a page an hour earlier with a clear message: the watchdog has given up. ```bash #!/usr/bin/env bash set -euo pipefail state=/var/lib/watchdog/restarts hour=$(date +%Y%m%d%H) count=$(grep -c "^$hour" "$state" 2>/dev/null || echo 0) if [ "$count" -ge 3 ]; then logger -p daemon.err "watchdog: $count restarts this hour, escalating" exit 1 fi echo "$hour" >> "$state" systemctl restart api ``` ## The restarts were not free A process restart drops whatever it was doing. Three restarts a night means three bursts of failed requests, retries and elevated latency that real users experienced. Those errors belong in the service's own reliability measurement — if they never appeared there, the measurement is not capturing what users saw, which is a second defect hiding behind the first. ## Diagnosing it now Before the next restart, capture state: heap or memory profile, open file descriptors, connection counts, thread dumps. A watchdog that restarts on failure destroys the evidence every time, which is another reason silent remediation is expensive — two months of restarts produced zero diagnostic artefacts. A mature remediator captures a snapshot *before* it acts, so the loop that buys time also collects the material needed to stop needing it. ## How to answer in an interview Say plainly that the automation was not instrumented, and that the missing alert is on the remediation rate rather than on the fault. Then add the two senior touches: the ceiling that makes the loop escalate instead of grinding, and the diagnostic capture that makes the borrowed time useful. Avoid the weak answer — "they should have fixed the root cause" — which is true of everything and describes no mechanism.

  • What threshold would you actually set for the escalation, and how would you pick it?
    Start from the observed baseline rather than a round number. If the service normally restarts zero times a week, then any restart deserves a ticket and three in an hour deserves a page. The useful property is that the threshold is a rate, so a fault that is accelerating trips it well before the remediation stops working — which is the whole point of watching the remediator instead of the fault.
  • Should the watchdog restarts count against the service's error budget?
    The restarts themselves are not the unit of measurement — the user-visible errors they cause are. If requests failed during each restart, those failures belong in the reliability measurement like any others. If they were genuinely invisible to users because another instance served the traffic, they cost nothing against the budget but should still raise a ticket, because the underlying fault is still growing.
  • How do you keep the remediation from destroying the evidence you need to fix the cause?
    Capture before you act. Have the remediator snapshot the cheap diagnostics — memory, descriptor counts, a thread dump, recent logs — into a durable location, then remediate. It costs a few seconds of extra downtime and it is the difference between two months of restarts producing a diagnosis and producing nothing at all.

It is a smoke alarm wired directly to the sprinklers with the bell disconnected: the fire keeps getting put out, nobody is ever told, and the first thing anyone learns about the wiring fault is when the water runs out.

saying these in an interview costs you the question

  • Self-healing means the problem is solved
  • A restart that users did not notice costs nothing
  • Alert on the fault; the automation handles the rest
  • The remediator should keep retrying until it works
  • Root cause analysis can wait since the loop is holding

context