skip to content

What makes a chaos experiment's abort condition a real control rather than a formality, and what belongs in its rollback plan?

level: seniorimportance: should knowfreq 48%

answer

  1. thresholds agreed before, not during
  2. automated halt beats a watched dashboard
  3. undo must not need what you broke
  4. hard time cap and dead-man's switch
  5. abort is a result, not a failure

basics

~20 s

An abort condition is real when it is a numeric threshold agreed before the run and wired to an automatic halt, backed by a one-action rollback that was tested first and works even if the fault broke the path you would normally use to undo it.

solid answer

~50 s

Three properties. **Pre-declared and numeric** — "customer error rate above twice baseline for 60 seconds" decided before injection, not "stop if it looks bad", because during a run everyone is invested and "give it one more minute" always wins. **Automated** — the halt fires from the same predicate that evaluates the steady-state hypothesis, so it doesn't depend on a human watching a dashboard; platforms like AWS Fault Injection Service let you attach a stop condition backed by a CloudWatch alarm for this. **Independently reversible** — the undo is one action, rehearsed before the fault went in, and does not route through whatever you just degraded. Add a hard time cap and a dead-man's switch so the fault expires by itself if the operator's session dies, plus a standing rule that the on-call may halt without debate. And treat an abort as a result: the hypothesis was falsified, which is the finding.

go deeper

for a junior

Know that an experiment must have a pre-agreed stopping rule and a way to undo the fault, and that both are decided before anything is injected.

for a middle

Be ready to state concrete thresholds — a metric, a multiple of baseline, a duration — and explain why the halt should be automated rather than left to whoever is watching.

for a senior

Show that you have thought about the undo path failing: TTL-expiring faults, out-of-band removal, rehearsed rollback, and the rule that the on-call can halt without debate.

for a principal

Own the guardrails as platform policy — no experiment runs without a stop condition and a tested revert, a central kill switch exists, and experiment-caused customer impact enters the normal incident process rather than being cleaned up quietly.

## The failure mode this guards against The classic chaos incident is not the injection — it is the *minutes after* the injection during which somebody was watching the graph, wasn't sure, and waited. Abort conditions exist because judgement in the moment is systematically worse than judgement beforehand: the people running an experiment want it to succeed, they have context that explains away the first anomaly, and the person best placed to notice is often the person concentrating on the injection rather than the customer. ## Pre-declared and numeric Write the thresholds down with the hypothesis, before anything is injected. A workable set has three layers: - **The steady-state predicate itself** — the metric leaves its band for longer than a defined interval. - **Guardrail metrics** the experiment isn't about but must not damage — customer-facing error rate, checkout success, queue depth, the rate at which the SLO's budget is being consumed. - **A hard time cap** — the run ends at T+10 minutes regardless of what anyone thinks. The rule for each is stated as a number and a duration: *error rate above 2× baseline for 60 consecutive seconds*. "Stop if it looks bad" is not a control, because it has no threshold, no detection window and no owner. ## Automated, not supervised A human watching a dashboard adds latency (notice, interpret, decide, act) and bias ("that spike is the deploy, not us"). The stop should be executed by the same code that evaluates the hypothesis: ```bash #!/usr/bin/env bash set -euo pipefail # Undo is registered BEFORE the fault goes in, and fires on any exit path. trap 'rollback_fault' EXIT INT TERM inject_fault deadline=$(( SECONDS + 600 )) # hard 10-minute cap while (( SECONDS < deadline )); do check_steady_state || { echo "steady state violated - aborting"; exit 1; } sleep 10 done ``` Two details in that sketch carry the weight. The `trap` is registered *before* injection, so a crash, a Ctrl-C or a dropped SSH session still reverts the fault. And the deadline is absolute, so an operator who walks away does not leave a fault running. Better still is a fault with its own **time-to-live**, so it expires at the injection point without anyone needing to send a command — a dead-man's switch rather than a promise to clean up. ## The rollback must not depend on what you broke The subtle and genuinely dangerous case: your experiment degrades a control plane, a service mesh, a network path or an identity provider — and the undo command needs that same component. Now the fault is stuck in place and you are in a real incident with your recovery tool disabled. Design an out-of-band undo: - Faults that self-revert on a lost heartbeat or an expiring TTL. - A removal path that is local to the affected node and needs no central coordination. - Credentials and access for the undo that don't traverse the degraded dependency. - Rehearsal: inject and remove the fault at a trivial scope first, so you have proof the removal works *before* the run that matters. If no independent undo exists, that is a reason not to run the experiment, not a risk to accept. ## Human authority alongside the automation Automation handles the thresholds you predicted. People handle the ones you didn't. So the run also needs: - A named owner watching, whose only job is the experiment. - A standing rule that the on-call engineer may halt the experiment immediately, without justifying it and without finding the owner first. Any friction there converts into minutes of customer impact. - Advance notice to the on-call that the run is happening, so a resulting alert is not triaged from scratch as an unexplained anomaly — and, equally, so nobody assumes an unrelated real incident is "just the chaos test". ## After the abort An abort is a **successful experiment**: the hypothesis was falsified, cheaply, on your schedule. Record the magnitude and the timing, not just the fact. Two follow-ups matter more than the fix itself: how long elapsed between the real onset and the halt (that number is your detection capability, measured for free), and whether existing alerting would have caught this had it happened unannounced at 03:00. If the experiment's own tooling noticed before your monitoring did, you have found a second, larger finding. And if customers were actually harmed, the run stops being an experiment and becomes an incident — declared and handled through the normal response path, not quietly cleaned up because it was self-inflicted.

  • Why isn't "an engineer watches the dashboard and stops it if things look bad" good enough?
    It has no threshold, so "bad" is negotiated live by people who want the run to succeed — and "give it another minute" reliably wins. It also adds notice-interpret-decide-act latency, and the watcher may be the person the fault just disconnected. Automate the halt against a number, and keep the human for the failure modes you didn't predict.
  • Your rollback command routes through the same service mesh control plane the experiment just degraded. What do you do?
    Don't run it until there is an out-of-band undo: a fault with a TTL that expires by itself, an agent that self-reverts when it loses its heartbeat, or a node-local removal path needing no central coordination. Rehearse injection and removal at trivial scope first. If no independent undo exists, that is a stop, not a risk to accept.
  • The experiment aborted after eight minutes. What is worth measuring besides the failure itself?
    The gap between real onset and the halt — that is your detection time, measured for free — and whether your normal alerting would have fired at all if this had happened unannounced overnight. If the experiment harness noticed before production monitoring did, the monitoring gap is the bigger finding.

saying these in an interview costs you the question

  • We'll stop it manually if the graphs look bad
  • Abort thresholds decided while the experiment is running
  • Rollback tested for the first time during the real run
  • An aborted experiment means the experiment failed
  • No time cap — we run it until we've seen enough

context