When does a one-off chaos experiment deserve to be automated into one that runs continuously, and what does that automation cost you?
answer
- one-off probes, scheduled runs regress-test
- automate what silently decays
- the harness becomes production-critical
- stale experiments manufacture false confidence
- a policy check may be cheaper than chaos
basics
~20 sAutomate when the property being tested is one that silently regresses as the system changes, the hypothesis already passes at the target scope, and the abort path runs without a human. The cost is a new production-critical system with its own on-call, maintenance and error-budget spend.
solid answer
~50 sA one-off experiment is a probe; a continuous one is a regression test for a reliability property. Promote an experiment when the property erodes quietly — someone removes a timeout, adds a synchronous dependency, resizes a pool — and when three preconditions hold: the hypothesis passes at the intended scope, the steady-state check is code rather than a person reading a graph, and the abort is automated and rehearsed. The costs are real. The harness becomes production-critical: it can cause outages, so it needs permissions, audit, a central kill switch and an owner. Experiments go stale as the architecture changes, and a stale experiment that tests a dependency nobody calls any more manufactures confidence. It spends error budget every run, and scheduled runs must be visible to on-call or they generate self-inflicted pages. Before automating, ask whether a structural fix or a build-time check gets the same guarantee cheaper.
go deeper
Know that some chaos experiments are run on a schedule rather than once, and that the automated version still needs the same steady-state check and abort conditions.
Be able to explain what continuous chaos protects against — reliability properties quietly regressing as code and config change — and why the abort must be code, not a person watching.
Show the operational consequences: the harness is production-critical, scheduled runs must be visible to on-call, and stale experiments produce green results that mean nothing.
Own the tradeoff and the guardrails — when a build-time or policy check buys the same assurance more cheaply, what the platform enforces centrally (scope caps, stop conditions, kill switch), and the maturity path from manual runs to a small set of default-on experiments.
## Probe versus regression test A single experiment answers *is this true today?* Running it continuously answers a different question: *is this still true?* That is worth paying for only when the property is one that decays without anyone noticing. Reliability properties decay constantly: a timeout gets relaxed to make a flaky test pass, a new synchronous call is added to a path that used to be fully async, a pool is resized during an unrelated incident, a fallback is removed as dead code, a retry budget quietly stops being enforced. Each is invisible in review and each falsifies a hypothesis that passed six months ago. ## Promotion criteria A reasonable bar before an experiment goes on a schedule: 1. **It passes reliably at the intended scope.** Automating an experiment that fails intermittently automates an outage. 2. **The steady-state check is executable.** If a human has to read a dashboard to know the result, it is not automatable, and the effort belongs in making the SLI measurable first. 3. **The abort is automatic and rehearsed**, with a tested out-of-band undo. 4. **The property is worth continuous verification** — a load-bearing claim like "we survive losing one zone", not a curiosity. 5. **Someone owns the experiment**, in the same sense someone owns a test suite. Unowned experiments rot. ## What automation actually costs **The harness becomes production-critical.** A system that can deliberately break production is one of the most dangerous services you run. It needs an owner and an on-call, a permissions model (who may target which services, at what blast radius), an audit trail of every injection, safety interlocks that enforce maximum scope centrally rather than trusting each experiment's config, and a global kill switch that a single human can hit. Getting this wrong produces an incident where the chaos platform is the cause and nobody can stop it. **Staleness manufactures false confidence.** An experiment that blackholes a dependency the service stopped calling two quarters ago passes every night and proves nothing, while the dashboard shows a reassuring row of green. Continuous experiments need periodic review against the current architecture, and the review is the part teams skip. **Ongoing error-budget spend.** Each run consumes a little reliability. Small per run, not small annually. It has to be inside the ceiling the team agreed for the programme, and the schedule must respond to budget state. **Interaction with on-call and alerting.** A scheduled run that fires a page is either noise the on-call learns to dismiss — the far worse outcome, because they will dismiss the real one too — or a signal that has to be correlated with the schedule. Scheduled chaos must be annotated where responders look, and there must be an unambiguous way to answer "is this us?" in seconds. ## Scheduling policy The schedule is where the judgement shows: - **Business hours only.** Netflix's originally published Chaos Monkey design terminated instances during the working day precisely so engineers were present to observe and respond. Overnight automation optimises for nobody noticing, which is the opposite of the goal. - **Suspend on budget exhaustion.** If real incidents have already spent the error budget, the experiments pause. - **Suspend during incidents, freezes, launches and peak periods.** Ideally enforced by the platform, not by a convention people remember. - **Don't compound.** No injection during an in-flight deploy to the same service, or the two effects are inseparable. - **Vary, but attributably.** Randomised targeting finds more, but every run must be recorded so any anomaly can be matched to an injection. ## The alternative you should price first The most useful principal-level instinct here is to ask whether continuous chaos is the cheapest way to get the guarantee. Sometimes it plainly isn't: - "Every outbound call has a timeout" is better enforced by a build-time check or a shared client library than by an experiment that discovers a missing one at 11:00 on a Tuesday. - "No service has a single replica" is a policy check against the deployment manifests. - "Failover works" often cannot be checked statically at all — that one earns its experiment. Continuous chaos is the right tool for **emergent, cross-service, state-dependent** properties that no static check can see. Where a lint rule, a policy gate or an architectural constraint gives the same assurance, use it: it is cheaper, faster, and it never causes an outage. ## The maturity path Most organisations that get this right walk the same road: manual experiments run by a small team; a shared harness with mandatory stop conditions and enforced blast-radius caps; opt-in scheduled experiments owned by service teams; and finally a small number of default-on experiments encoding the platform's load-bearing claims. Jumping straight to default-on across an estate is how a chaos programme gets banned after its first self-inflicted major incident — and the ban usually outlasts the people who caused it.
- What makes a continuously running chaos experiment dangerous in a way a one-off is not?Nobody is watching. A one-off has an owner at a keyboard; a scheduled run relies entirely on its automated abort, and on that abort still being correct months after the architecture moved. It also accumulates: dozens of scheduled experiments across an estate become a system that can cause an outage with no human in the loop, which is why the platform needs enforced blast-radius caps and a global kill switch.
- How would you stop scheduled chaos runs from generating pager noise?Annotate every run where responders look, so "is this us?" is answered in seconds; keep runs inside business hours; and suppress only alerts scoped to the injected fault, never the service's real customer-facing alerts. If a run reliably pages, that is a finding to fix, not an alert to silence — and an on-call trained to dismiss chaos pages will dismiss the real one too.
- Give an example of a reliability property that should not be verified by continuous chaos."Every outbound HTTP call sets a timeout." A build-time check or a shared client library enforces that before the code ships, deterministically and at zero production risk. Chaos earns its cost on emergent, state-dependent, cross-service properties — zone failover, cache stampedes, load redistribution — that no static analysis can see.
saying these in an interview costs you the question
- Automate every experiment; more chaos is always better
- Run scheduled chaos overnight so fewer customers are affected
- The chaos platform is a tool, not a production system
- A passing nightly experiment always means the property still holds
- Randomised targeting with no record of what was injected