A colleague proposes to "kill a random production server on Friday afternoon and see what breaks" and calls it chaos engineering. What does a real chaos experiment have that this proposal is missing?
answer
- controlled test, not a stunt
- measure normal before breaking anything
- one falsifiable sentence, one variable
- smallest scope that still falsifies
- abort thresholds agreed before injection
basics
~20 sA chaos experiment adds four things to breaking something: a measured steady state, a falsifiable hypothesis about it, one fault scoped to the smallest population that can still test it, and pre-agreed abort conditions with a rollback.
solid answer
~50 sKilling a box on a Friday is an outage you caused; an experiment is a controlled test. First you measure a steady state — a user-visible number such as successful checkouts per minute, with its normal band. Then you write a hypothesis that could turn out false: "while one of six API replicas is gone, checkout success stays within 5% of baseline." Then you scope the fault to the smallest population that could still falsify it — one replica, ten minutes — and pre-agree the abort thresholds and the one-action rollback *before* anything is injected. Finally you run it when the owning team is at their desks, not as everyone leaves for the weekend. And if you already believe it will break, don't experiment: fix it first, then run the experiment to verify the fix.
go deeper
Be ready to name the four parts — steady state, hypothesis, bounded fault, abort plan — and say plainly that the Friday proposal has none of them, so it is a scheduled outage rather than a test.
Expect to be asked to actually write the hypothesis: which metric you would pick for this service, what its normal band is, and how you would scope the fault so the result is attributable to it.
Show scheduling and staffing judgment: run when the owners are present, treat a known weakness as work to fix rather than to demonstrate, and be explicit that customer impact converts the run into an incident.
Own the programme view — which reliability claims are worth spending experiment time and error budget on, and what platform guardrails stop teams from running unbounded experiments in the first place.
## Experiment, not stunt Chaos engineering is the practice of running deliberate, controlled experiments on a system to build confidence that it withstands turbulent conditions in production. The operative word is *experiment*. You hold a belief about the system, you design a test that could prove that belief wrong, and you run the test under conditions where being wrong is survivable. "Kill a box and see what happens" has the injection and none of the discipline, so it produces an incident rather than a finding — and worse, an incident nobody can interpret, because no one wrote down beforehand what was supposed to happen. ## The four parts **1. A measured steady state.** Before you touch anything you need a number that describes the system working normally, measured on the output side — orders per minute, successful stream starts, checkout success rate — together with its normal range. Without a baseline, "it looked fine" and "it looked bad" are both opinions. **2. A falsifiable hypothesis.** One sentence of the form: *while <event> is happening, <steady-state metric> stays within <band>*. It has to be capable of being wrong. "The system will handle it gracefully" is not a hypothesis, because no observation could contradict it. **3. One varied event, with a bounded blast radius.** Change one thing so the result is attributable. Scope it to the smallest population that could still show the effect — one instance rather than a whole tier, a small slice of traffic rather than all of it, ten minutes rather than an afternoon. And prefer a fault you can undo; an experiment that destroys data is not an experiment. **4. Abort conditions and a rollback.** Decide the numeric thresholds that stop the run, and prove the undo works, before injection. Halfway through a run is the worst possible moment to be inventing either. ## Why the Friday afternoon detail matters Run experiments when the people who own the service are present, alert, and able to fix what they find. Friday afternoon is the opposite: staffing thins out, anyone paged is paged into their weekend, and a weakness discovered at 17:00 Friday either gets fixed under duress or waits until Monday with the risk still live. Traffic shape differs too, so the result may not generalise. This is not squeamishness — Netflix's Chaos Monkey was deliberately designed to terminate instances during working hours for exactly this reason. ## "We already know it will break" Chaos experiments are for verifying things you believe are *true*. If the team already suspects the service cannot survive losing that node, running the experiment mostly buys you a self-inflicted outage confirming a known gap. Spend the effort on the fix, then run the experiment afterwards to verify it. The one honest exception is evidence: if a known weakness cannot get prioritised without a demonstration, a tightly scoped experiment can be the cheapest way to make the risk concrete — at the smallest blast radius that still makes the point. ## What counts as a result Both outcomes are valuable, and neither is "nothing happened": - **Hypothesis survives** — you now have dated evidence that the property holds under that specific condition at that specific scale. That is a claim you can rely on, and a candidate to run repeatedly. - **Hypothesis falsified** — you found a real weakness before a customer did, at a scope and a time of your choosing, with the experts already watching. That is the best possible way to discover it. The failure outcome is a third one: the experiment ran and you cannot tell which happened, because there was no baseline or no defined band. ## The proposal, rewritten - **Steady state:** successful checkouts per minute, baseline ~420, normal band ±5% over a ten-minute window. - **Hypothesis:** with one of six API replicas terminated, checkout rate stays inside that band and no customer-facing error rate doubles. - **Scope:** one replica, ten minutes, Tuesday at 10:00, owning team watching. - **Abort:** checkout rate outside the band for 60 seconds, or customer error rate above twice baseline, halts the run automatically and restores the instance. - **Aftermath:** result recorded either way; if it falsified, the finding becomes owned work. Same injection, entirely different activity. The difference between the two is not how brave the team is — it is whether the run can produce a defensible answer.
- If the team already suspects the service can't survive losing that node, is the experiment still worth running?Usually no. Chaos verifies beliefs you think are true; where you already know the gap, running the experiment just buys a self-inflicted outage that confirms it. Fix it, then experiment to verify the fix holds. The exception is when the risk can't get prioritised without evidence — then run it at the smallest blast radius that still makes the point undeniable.
- The experiment ran and nothing broke. What is the value of that?You converted an assumption into dated evidence: this property held under this specific fault, at this scope, on this version of the system. That is what lets you rely on it in a design review or an incident. It also makes the experiment a candidate to repeat at a wider scope or on a schedule, so the property is re-checked as the architecture drifts.
saying these in an interview costs you the question
- Chaos engineering means randomly breaking things in production
- Skipping the hypothesis — just inject and observe what happens
- A bigger blast radius finds more bugs faster
- A successful experiment is one where something broke
- Running it out of hours so fewer customers notice