skip to content

Principles of Chaos Experiments

Chaos engineering as disciplined experimentation, not random breakage. Interviewers ask how you would introduce chaos safely — hypothesis, blast radius, and abort conditions are the expected structure.

on this pageshow

questions

6

In a chaos experiment, what is the steady-state hypothesis, and how do you choose the metric it is stated in?

level: middleimportance: must knowfreq 70%

answer

  1. measure output, not internals
  2. normal band drawn from real variance
  3. falsifiable sentence, not a hope
  4. hypothesis about the system, not the fault
  5. concurrent control beats previous hour

basics

~20 s

The steady-state hypothesis is a falsifiable claim that a measured, user-visible output of the system stays inside a defined band while a fault is present. State it on output the customer feels, not on internal resource metrics.

solid answer

~50 s

Steady state is the system's normal *output* behaviour expressed as a number with a tolerance band — successful checkouts per minute, stream starts per second, request success rate — measured over a baseline window before anything is injected. The hypothesis is one falsifiable sentence built on it: "while one payment-gateway instance is unavailable, checkout success rate stays within 5% of baseline." Choose an output metric rather than CPU, memory or thread counts, because internals move for a dozen unrelated reasons and normal-looking internals tell you nothing about whether customers were harmed. Two design details matter: pick the band from the metric's real variance, so ordinary noise cannot falsify you, and prefer comparing a treated group against a concurrent control group over comparing to the previous hour — that removes time-of-day and deploy confounders. If no observation could contradict the sentence, it isn't a hypothesis.

code

python · 11 lines
python
BASELINE_ORDERS_PER_MIN = 420.0
TOLERANCE = 0.05  # hypothesis holds while the order rate stays within 5% of baseline


def steady_state_holds(observed_orders_per_min: float) -> bool:
    drop = (BASELINE_ORDERS_PER_MIN - observed_orders_per_min) / BASELINE_ORDERS_PER_MIN
    return drop <= TOLERANCE


print(steady_state_holds(410.0))  # True  -> inside the band, hypothesis survives
print(steady_state_holds(300.0))  # False -> abort and roll the fault back

go deeper

for a junior

Know that steady state is a normal-behaviour metric measured before the fault, and be able to say the hypothesis must be a sentence that could turn out false.

for a middle

Be ready to write one for a service on the spot: name the output metric, where the baseline comes from, the tolerance band, and why you rejected CPU as the steady-state signal.

for a senior

Demonstrate that you know how experiments produce false confidence — bands wider than the effect, aggregates that hide a broken cohort, and before/after comparisons confounded by a concurrent deploy.

for a principal

Own the standard: which SLIs across the estate are trustworthy enough to be steady-state metrics at all, and whether teams have the instrumentation to make a hypothesis measurable before they are allowed to inject anything.

## What "steady state" means Steady state is not "the system is healthy". It is a specific, measurable property of the system's **output** that holds while the system is doing its job normally, expressed as a number plus the range that number normally occupies. For a storefront it might be *successful checkouts per minute, ~420, normally within ±5% over a ten-minute window*. For a video service, stream starts per second. For an API, the ratio of successful to total requests. The reason the definition insists on output is that a chaos experiment is asking a question about the *system*, not about a machine. Whether a host's CPU stayed under 80% is uninteresting if orders stopped; whether CPU spiked is uninteresting if orders never wavered. ## The hypothesis is one falsifiable sentence The canonical shape: > While **<event>** is occurring, **<steady-state metric>** remains within **<band>**. Three properties make it usable: **It can be wrong.** "The system degrades gracefully" cannot be contradicted by any measurement, so it teaches nothing. "Checkout success stays within 5% of the 420/min baseline" can be. **It is about the system, not the fault.** A frequent beginner error is to state the hypothesis on the injection itself — "the killed instance is replaced within 30 seconds". That measures your orchestrator's reaction, which you probably already know. The interesting question is whether users noticed while it happened. **It carries a band, and the band comes from real variance.** If the metric naturally swings ±15% hour to hour and you assert ±5%, the experiment will "fail" on ordinary noise and you will learn to distrust it. If it swings ±1% and you assert ±15%, you have built an experiment that cannot detect the regression you care about. Derive the band from the baseline window, and check that the effect size you would consider a failure is larger than it. ## Choosing the metric A usable rank order: 1. **Business output** — orders, sign-ups, plays, messages delivered. Closest to what harm means, hardest to argue with. 2. **User-facing service level** — request success rate or latency at a percentile on the critical path. 3. **Internal resource metrics** — CPU, memory, queue length, connection counts. These belong in the experiment as *guardrails* that can abort the run, not as the steady state it is judged on. Two cautions on aggregates. First, a top-line number can hide a total failure for a small cohort: p99 latency across all traffic may not move at all when one shard, one region or one tenant is completely broken, so scope the metric to the population the fault actually reaches. Second, a low-volume metric may be too noisy to say anything in a ten-minute window — sometimes the honest fix is a longer run rather than a different metric. ## Control group beats before-and-after Comparing the experiment window against the previous hour is a weak design: traffic shape, a concurrent deploy, a marketing email or a third-party slowdown all confound it. The stronger design splits comparable traffic into a **control** group and a **treatment** group that receives the fault, and compares the two *at the same moment*. Both see identical conditions, so a difference is attributable to the injected variable. Netflix built this into its chaos platform for exactly that reason, and it also means a smaller sample can support a conclusion. ## Writing the check down Make the hypothesis executable rather than a sentence someone eyeballs on a dashboard. If it is a predicate, the same predicate can drive the abort: ```python def steady_state_holds(observed: float, baseline: float, tolerance: float = 0.05) -> bool: return (baseline - observed) / baseline <= tolerance ``` Once the check is code, "is the hypothesis still true?" is answered continuously during the run instead of after it, and the same function can be reused when the experiment is repeated months later. ## Reading the result - **Held, with a band that could have detected the effect** — real evidence about a real property. - **Falsified** — you found the weakness. Note the magnitude, not just the yes/no: a 6% dip and a 60% dip are different findings. - **Held, but the band was wider than any plausible impact** — a null result that proves nothing. This is the outcome that quietly produces false confidence, and it is the one to check for before you declare success. The last case is why the hypothesis is written first. Written afterwards, it always fits.

  • Why prefer a concurrent control group over comparing the experiment window to the previous hour?
    Because time is a confounder. Traffic shape, a deploy, a marketing send or a third-party slowdown can all move the metric between the two windows, so any difference is ambiguous. Splitting comparable traffic into control and treatment groups running at the same moment isolates the injected variable, and lets a smaller sample support a conclusion.
  • Your steady-state metric is p99 latency, the experiment shows no change, yet support tickets spike during the run. What went wrong?
    The metric didn't cover the harmed population. A global p99 can be flat while one shard, region or tenant fails completely, because that cohort is a small share of the aggregate. Scope the steady-state metric to the population the fault actually reaches, and pair a latency metric with a success-rate metric — a request that errors fast never appears as slow.
  • How long should the baseline window be before you inject anything?
    Long enough to capture the metric's ordinary variance at that time of day — often the same duration as the planned run, and never shorter. The baseline exists to set the tolerance band; a two-minute baseline on a metric with a ten-minute cycle produces a band that is either meaninglessly wide or guaranteed to be breached.

saying these in an interview costs you the question

  • Steady state means CPU and memory look normal
  • The hypothesis is "let's see what happens"
  • Stating the hypothesis on the fault, not the system's output
  • No tolerance band, so any wobble counts as failure
  • Comparing against yesterday and calling it a control

context

open as a page

You are scoping the first production chaos experiment for a payment service. Along which dimensions do you minimize the blast radius, and how do you decide when to widen it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Minimize along population, duration, severity and reversibility: smallest slice of traffic or one instance, a hard time cap, the mildest fault that still tests the claim, and an undo. Widen one dimension at a time only after the hypothesis holds.

open as a page

A colleague proposes to "kill a random production server on Friday afternoon and see what breaks" and calls it chaos engineering. What does a real chaos experiment have that this proposal is missing?

level: juniorimportance: should knowfreq 62%

basics

~20 s

A chaos experiment adds four things to breaking something: a measured steady state, a falsifiable hypothesis about it, one fault scoped to the smallest population that can still test it, and pre-agreed abort conditions with a rollback.

open as a page

What makes a chaos experiment's abort condition a real control rather than a formality, and what belongs in its rollback plan?

level: seniorimportance: should knowfreq 48%

basics

~20 s

An abort condition is real when it is a numeric threshold agreed before the run and wired to an automatic halt, backed by a one-action rollback that was tested first and works even if the fault broke the path you would normally use to undo it.

open as a page

Chaos engineering's principles call for running experiments in production. What does chaos testing only in a staging environment fail to catch, and when is staying out of production the right call?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Staging lacks production's traffic volume and mix, data size, real dependency versions, scaling settings and — above all — its configuration, which is where most failures live. A green staging run is evidence about staging. Stay out of production when the fault is irreversible or you cannot yet measure harm.

open as a page

When does a one-off chaos experiment deserve to be automated into one that runs continuously, and what does that automation cost you?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Automate when the property being tested is one that silently regresses as the system changes, the hypothesis already passes at the target scope, and the abort path runs without a human. The cost is a new production-critical system with its own on-call, maintenance and error-budget spend.

open as a page