In a chaos experiment, what is the steady-state hypothesis, and how do you choose the metric it is stated in?
answer
- measure output, not internals
- normal band drawn from real variance
- falsifiable sentence, not a hope
- hypothesis about the system, not the fault
- concurrent control beats previous hour
basics
~20 sThe steady-state hypothesis is a falsifiable claim that a measured, user-visible output of the system stays inside a defined band while a fault is present. State it on output the customer feels, not on internal resource metrics.
solid answer
~50 sSteady state is the system's normal *output* behaviour expressed as a number with a tolerance band — successful checkouts per minute, stream starts per second, request success rate — measured over a baseline window before anything is injected. The hypothesis is one falsifiable sentence built on it: "while one payment-gateway instance is unavailable, checkout success rate stays within 5% of baseline." Choose an output metric rather than CPU, memory or thread counts, because internals move for a dozen unrelated reasons and normal-looking internals tell you nothing about whether customers were harmed. Two design details matter: pick the band from the metric's real variance, so ordinary noise cannot falsify you, and prefer comparing a treated group against a concurrent control group over comparing to the previous hour — that removes time-of-day and deploy confounders. If no observation could contradict the sentence, it isn't a hypothesis.
code
python · 11 linesBASELINE_ORDERS_PER_MIN = 420.0
TOLERANCE = 0.05 # hypothesis holds while the order rate stays within 5% of baseline
def steady_state_holds(observed_orders_per_min: float) -> bool:
drop = (BASELINE_ORDERS_PER_MIN - observed_orders_per_min) / BASELINE_ORDERS_PER_MIN
return drop <= TOLERANCE
print(steady_state_holds(410.0)) # True -> inside the band, hypothesis survives
print(steady_state_holds(300.0)) # False -> abort and roll the fault backgo deeper
Know that steady state is a normal-behaviour metric measured before the fault, and be able to say the hypothesis must be a sentence that could turn out false.
Be ready to write one for a service on the spot: name the output metric, where the baseline comes from, the tolerance band, and why you rejected CPU as the steady-state signal.
Demonstrate that you know how experiments produce false confidence — bands wider than the effect, aggregates that hide a broken cohort, and before/after comparisons confounded by a concurrent deploy.
Own the standard: which SLIs across the estate are trustworthy enough to be steady-state metrics at all, and whether teams have the instrumentation to make a hypothesis measurable before they are allowed to inject anything.
## What "steady state" means Steady state is not "the system is healthy". It is a specific, measurable property of the system's **output** that holds while the system is doing its job normally, expressed as a number plus the range that number normally occupies. For a storefront it might be *successful checkouts per minute, ~420, normally within ±5% over a ten-minute window*. For a video service, stream starts per second. For an API, the ratio of successful to total requests. The reason the definition insists on output is that a chaos experiment is asking a question about the *system*, not about a machine. Whether a host's CPU stayed under 80% is uninteresting if orders stopped; whether CPU spiked is uninteresting if orders never wavered. ## The hypothesis is one falsifiable sentence The canonical shape: > While **<event>** is occurring, **<steady-state metric>** remains within **<band>**. Three properties make it usable: **It can be wrong.** "The system degrades gracefully" cannot be contradicted by any measurement, so it teaches nothing. "Checkout success stays within 5% of the 420/min baseline" can be. **It is about the system, not the fault.** A frequent beginner error is to state the hypothesis on the injection itself — "the killed instance is replaced within 30 seconds". That measures your orchestrator's reaction, which you probably already know. The interesting question is whether users noticed while it happened. **It carries a band, and the band comes from real variance.** If the metric naturally swings ±15% hour to hour and you assert ±5%, the experiment will "fail" on ordinary noise and you will learn to distrust it. If it swings ±1% and you assert ±15%, you have built an experiment that cannot detect the regression you care about. Derive the band from the baseline window, and check that the effect size you would consider a failure is larger than it. ## Choosing the metric A usable rank order: 1. **Business output** — orders, sign-ups, plays, messages delivered. Closest to what harm means, hardest to argue with. 2. **User-facing service level** — request success rate or latency at a percentile on the critical path. 3. **Internal resource metrics** — CPU, memory, queue length, connection counts. These belong in the experiment as *guardrails* that can abort the run, not as the steady state it is judged on. Two cautions on aggregates. First, a top-line number can hide a total failure for a small cohort: p99 latency across all traffic may not move at all when one shard, one region or one tenant is completely broken, so scope the metric to the population the fault actually reaches. Second, a low-volume metric may be too noisy to say anything in a ten-minute window — sometimes the honest fix is a longer run rather than a different metric. ## Control group beats before-and-after Comparing the experiment window against the previous hour is a weak design: traffic shape, a concurrent deploy, a marketing email or a third-party slowdown all confound it. The stronger design splits comparable traffic into a **control** group and a **treatment** group that receives the fault, and compares the two *at the same moment*. Both see identical conditions, so a difference is attributable to the injected variable. Netflix built this into its chaos platform for exactly that reason, and it also means a smaller sample can support a conclusion. ## Writing the check down Make the hypothesis executable rather than a sentence someone eyeballs on a dashboard. If it is a predicate, the same predicate can drive the abort: ```python def steady_state_holds(observed: float, baseline: float, tolerance: float = 0.05) -> bool: return (baseline - observed) / baseline <= tolerance ``` Once the check is code, "is the hypothesis still true?" is answered continuously during the run instead of after it, and the same function can be reused when the experiment is repeated months later. ## Reading the result - **Held, with a band that could have detected the effect** — real evidence about a real property. - **Falsified** — you found the weakness. Note the magnitude, not just the yes/no: a 6% dip and a 60% dip are different findings. - **Held, but the band was wider than any plausible impact** — a null result that proves nothing. This is the outcome that quietly produces false confidence, and it is the one to check for before you declare success. The last case is why the hypothesis is written first. Written afterwards, it always fits.
- Why prefer a concurrent control group over comparing the experiment window to the previous hour?Because time is a confounder. Traffic shape, a deploy, a marketing send or a third-party slowdown can all move the metric between the two windows, so any difference is ambiguous. Splitting comparable traffic into control and treatment groups running at the same moment isolates the injected variable, and lets a smaller sample support a conclusion.
- Your steady-state metric is p99 latency, the experiment shows no change, yet support tickets spike during the run. What went wrong?The metric didn't cover the harmed population. A global p99 can be flat while one shard, region or tenant fails completely, because that cohort is a small share of the aggregate. Scope the steady-state metric to the population the fault actually reaches, and pair a latency metric with a success-rate metric — a request that errors fast never appears as slow.
- How long should the baseline window be before you inject anything?Long enough to capture the metric's ordinary variance at that time of day — often the same duration as the planned run, and never shorter. The baseline exists to set the tolerance band; a two-minute baseline on a metric with a ten-minute cycle produces a band that is either meaninglessly wide or guaranteed to be breached.
saying these in an interview costs you the question
- Steady state means CPU and memory look normal
- The hypothesis is "let's see what happens"
- Stating the hypothesis on the fault, not the system's output
- No tolerance band, so any wobble counts as failure
- Comparing against yesterday and calling it a control