You are scoping the first production chaos experiment for a payment service. Along which dimensions do you minimize the blast radius, and how do you decide when to widen it?
answer
- four dials: who, how long, how hard, undoable
- smallest scope that can still falsify
- never a fault you cannot undo
- widen one dimension per run
- budget ceiling agreed before the programme
basics
~20 sMinimize along population, duration, severity and reversibility: smallest slice of traffic or one instance, a hard time cap, the mildest fault that still tests the claim, and an undo. Widen one dimension at a time only after the hypothesis holds.
solid answer
~50 sThere are four independent dials. **Population** — one instance, one shard, one internal tenant, or 1% of traffic rather than a whole tier. **Duration** — a hard cap of minutes, enforced automatically, not "until we've seen enough". **Severity** — 100 ms of added latency before a full dependency blackhole. **Reversibility** — never run a fault you cannot undo, and never one that touches money movement or destroys data. Time of day belongs here too: business hours with the owners watching, away from peak, launches and freezes. The floor is set by detectability — a slice too small to produce a signal above noise teaches nothing, so pick the smallest scope that could still falsify the hypothesis, not the smallest scope full stop. Widen one dial at a time after each pass, and cap the whole programme by how much error budget you are willing to spend.
go deeper
Know that a chaos experiment is deliberately scoped small — one instance or a slice of traffic, for a few minutes — and that the fault must be something you can undo.
Be able to list the dials — population, duration, severity, reversibility — and explain why you change only one of them between runs so a failure stays attributable.
Show the judgement: pick the smallest scope that could still falsify the hypothesis, recognise a null result caused by an undersized sample, and refuse irreversible faults on a money path outright.
Own the ceiling — how much error budget the programme may spend, what maximum blast radius any team may self-authorise, and where the platform enforces that limit rather than trusting each team to.
## Why the blast radius is the whole safety argument An experiment is only worth running if it might falsify your hypothesis — which means it might hurt. The blast radius is what converts "might hurt" into "hurt an amount we chose in advance". Every other safeguard is secondary: abort conditions limit how long the damage lasts, but the blast radius limits how much damage per unit time is even possible. ## The four dials **Population.** The strongest dial, and the first to reach for. Options, roughly in order of increasing exposure: an internal or synthetic client only; employees or a beta cohort; one tenant on a multi-tenant platform; one shard or partition; one instance out of N; a percentage of live traffic. If the service runs with N+1 or N+2 redundancy, removing one instance is by definition something the design already claims to tolerate — which is precisely the claim worth testing first. **Duration.** A ten-minute run with a hard, automatically enforced cap. Duration is often the cheapest dial to widen later, and the one people most often forget to bound: an experiment left running because everyone got distracted is a self-inflicted incident with no owner. **Severity.** The same fault type has a dial inside it. 100 ms of added latency is a different experiment from 2 seconds; degrading a dependency is different from blackholing it. Start at the mildest setting that could still move the steady-state metric. **Reversibility.** A dial with essentially one acceptable setting. The fault must be removable by a single action, and that removal must be tested before injection. Faults that destroy data, move money, or send real notifications to real customers are not chaos experiments; they are damage. Time of day is not a fifth dial so much as a precondition: business hours, owning team present, not during peak season, not mid-launch, not during a change freeze, and not while an unrelated incident is in progress. ## The floor: small enough to be safe, large enough to be readable Minimising blast radius has a lower bound that people miss. If you expose 0.05% of traffic for two minutes and report "no impact detected", the honest reading is usually *this experiment could not have detected the impact*, not *there is no impact*. The signal has to clear the noise in your steady-state metric. So the rule is the smallest scope that could **still falsify the hypothesis**. When the population you can safely expose is too small to produce a readable signal, the usual moves are to extend duration rather than population, to narrow the steady-state metric to just the exposed cohort so its denominator is small, or to run against a concurrent control group so the comparison is like-for-like instead of against a noisy global baseline. ## Widening: one dial per run After a pass, expand deliberately: 1. Population, in steps — 1% → 5% → 25% → one full zone. 2. Then duration. 3. Then severity. One dial at a time, because failures are frequently **nonlinear**. Killing one instance out of six proves nothing about killing three, because the remaining capacity crosses a threshold: connection pools saturate, autoscaling reacts too slowly, a shared downstream that absorbed the redirected load at 1% collapses at 25%, caches that stayed warm go cold. If you widen two dials at once and the run fails, you cannot attribute the failure. Before each widening, apply the honest test: *if this run falsifies the hypothesis at this new scope, am I willing to have caused that impact?* If the answer is no, you are not ready to widen — you are gambling. ## The programme-level ceiling Beyond the single run, decide in advance what fraction of the service's error budget the chaos programme may consume in a period, and treat it as a hard ceiling rather than a guideline. That converts an unbounded appetite for confidence into a number the product owner can agree to, and it gives you an unambiguous rule for when to stop: if the budget is already spent on real incidents, the experiments pause. It also makes the tradeoff visible in the right direction — an experiment that finds a genuine weakness usually saves far more budget than it spends, but only if the spending is deliberate. ## For the payment service specifically Payment is the case where reversibility dominates everything. Do not inject into the money-movement path itself on the first run. Scope to a read-side or an ancillary dependency — the fraud-scoring call, the receipt service, the rate lookup — assert a hypothesis about checkout success rate, and pick a fault the payment provider will never see. Then widen only when you have evidence, from the earlier runs, that the failure handling around the critical path behaves as documented.
- The experiment passed at 1% of traffic. Do you go straight to 100%?No — widen one dial at a time, re-testing the hypothesis at each step. Failures are nonlinear: a shared downstream that absorbed redirected load at 1% can collapse at 25%, caches go cold, connection pools saturate, and autoscaling that kept up at small scale lags at large. Each step is also a fresh decision about whether you would accept the impact if it fails.
- How do you know when the blast radius is too small to be useful?When the effect you would call a failure is smaller than the noise in your steady-state metric over the run window — then "no impact detected" is uninformative. Fix it by extending duration rather than exposure, by scoping the metric to just the exposed cohort so the denominator is small, or by running a concurrent control group so the comparison is like-for-like.
- Where does the ceiling on blast radius ultimately come from?From the impact the business has agreed to accept — expressed as a fraction of the error budget the chaos programme may consume in a period. That makes the tradeoff explicit rather than an engineer's private judgement, gives a clear rule for pausing experiments when real incidents have already spent the budget, and keeps the programme from expanding until something expensive happens.
saying these in an interview costs you the question
- Minimizing blast radius just means running it in staging
- Bigger blast radius finds more problems, so start big
- We can always stop it manually if it goes wrong
- Any experiment is safe as long as it's short
- Passing at 1% means it will pass at full scale