skip to content

Chaos engineering's principles call for running experiments in production. What does chaos testing only in a staging environment fail to catch, and when is staying out of production the right call?

level: seniorimportance: should knowfreq 52%

answer

  1. staging tests staging's configuration
  2. volume, data skew, warm caches, real third parties
  3. autoscaling and quota settings differ
  4. irreversible faults never go to production
  5. no usable SLI means not yet

basics

~20 s

Staging lacks production's traffic volume and mix, data size, real dependency versions, scaling settings and — above all — its configuration, which is where most failures live. A green staging run is evidence about staging. Stay out of production when the fault is irreversible or you cannot yet measure harm.

solid answer

~50 s

The principle is "run close to production" because the failure modes worth finding are properties of production's *state*, not its code: real traffic concurrency and mix, real data volume and skew, warm caches, real third parties and their timeouts, real autoscaling and quota settings, cross-zone latency, and the config that differs from staging in ways nobody has enumerated. Staging reproduces the code and almost none of that, so a passing staging experiment supports a claim about staging. It is still worth running — as a rehearsal for the experiment harness and the abort path, cheaply, before you point it at customers. Legitimate reasons to stay out: the fault is irreversible or touches money or personal data; you have no SLI good enough to define steady state, so you couldn't detect harm; a known unfixed weakness; peak season or a change freeze. Otherwise, go to production at 1%.

go deeper

for a junior

Know that chaos experiments aim at production because staging doesn't have real traffic, real data or real dependencies — and that they are still scoped to a small slice of it.

for a middle

Be able to list concrete production-only properties — data volume and skew, warm caches, third-party behaviour, autoscaling settings, configuration — instead of saying production is "more realistic".

for a senior

Show both sides of the decision: what a green staging run does and doesn't license you to claim, and the specific conditions (irreversibility, no usable SLI, freeze, exhausted budget) under which you would refuse to go to production.

for a principal

Own the entry criteria for the estate — which services have measurement good enough to experiment on at all, what sign-off a production run needs, and how regulatory constraints are handled before a team designs the experiment.

## What the principle actually claims "Run experiments in production" is often heard as bravado. It isn't. The claim is narrow and empirical: the behaviours a chaos experiment is trying to observe are properties of the running production system, and most of them do not exist anywhere else. Testing elsewhere doesn't make the experiment safer in any meaningful sense — it makes it answer a different question. ## What staging does not have **Traffic volume and mix.** Concurrency-driven failures — lock contention, connection-pool exhaustion, thread starvation, queue growth — appear at a load level staging never reaches. And the *mix* matters as much as the volume: the odd 2% of requests that hit the expensive path are usually absent from a synthetic load profile. **Data size and skew.** A query that is instant against 10k rows and catastrophic against 200M is a production-only failure. So is the one hot tenant whose partition is fifty times the median. **Real dependencies.** Staging typically points at mocks, sandboxes, or the same third party's test endpoint with different capacity, different rate limits and different latency. Payment providers, identity providers and geolocation services behave differently in their sandboxes almost by design. **Cache state.** Production runs warm. A cold-cache stampede after an instance is replaced is a real and common failure mode that a staging environment with no meaningful cache population cannot exhibit. **Capacity and elasticity settings.** Instance sizes, replica counts, autoscaling thresholds and cooldowns, cloud quotas — these are usually different by an order of magnitude, and the interesting question (does the system scale fast enough to absorb losing a zone?) is entirely determined by them. **Topology.** Cross-zone latency, real DNS and load-balancer behaviour, real network policy. **Configuration.** The big one. A large share of production incidents are configuration, and configuration is precisely what differs between environments in undocumented ways. An experiment in staging exercises staging's config. So the honest summary: a passing staging experiment tells you the *experiment* works and that the code path handles the fault in principle. It does not tell you production survives it. ## When staying out of production is right This is a decision with a cost on both sides, not a dogma. Reasons to stay out, or to defer: - **Irreversibility.** The fault destroys data, moves money, or sends real communications to real people. No blast radius makes that acceptable. - **You can't measure harm.** If there is no SLI good enough to express a steady state, you cannot tell whether the experiment hurt anyone, so you cannot run it responsibly. Building the measurement is the prerequisite, and it is often the more valuable work anyway. - **A known unfixed weakness.** Fix it first; experiment afterwards to verify. - **Timing.** Peak season, a live launch, an active incident, a change freeze, or an already-exhausted error budget. - **Regulatory or contractual constraints** on deliberately degrading a regulated workload — real, and worth surfacing early rather than discovering mid-review. - **The harness is unproven.** Which is exactly what staging is for. ## The sensible sequence 1. **Staging run** — debug the experiment tooling, confirm the fault injects and reverts, confirm the abort predicate fires. Treat the output as evidence about the *harness*, not about production resilience. 2. **Production, minimum scope** — 1% of traffic or a single instance, ten minutes, owners watching, automated abort armed. 3. **Widen deliberately**, one dial at a time. Notice that "run in production" and "expose everyone" are unrelated ideas. Running in production at 1% of traffic with an automated halt is a far smaller risk than most routine deploys, and it is the only place the answer is real. ## Framing it in an interview The strong answer names the specific production-only properties (data volume, config, third parties, cache warmth, autoscaling settings) rather than asserting that production is "more realistic", and then immediately supplies the guardrails — blast radius, abort, reversibility, timing — so the position reads as considered rather than reckless. The weak answer is either "we'd never do that in production" or "just do it in prod, that's the point". Both skip the decision.

  • What must be true before a team's very first production chaos experiment?
    A steady-state metric they can actually measure, alerting that would notice harm, a fault with a tested one-action undo, an automated abort armed on a pre-agreed threshold, a scoped population, the on-call informed, and an accepted budget for the impact. Missing measurement is the usual blocker — and building it is worth more than the experiment.
  • Is there any value in running chaos experiments in staging at all?
    Yes, but for a different purpose: rehearsing the harness. Staging is where you confirm the fault injects and reverts cleanly, the abort predicate fires on the right threshold, and permissions and tooling work — cheaply, before any of it is pointed at customers. Just don't read a green staging run as evidence that production survives the same fault.
  • A team argues their staging environment is a full production clone, so production experiments are unnecessary. How do you respond?
    Ask what traffic it serves and what data it holds. Clones match topology, rarely load, almost never data volume and skew, and never the live third parties or warm caches. Also ask when the config last diverged — if nobody can answer, that gap is exactly where production incidents come from, and the clone cannot test it.

saying these in an interview costs you the question

  • Production chaos is reckless; staging proves the same thing
  • Our staging is a full clone, so results transfer
  • Running in production means exposing all users
  • Real dependencies behave the same in their sandboxes
  • Run it in production before you can measure the impact

context