skip to content

Fault Injection Techniques

The catalog of failures worth injecting and what each one proves. Interviewers ask which faults you would test first for a given architecture — dependency and zone failures usually beat killing random pods.

on this pageshow

questions

5

You are starting a chaos program for a payment API that runs on Kubernetes across three availability zones and calls twelve downstream services. Which faults would you inject first, and why not start by killing random pods?

level: seniorimportance: must knowfreq 62%

answer

  1. order by untested assumption, not drama
  2. twelve dependencies, one claim unverified
  3. is the optional call really optional
  4. three zones, one gone, 1.5x each
  5. the platform already kills pods for you

basics

~20 s

Start with dependency faults — blackhole each downstream one at a time to check that every call you call optional really is — then latency on the busiest ones, then losing one whole zone. Random pod kills come last because rolling deploys and node scale-down already kill pods daily.

solid answer

~50 s

I order experiments by untested risk, not by novelty. First, dependency faults: blackhole each of the twelve downstreams in turn and check the claim you have never verified — that the eight you call non-critical really are optional. Teams routinely find an "optional" fraud-scoring or analytics call sitting synchronously on the payment path with no timeout. Second, latency rather than errors on the critical few, injected just above the configured timeout, because slow calls exhaust shared pools and stall unrelated endpoints. Third, correlated loss: take one availability zone out. With three zones, losing one moves each survivor from about a third of traffic to a half — a 1.5x step — so it is the only experiment that tests capacity headroom and cross-zone data paths. Fourth, a database primary failover. Random pod kills sit last: every deploy and node scale-down already runs that experiment for free.

go deeper

for a junior

Be able to say that the things a service calls fail more often than the service itself, so testing what happens when a dependency is unreachable comes before testing what happens when your own instance dies.

for a middle

Explain why blackholing a dependency is harsher than returning an error, and why latency saturates shared pools while a fast error does not.

for a senior

Carry the numbers: three zones means a 1.5x step onto each survivor, so steady-state utilization must stay under about two-thirds, and say what you would change when the experiment shows saturation.

for a principal

Own the program design — the sequence teams follow by default, what must be verified before a service is allowed on the money path, and how you keep the portfolio from drifting into experiments the platform already passes for free.

## The ordering principle The question behind "what would you inject first" is not "what is most dramatic" but **where is the largest untested assumption, and what does it cost to be wrong there**. For a payment API, being wrong costs money and trust, so both parts of that sentence matter: you go after big unverified claims, and you scope each experiment so a discovered defect is a small event rather than a large one. Rank candidate faults by two axes: how likely the mechanism is to be exercised in a real incident, and how confident you are today that it works. Anything the platform already exercises daily scores near zero on the second axis, no matter how alarming it sounds. ## First: dependency faults, one downstream at a time Twelve downstreams is the story here. Almost every team can name which of them are "critical" and which are "nice to have" — and almost no team has verified the second list. Blackhole one dependency at a time (drop its traffic silently rather than returning an error, because a hang is the more punishing and more realistic fault) and assert the hypothesis explicitly: *a payment still completes successfully, within the normal latency budget, with the non-critical dependency unreachable*. The classic findings are worth naming in an interview because they are what actually happens: - An "optional" fraud-signal, analytics, or personalization call turns out to be synchronous on the request path with no timeout, so an optional dependency can take payments to zero. - The fallback exists but is never reached, because the exception type thrown on a socket hang is not the one the catch block handles. - The call is properly optional but its failure is counted as a payment failure, so an outage in a cosmetic service burns the payment SLO's error budget. Run this for all twelve, not just the interesting ones. The value is in the ones you were confident about. ## Second: latency on the critical few Having established what is truly required, inject delay — not errors — into the required dependencies. Errors return fast and stay contained in one request; latency holds a thread and a connection, and by `concurrency = throughput x latency` a 20x slowdown demands 20x the concurrent slots. That is how a single slow dependency saturates a shared pool and stalls endpoints that never call it. Inject at a value just above the configured timeout so the timeout path, the retry policy, and the fallback all actually fire, and inject on a fraction of calls as well as on all of them, since a single degraded replica behind a load balancer is more common than a clean outage. ## Third: lose an availability zone This is the highest-value structural experiment and the one instance kill can never approximate, because the fault is **correlated**. Three zones, one gone, means the survivors go from carrying roughly 33% of traffic each to roughly 50% each — a 1.5x step change applied instantly. If each zone was running at 70% of capacity, the survivors need 105%, and you have converted a zone failure into a total outage. Sizing for single-zone loss across three zones means keeping steady-state utilization under about two-thirds; that is what N+1 redundancy means in concrete numbers here. Zone loss also probes things nothing else does: whether a stateful dependency's replicas were actually spread across zones rather than co-located by an unnoticed scheduling constraint, whether cross-zone calls exist that you believed were zone-local, whether autoscaling can acquire capacity in the surviving zones quickly enough, and whether anything holds a hardcoded zone-specific endpoint. Inject it by cutting network reachability to the zone rather than by deleting resources — it is more faithful to a real zone impairment and it is reversible in one step. ## Fourth: stateful failover Fail over the database primary, or force a leader re-election in whatever holds consensus. This is the fault with the longest tail of surprises: connection pools that cache a dead endpoint until the process restarts, a failover that takes 40 seconds while callers time out at 2, writes accepted during the window and lost, and application code that reconnects but never re-issues the transaction it was in the middle of. For a payment system this is the one where correctness, not just availability, is at stake. ## Last: random pod termination On Kubernetes, pods are already terminated constantly — every rolling deploy replaces every pod, node autoscaling drains hosts, and preemption reclaims capacity. You pass this experiment daily whether you run it or not, so a deliberate pod kill mostly confirms something already continuously verified. It is not worthless, but it is a low-yield place to spend your first quarter of a chaos program, and choosing it first is the tell that a candidate has read about chaos engineering rather than run it. ## Doing this on a payment path without breaking payments Two practical commitments belong in the answer. Every injection must be **revocable in one action and self-expiring** — a rule that outlives the agent that applied it is an incident, so give it a TTL. And you scope the first run of each experiment narrowly: a subset of traffic or a single zone's replicas before all of them, with the owning team watching live. That is not timidity; on a money path it is what makes running in production defensible at all.

  • How would you word the hypothesis for the first dependency experiment so the result is unambiguous?
    As a measurable prediction with a threshold, not a vibe. For example: with the fraud-scoring service unreachable, the payment success rate stays above 99.9% and p99 checkout latency stays under 800 ms for the injection window. That gives a clear pass or fail, names the metric to watch, and defines when to stop early. "We think it degrades gracefully" cannot be falsified.
  • The zone-loss experiment shows the surviving zones saturating. What do you actually change?
    Either add headroom or reduce what has to survive. Concretely: cap steady-state utilization so any one zone can be lost — under about two-thirds with three zones — pre-provision rather than relying on autoscaling to react within the outage, and check that scaling can actually acquire that capacity. If neither is affordable, define what you shed first so the degradation is chosen rather than random.
  • Your twelve dependencies include one that is genuinely critical and has no alternative. Is injecting a fault into it worth it?
    Yes, but the hypothesis changes. You are no longer testing whether payments survive, because they will not. You are testing failure quality: does the service fail fast instead of hanging, does it return a correct status the client can act on, does it stop retrying into a dead dependency, does the right alert fire, and does the runbook match reality. Recovery behaviour after the fault is lifted matters just as much.

saying these in an interview costs you the question

  • Chaos engineering starts with randomly killing pods
  • Every dependency we call optional is already optional
  • Losing one of three zones adds a third to each survivor
  • Autoscaling will cover the capacity gap during a zone loss
  • Injecting errors is enough; latency is a subtler version of the same test

context

open as a page

Netflix's Chaos Monkey randomly terminates running instances in production. What reliability property does that specifically verify, and which common failure modes does randomly killing an instance never exercise?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Killing a random instance verifies only that losing one replica is a non-event: it drops out of the load balancer, traffic reroutes, capacity is replaced. It never exercises partial failure such as slow dependencies, resource exhaustion, or a whole zone going away.

open as a page

When injecting a fault into a service's call to a downstream dependency, what is the practical difference between injecting added latency and injecting error responses, and why does latency injection usually uncover more bugs?

level: middleimportance: should knowfreq 50%

basics

~20 s

Error injection returns a failure fast, so it tests the error-handling branch. Latency injection holds each call open, so it consumes threads, connections, and memory across the whole service. Latency finds more bugs because slowness spreads to unrelated work while a fast error stays local.

open as a page

Beyond terminating processes and injecting latency, resource-exhaustion faults deliberately starve a host or container of CPU, memory, disk space, or file descriptors. What does each of those prove that a process kill does not?

level: middleimportance: should knowfreq 40%

basics

~20 s

Each starves a different resource and produces a different degradation. CPU starvation causes queuing and timeouts while the process stays healthy; memory pressure triggers GC thrash or an out-of-memory kill; a full disk breaks writes and logging; exhausted file descriptors block new connections. All keep the process alive and misbehaving.

open as a page

To make a downstream dependency fail during a chaos experiment, you can inject the fault in the caller's own client code, at a sidecar or mesh proxy, or at the host network layer with firewall or traffic-control rules. How do you choose, and what does each level change about what the experiment can prove?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Choose the lowest layer that can produce the fault your hypothesis is about, and the highest that still lets you target and revoke it safely. Client-code injection targets precisely but skips DNS, connect, and TLS. Proxy injection is realistic on the wire and centrally revocable. Network rules are the most faithful and the hardest to undo.

open as a page