skip to content

To make a downstream dependency fail during a chaos experiment, you can inject the fault in the caller's own client code, at a sidecar or mesh proxy, or at the host network layer with firewall or traffic-control rules. How do you choose, and what does each level change about what the experiment can prove?

level: principalimportance: nice to knowfreq 30%

answer

  1. fidelity up, control down
  2. above the socket sees no handshake
  3. the pool bug hides above the pool
  4. the mesh misses your database traffic
  5. who removes it if you vanish

basics

~20 s

Choose the lowest layer that can produce the fault your hypothesis is about, and the highest that still lets you target and revoke it safely. Client-code injection targets precisely but skips DNS, connect, and TLS. Proxy injection is realistic on the wire and centrally revocable. Network rules are the most faithful and the hardest to undo.

solid answer

~50 s

Each level trades fidelity against control. Injecting in the caller's client code gives per-call, per-tenant precision and is trivially revocable with a flag, but the fault is one your own code fabricated above the socket, so it never exercises DNS resolution, connection establishment, TLS, or the connection pool itself — and if the bug lives in the pool, you will not find it. A sidecar or mesh proxy injects real delays and real HTTP status codes on the wire, per route and per percentage, with no application change and central revocation; its limits are that it only covers traffic that traverses it, which usually excludes direct database connections and non-HTTP protocols. Host network rules reach everything — packet loss, silent blackholes, asymmetric partitions, failures during the handshake — but target coarsely and are the hardest to undo, so they must carry a self-expiring bound.

code

bash · 6 lines
bash
# Network-level injection with a deadman: the rule removes itself
# even if the shell that created it dies.
iptables -A OUTPUT -p tcp -d 10.0.3.17 --dport 5432 -j DROP
at now + 5 minutes <<'EOF'
iptables -D OUTPUT -p tcp -d 10.0.3.17 --dport 5432 -j DROP
EOF

go deeper

for a junior

Know there is more than one place to break a dependency — in your own code, in a proxy in front of it, or in the network — and that the closer to the network you go, the more realistic and less targeted it becomes.

for a middle

Explain what an application-level injection skips: DNS, connect, TLS and the connection pool, and therefore which classes of defect it structurally cannot find.

for a senior

Match the level to the hypothesis and defend the safety story — narrow scope, central revocation, and a self-expiring rule so an abandoned injection cannot become an incident.

for a principal

Own the standard for the estate: which level teams may use by default, who holds privileged network injection, how experiments are revoked centrally, and how a program earns its way down the ladder as tooling matures.

## The choice is fidelity versus control There is a ladder from "inside the application" to "in the network", and it runs consistently in one direction: as fidelity to real-world failure rises, precision of targeting and ease of revocation fall. Picking a level is picking a point on that trade, and the right point depends entirely on what your hypothesis says. ## Level 1: inside the caller A hook in the HTTP or RPC client — an interceptor, a wrapper, a flagged branch — decides that this particular call will throw or sleep. **What it buys.** The finest targeting available: one endpoint, one tenant, one percentage of calls, one shard. It needs no infrastructure privileges, so a developer can run it in a service they own. Revocation is a flag flip and takes effect immediately. **What it cannot prove.** Everything below the injection point is skipped. There is no DNS lookup to fail, no TCP handshake to hang, no TLS negotiation to break, and no connection actually taken from the pool. If your defect is that the pool has 20 slots and no acquisition timeout, injecting a sleep *above* the pool will never reveal it. You are also testing your own idea of how the dependency fails, which is exactly the assumption that tends to be wrong: teams inject a clean exception and discover in the real incident that the library threw a different type from a socket read, and their catch block never matched. ## Level 2: sidecar or mesh proxy A proxy in the request path applies a delay or returns an error status for matching routes — a service mesh's fault-injection rules are the common form, typically expressible as a fixed delay or an abort with a chosen HTTP status, applied to a percentage of matching requests. **What it buys.** The caller sees a genuine wire event: a real socket held open for the delay, a real status code parsed by the real client library, real pool occupancy. No application change is required, so it works uniformly across services written in different languages. Rules live in one place, so an operator can revoke every experiment in the estate with one action — which matters enormously when the person who started an experiment is unreachable. **What it cannot prove.** Only traffic that traverses the proxy is affected. Direct database connections, cache protocols, message brokers, and anything that bypasses the mesh are untouched, and those are frequently the dependencies you most want to test. Proxies also normalise: an abort is a well-formed HTTP response, not a half-written body or a connection reset mid-stream. And the proxy is now part of the experiment — a delay it imposes exercises its own buffering and timeout behaviour alongside your application's. For most organisations this is the right default level for a chaos program: enough realism to be worth running, enough central control to be safe. ## Level 3: host and network Traffic-control rules add delay, jitter, reordering, or packet loss; firewall rules drop or reject packets; security-group or network-ACL changes remove reachability wholesale. **What it buys.** The highest fidelity available short of unplugging something. It is the only level that can produce a true blackhole, where the caller hangs on connect or read until its own timeout — or, if it has none, until the OS gives up minutes later. It reaches DNS, handshake, and TLS failures. It can create asymmetric partitions, where A reaches B but B cannot reach A, which is the shape that breaks consensus systems and produces the most interesting split-brain behaviour. It works for every protocol, including the ones the mesh does not see. **What it costs.** Targeting is coarse — a rule usually applies per host and port, hitting every caller on that machine rather than the one you meant. It needs privileged access, which is a governance question as much as a technical one. And it is the most dangerous to leave behind: a rule that outlives the process that applied it is now an unexplained outage, and a host you have blackholed may not accept the connection you need in order to fix it. Every network-level injection needs a self-expiring mechanism — a timer that removes the rule regardless of what happens to the agent — and a recovery path that does not depend on the path you just severed. ## Level 4: the platform API Above all of these sits the provider's own control plane: stop these instances, detach that volume, apply a network ACL that isolates a subnet or a zone. Managed fault-injection services exist for exactly this. It is the coarsest level and the closest to the real event, and it is how you run a zone-loss experiment — cutting reachability at the network boundary rather than deleting anything. ## The decision rule Ask what the hypothesis is *about*, then pick the shallowest layer that cannot fake it: - "Our fallback returns cached data when recommendations return 503" → client or proxy is fine; the status code is the whole point. - "Our connection pool does not deadlock when the dependency hangs" → must be at least at the proxy, and preferably a network drop, because pool behaviour is below a client-level injection. - "A zone loss does not exhaust the survivors" → platform or network level; nothing smaller produces correlated loss. - "Our cluster does not split-brain under an asymmetric partition" → network level only. Then apply the second rule: at whatever level you chose, the injection must be **scoped as narrowly as the hypothesis allows** and **removable by someone who was not involved in starting it**. A program that begins at the proxy level and earns its way down to the network as trust and tooling mature is the pattern that survives contact with a real production estate.

  • Your hypothesis is that the client's connection pool does not deadlock when a dependency hangs. Why is injecting the delay in application code inadequate?
    Because a sleep or thrown exception in your own client wrapper usually sits above the pool, so no connection is ever checked out and held. The pool never fills, the acquisition path is never contended, and the acquisition timeout is never tested. To exercise it you need the fault below the pool — a proxy holding the response, or dropped packets so a real connection stays occupied until the socket times out.
  • What makes an asymmetric partition worth injecting, and where must it be injected?
    It breaks the assumption that reachability is mutual. Node A believes B is alive because its calls succeed, while B considers A gone and elects a new leader, so both sides act on incompatible views — the classic split-brain setup. Only network-level rules can produce it, since dropping traffic in one direction requires per-direction control that no application-level or proxy abort provides.
  • Why is a self-expiring bound more important for network-level injection than for a flag in the caller's code?
    Because the blast radius outlives the operator. A flag lives in a system you can still reach and flip from anywhere. A firewall rule persists on the host, may block the management path you would use to remove it, and stays active if the agent that applied it is killed or the operator loses connectivity. A timer that removes the rule regardless converts a potential outage into a bounded experiment.

saying these in an interview costs you the question

  • Injecting an exception in client code is equivalent to the dependency failing
  • A service mesh can inject faults into every dependency, including databases
  • Network-level rules can always be removed remotely when needed
  • Blackholing and returning an error produce the same caller behaviour
  • Higher fidelity injection is always the better choice

context