When injecting a fault into a service's call to a downstream dependency, what is the practical difference between injecting added latency and injecting error responses, and why does latency injection usually uncover more bugs?
answer
- fast failure versus held resource
- errors stay local, slowness spreads
- concurrency equals throughput times latency
- inject just past the configured timeout
- the call with no timeout at all
basics
~20 sError injection returns a failure fast, so it tests the error-handling branch. Latency injection holds each call open, so it consumes threads, connections, and memory across the whole service. Latency finds more bugs because slowness spreads to unrelated work while a fast error stays local.
solid answer
~50 sAn injected error returns immediately: the caller gets a 503 or a refused connection, takes its error path, and the request is finished. It tests whether the fallback, the retry policy, and the error accounting are correct — a narrow, valuable check. Injected latency does something structurally different: it makes every in-flight call *occupy* something. By Little's law, required concurrency equals throughput times latency, so a dependency that goes from 100 ms to 2 seconds at 200 requests per second needs 400 concurrent slots instead of 20. Your pool of 50 threads or connections saturates, and requests that never touch that dependency start queuing behind it. That is how a single slow downstream turns into a full outage. The most informative value to inject is the one just above your configured timeout — which is also how you discover which calls have no timeout at all.
code
bash · 6 lines# Add 500ms +/- 100ms of egress delay on a host, then remove it.
# Every outbound call from this host now answers late but correctly.
tc qdisc add dev eth0 root netem delay 500ms 100ms distribution normal
# Always pair an injection with its removal, and give it a deadline.
sleep 120; tc qdisc del dev eth0 rootgo deeper
Know that an injected error returns immediately while injected latency keeps the call open, and that holding calls open ties up threads and connections other requests need.
Do the arithmetic out loud: concurrency equals throughput times latency, so name the pool size that saturates and explain how one slow dependency stalls unrelated endpoints.
Show how you pick the injected value against the configured timeout, and describe what you would change after finding an unbounded call — isolation, a bounded wait, and a retry budget.
Own the standard: which dependencies must be isolated by default across the estate, what a required timeout policy looks like, and how you verify the policy holds rather than trusting per-team configuration.
## Two faults that feel similar and behave nothing alike Both experiments target the same dependency, but they exercise different machinery. **Error injection** makes the dependency answer wrongly and quickly: an HTTP 500 or 503, a connection refused, a gRPC `UNAVAILABLE`. The call completes almost instantly. The resources it held are released. What is under test is the code path *after* the failure: is there a fallback, is it correct, does the retry policy fire, does the failure get counted against the right metric, does the user see something sane? **Latency injection** makes the dependency answer *correctly but late*: a fixed or distributed delay before the normal response. Nothing errors. Every call simply stays open longer, and while it is open it holds a thread, a connection from the pool, a buffer, and a slot in whatever queue sits in front of it. ## Why slow beats broken, quantitatively Little's law is the tool for this: `concurrency = throughput x latency`. A service handling 200 requests per second against a dependency answering in 100 ms needs 200 x 0.1 = **20** concurrent slots. Inject 2 seconds of delay and the same throughput needs 200 x 2 = **400** slots. If the client pool is sized at 50, the other 350 requests queue, then time out, then — if a retry policy is naive — come back and make it worse. The critical consequence is that saturation is **shared**. A fast error is contained inside the one request that touched the failing dependency. A slow response consumes a resource that other, unrelated requests need. If the whole application shares one servlet thread pool or one HTTP connection pool, an outage in a non-critical recommendations call can stall checkout. That is the failure this experiment is designed to expose, and it is exactly the argument for per-dependency isolation of pools. ## Choosing the delay value: aim just past the timeout A random delay teaches you little. Pick values relative to the caller's own configuration: - **Just under the timeout** (say 90% of it) — proves the timeout is not accidentally firing on normal-ish slowness, and shows what the latency budget upstream looks like when a dependency is merely unhealthy. - **Just over the timeout** — the highest-value single value. It forces the timeout to fire and reveals whether the timeout path, the retry policy, and the fallback all behave. - **Far past the timeout, or indefinite** — this is where you discover calls with **no timeout configured**, which is the single most common finding of a latency experiment. If the injected delay is 60 seconds and requests hang for 60 seconds, some layer has no bound. Many client libraries default to an unbounded read timeout, so this is not a hypothetical. Also vary the *shape*: a fixed delay on 100% of calls is a clean signal, but a delay applied to a small percentage of calls models a partially degraded dependency — one bad replica behind a load balancer — and that reveals whether your alerting and your outlier ejection notice a tail problem at all, which an all-or-nothing outage would mask. ## What error injection is still uniquely good for Latency does not replace error injection; it answers a different question. Errors are the right fault when you want to test: - **Which errors are retried.** Retrying a 500 may be correct; retrying a 400 is a bug; retrying a non-idempotent write after a timeout can double-charge a customer. - **Correct status semantics.** Does a downstream 503 surface as a 503 or does it leak as a 500 and pollute your own error budget? - **Fallback correctness.** A fallback value that is silently wrong is worse than an error, and only an error path exercises it. - **Error-rate alerting.** Injecting a 5% error rate is the honest way to check whether your alerting actually fires at the threshold you claim. ## Blackhole: the third option people forget Between "fast error" and "slow success" sits the nastiest real-world fault: packets dropped with no response at all. A `REJECT` produces an immediate connection-refused — that is error injection at the network layer. A `DROP` produces a blackhole: the caller's TCP connect or read hangs until *its own* timeout, or until the OS TCP retransmission limit, which can be minutes. Blackholing is the most faithful model of a hung dependency, a lost network path, or an overloaded box that accepted the connection and never got to it, and it is what you inject when you want to know whether connection timeouts and read timeouts are both set. ## The order to run them in If you can only run one experiment against a dependency, run latency. It subsumes much of what error injection tests (once the timeout fires you are on the error path anyway) *and* it reaches the resource-saturation behaviour that errors never touch. Run explicit error injection afterwards to check the semantics of specific status codes and the correctness of fallbacks.
- You inject 30 seconds of latency into one dependency and the whole service stops answering, including endpoints that never call it. What does that tell you?That the slow calls and the unrelated requests share one exhaustible resource — typically a single HTTP connection pool, a shared thread pool, or the server's request-handling threads. The fix is isolation: give that dependency its own bounded pool so its saturation cannot consume slots the rest of the service needs, and bound the wait so callers fail fast instead of queuing.
- Why is injecting latency on only 5% of calls sometimes more informative than injecting it on all of them?It models a partially degraded dependency — one bad replica behind a load balancer — which is far more common than a total outage. It tests whether you detect tail latency at all: p50 barely moves while p99 explodes, so dashboards and alerts built on averages stay green. It also exercises outlier ejection and per-endpoint retries rather than an obvious all-or-nothing failure.
saying these in an interview costs you the question
- Injecting errors and injecting latency test the same thing
- A slow dependency is less dangerous than a failing one
- Timeouts always exist by default in HTTP clients
- Percentiles do not matter if the average looks fine
- Retrying a timed-out call is always safe