Netflix's Chaos Monkey randomly terminates running instances in production. What reliability property does that specifically verify, and which common failure modes does randomly killing an instance never exercise?
answer
- one replica gone, nothing else
- fail-stop is the friendliest failure
- proves replaceable, not resilient
- grey failure stays invisible
- your deploys already kill pods daily
basics
~20 sKilling a random instance verifies only that losing one replica is a non-event: it drops out of the load balancer, traffic reroutes, capacity is replaced. It never exercises partial failure such as slow dependencies, resource exhaustion, or a whole zone going away.
solid answer
~50 sChaos Monkey injects one fault: a clean, sudden process or instance death. It proves the service is genuinely replaceable — health checks evict the dead instance quickly, service discovery stops routing to it, surviving replicas absorb the traffic, the scheduler replaces it, and nothing important was pinned to that host (no in-memory session state, no singleton cron, no leader that cannot re-elect). That is a real and useful property, but it is the *friendliest* failure a distributed system has: fail-stop. The expensive outages are grey failures — a dependency that answers in 8 seconds instead of 80 milliseconds, a replica returning 500s while still passing its health check, a disk filling, an entire availability zone degrading at once. None of those look like a missing instance. On a modern platform that reschedules pods during every deploy and every node scale-down, you are arguably already passing the instance-kill test daily for free.
go deeper
Be ready to say plainly what a random instance termination checks: the replica leaves the load balancer, other replicas take the traffic, a new one is started, and no data lived only on that box.
Explain the detection-to-replacement chain with timings — health check interval times failure threshold sets how long dead endpoints keep receiving traffic — and contrast fail-stop with grey failure.
Show judgment about yield: on a platform that recycles pods every deploy you already pass this test daily, so argue for dependency latency and zone loss as the experiments that carry real untested risk.
Own the portfolio question — which fault types a chaos program should fund first across many teams, and how to avoid a program that reports green because it only ever injects the failure mode the platform already handles.
## What the tool actually does Chaos Monkey, built at Netflix and open-sourced as part of the Simian Army, picks instances at random from a group and terminates them. Two design choices matter more than the tool itself: it runs **in production**, and it runs **during business hours**, so a human is awake to see the consequence. Everything else — the schedule, the opt-out list, the group targeting — is packaging around a single fault type: an instance disappears without warning. ## The property it verifies: replaceability A sudden termination is a *fail-stop* fault. The instance is working, then it is gone; it never lies, never answers slowly, never returns partial results. For the service to shrug that off, a chain of things must all be true: - **Detection.** The load balancer or service-discovery layer must notice within seconds, not minutes. If health checks run every 30 seconds with a threshold of 3, you are advertising a dead endpoint for a minute and a half. - **Request loss is survivable.** Every in-flight request on that instance dies. Callers must retry, and the operation must be safe to retry. - **Capacity absorbs the loss.** The surviving replicas must handle the redistributed traffic without tipping over — which is really a headroom question in disguise. - **Replacement is automatic.** An autoscaling group or a controller notices the shortfall and starts a new instance, and that instance warms up and joins on its own. - **Nothing was pinned there.** No in-memory session that only lived on that host, no local scratch file another request expected to find, no singleton scheduled job, no leader that cannot be re-elected quickly. That last bullet is the one that catches real teams. "We're stateless" is a claim that decays quietly — someone caches a user session locally to shave a redis call, someone runs the nightly reconciliation on whichever box holds a lock file. Random termination turns statelessness from an architecture-diagram assertion into a property that is continuously re-verified. ## What it structurally cannot see Because it only produces fail-stop, everything in the *grey* zone is invisible to it: - **Latency.** A dependency that slows from 50 ms to 5 s exhausts thread pools and connection pools and takes down callers that a dead instance never would. - **Errors without death.** A replica that returns 500s while still answering its health check keeps receiving traffic. Instance kill is the opposite scenario — the instance stops answering everything, including the health check. - **Blackholing.** Packets dropped silently, so the caller hangs until its timeout instead of getting a fast connection refused. - **Resource exhaustion.** Disk full, memory pressure, file descriptors exhausted, CPU saturated. The process is alive and degrading, which is far harder to handle than absent. - **Correlated failure.** Losing one instance tests nothing about losing thirty at once. Zone-level and region-level loss is a different experiment with different math, because it removes a third or a half of capacity simultaneously and puts the survivors under a step change in load. - **Dependency failure.** Your service is fine; the thing it calls is not. Most real incidents live here. ## Why it looks weaker in 2020s infrastructure than in 2011 When Chaos Monkey appeared, terminating a VM was a rare and scary event. On a container platform, pods are killed constantly and for free: every rolling deploy replaces every pod, node autoscaling drains hosts, preemptible or spot capacity is reclaimed with a short warning, and evictions happen under node pressure. If your platform does that weekly, a random pod kill is a test you are already passing continuously and it produces very little new information. Saying this in an interview signals that you understand the tool as one fault type in a catalogue rather than as the definition of chaos engineering. ## When it is still the right first experiment It earns its place when the fail-stop assumption is genuinely untested: a newly stateful service, a service with a singleton or leader, long-lived connections such as WebSocket or streaming gRPC (where clients must reconnect and rebalance rather than just retry a request), legacy VM fleets that are only replaced during releases, or anything with a slow warm-up where losing a replica means the replacement is cold and slow for minutes. In those cases, "one instance vanishes" is a hypothesis you have not actually checked. ## How to answer it Name the property (replaceability under fail-stop), name what it therefore cannot reach (grey failure, saturation, correlated loss, dependency failure), and finish with the judgment: on a platform that already recycles instances constantly, instance kill is a low-yield experiment and dependency latency or zone loss is where the untested risk actually sits.
- If a service passes instance-kill experiments consistently, what would you inject next and why?Dependency faults, and latency before errors. Real incidents come from things the service calls degrading, not from a replica vanishing. Latency injection exhausts thread and connection pools and reveals missing or overlong timeouts, which a clean kill never touches. After that, correlated loss — take out a whole zone — because that is the only fault that also tests capacity headroom.
- Chaos Monkey deliberately runs during business hours rather than overnight. Why is that a feature and not an oversight?An experiment is only useful if someone is watching. Running in working hours means the on-call engineer and the owning team are awake, the impact is observed within seconds, and a bad result can be stopped immediately. Overnight chaos converts a controlled experiment into an ordinary outage: same damage, slower detection, sleep-deprived responders, and much less learning.
- A team says their instance-kill experiment passed, but during the real node failure last month they had a five-minute outage. What likely differed?A graceful terminate is not what a node failure does. Termination usually sends SIGTERM and lets connections drain; a node dying takes the kubelet and the health reporter with it, so eviction depends on node-heartbeat timeouts measured in tens of seconds, and sockets hang rather than closing. To model that, drop the traffic silently or halt the machine rather than asking the process to exit.
saying these in an interview costs you the question
- Chaos engineering means randomly killing things in production
- If instance kills pass, the service is resilient
- A killed instance and a slow instance look the same to callers
- Random kills also test capacity headroom
- Chaos experiments should run at night to limit impact