Beyond terminating processes and injecting latency, resource-exhaustion faults deliberately starve a host or container of CPU, memory, disk space, or file descriptors. What does each of those prove that a process kill does not?
answer
- alive and wrong, not absent
- throttled, not killed
- slow collapse before the OOM kill
- old connections fine, new ones refused
- how much runway does the alert buy
basics
~20 sEach starves a different resource and produces a different degradation. CPU starvation causes queuing and timeouts while the process stays healthy; memory pressure triggers GC thrash or an out-of-memory kill; a full disk breaks writes and logging; exhausted file descriptors block new connections. All keep the process alive and misbehaving.
solid answer
~50 sA process kill produces a clean absence. Resource exhaustion produces a process that is present, passing health checks, and wrong — which is far harder to handle. CPU starvation shows what happens at saturation: queues build, latency climbs non-linearly past the knee, and throttling on a container limit causes stalls that look like network problems. Memory pressure exposes whether the container dies at its limit with an out-of-memory kill or slowly garbage-collects itself to a standstill first. Disk fill is the one almost nobody has tested and the one that produces the strangest outages: logging blocks, temp files fail, a database refuses writes. File-descriptor or connection exhaustion stops new connections while existing ones keep working, so the service reports healthy while rejecting all new traffic. Each has a distinct signal, and each tests whether you alert on the resource before it runs out.
code
bash · 5 lines# Saturate 4 cores at ~90% load for one minute, then stop on its own.
stress-ng --cpu 4 --cpu-load 90 --timeout 60s
# Push memory toward the container limit for one minute.
stress-ng --vm 2 --vm-bytes 90% --timeout 60sgo deeper
Know that starving CPU, memory, disk, or file descriptors leaves the process running and misbehaving rather than making it disappear, and that this is harder for load balancers to route around.
Explain each resource's distinct signature: throttling stalls under a CPU limit, garbage-collection thrash before an out-of-memory kill, blocked writes on a full disk, and refused new connections when descriptors run out.
Frame the hypothesis around detection and runway — does the alert fire early enough to act, is the symptom diagnosable in minutes, and does the instance shed load instead of accepting work it cannot finish.
Own the platform defaults: which saturation signals every service must expose, what limits and thresholds are set centrally, and how recovery paths avoid depending on the very resource that ran out.
## Why exhaustion is a separate class of fault Killing a process is *fail-stop*: the thing is gone, everyone detects it, traffic reroutes. Resource exhaustion is *fail-slow*. The process is running, it answers its health check, and it is doing its job badly or partially. Orchestrators and load balancers are built to route around absence, not around degradation, so exhaustion faults test the parts of your system that only work when something is broken in a nuanced way. ## CPU starvation Burn CPU on the host or inside the container and watch what saturation actually does. The important behaviour is non-linear: queueing means latency stays roughly flat as utilization rises, then climbs steeply past the knee of the curve — typically somewhere in the 70–85% range depending on variability — so a service that looks fine at 65% can be unrecoverable at 85%. In a container there is a second mechanism worth knowing: exceeding a CPU limit does not kill anything, it *throttles*. The kernel simply stops scheduling the process for the remainder of each accounting period. From inside, that looks like inexplicable multi-millisecond stalls — garbage-collector pauses that make no sense, timers firing late, health checks occasionally timing out — and teams routinely misdiagnose it as a network problem. Injecting CPU pressure deliberately is how you learn to recognise the signature and confirm that you alert on throttling, not just on CPU usage. ## Memory pressure Allocate until the container approaches its limit. Two very different outcomes are possible and you want to know which one you get. If the runtime hits the hard limit, the kernel's out-of-memory killer terminates the process — abrupt, but at least it is fail-stop and the orchestrator restarts it. The nastier outcome is the long approach: a managed runtime that garbage-collects more and more frequently, spending an increasing fraction of every second in collection, so throughput collapses while the process remains alive and "healthy" for minutes. Injecting this proves whether your limits produce a fast, visible death or a slow invisible one, and whether the restart that follows is safe — a process killed mid-write may leave state that the next start must reconcile. ## Disk fill This is the least-tested and most consistently surprising fault. Fill the volume and the effects fan out in ways that have nothing to do with the application's logic: log writes block or throw, temporary files cannot be created, an embedded database or a broker refuses writes and may go read-only, a checkpoint or a WAL cannot be flushed, and — a favourite — the agent that ships logs off the box was itself the thing that would have freed the space. The real lesson is usually about *time*, not the failure: how much runway do you have between the disk-usage alert firing and the disk being full? An alert at 90% on a volume growing 3% an hour gives three hours; the same alert on a volume that just started receiving debug-level logging gives twenty minutes. Injecting a controlled fill is how you find out whether the runbook, the alert threshold, and the rate of growth are compatible. ## File descriptors and connection exhaustion Exhausting descriptors produces a uniquely misleading state: existing connections keep working perfectly, and every *new* connection fails. Health checks that reuse a connection stay green. Metrics still flow. Meanwhile new clients cannot connect at all. This is the fault that proves whether your monitoring covers *new-connection success* rather than only the health of established traffic, and whether descriptor limits are alerted on as a percentage of the limit rather than watched as a raw number nobody reads. The same shape applies one level up, to a connection or thread pool that is fully checked out: the service is alive and rejecting or queueing everything. ## What to assert in the experiment For each of these, the hypothesis is rarely "nothing bad happens" — something bad will happen. Useful hypotheses are about **detection and containment**: - The resource alert fires *before* the resource is exhausted, with enough runway for a human to act. - Degradation is bounded: the affected instance leaves rotation or sheds load rather than accepting work it cannot finish. - Recovery is automatic once the pressure is removed, without a manual restart. - The signal is diagnosable — a responder looking at the dashboards can tell which resource ran out within a couple of minutes. That last point is the honest reason to run these experiments even when you are fairly sure of the outcome: they are how you find out whether the symptom is *recognisable*, and recognition is most of mean time to recovery. ## Safety Exhaustion faults are the easiest to leave running by accident and the hardest to undo remotely — a box with no descriptors left may refuse your SSH connection, and a full disk can block the very tooling you would use to clean it. Every injection needs a self-expiring bound and a recovery path that does not depend on the resource you just consumed. Reserving a small pre-allocated file you can delete to reclaim disk space is an old trick that exists for exactly this reason.
- A container is well under its CPU limit on average but users report random multi-hundred-millisecond stalls. What would you suspect and how does injecting CPU pressure help?CPU throttling against the limit within short accounting periods — average utilization hides bursts that exhaust the quota, and the kernel simply stops scheduling the process for the rest of the period. Injecting a controlled burn reproduces the signature so you can confirm it, and it verifies that you alert on throttled time rather than on average utilization, which is the metric that stays reassuringly low.
- Why is exhausting file descriptors particularly dangerous for health-check-based detection?Established connections keep working, so a health check that reuses an open connection, or one performed locally, keeps returning success while every new client connection is refused. The instance stays in rotation and quietly rejects all new traffic. The fix is to monitor new-connection success and descriptor usage as a percentage of the limit, rather than inferring health from traffic that is already connected.
saying these in an interview costs you the question
- Resource exhaustion just crashes the process, so it is the same as a kill
- Exceeding a container CPU limit kills the container
- Health checks will catch a saturated instance
- Disk usage alerts always leave enough time to react
- Average CPU utilization is enough to rule out CPU problems