skip to content

What experiment during a ticket on-sale would prove a resilience claim - one failing dependency, the rest of the system unaffected - is real?

level: seniorimportance: should knowfreq 48%

answer

  1. under load, never on an idle system
  2. two paths, one of them untouched
  3. slow is harder than dead
  4. the control path is the evidence
  5. recovery without a human, and timed

basics

~20 s

Inject a slow dependency under the on-sale load and compare two request paths: one that never touches it and one that always does. Containment means the untouched path keeps its response-time bound and the affected path still answers inside one.

solid answer

~50 s

Run it under load, because containment failures are capacity failures and they do not appear on an idle system. Pick two request paths - a control that never touches the chosen dependency and an affected path that always does - and record a high percentile for each before, during and after the injection. Make the injected failure a slow answer rather than a dead one: a refusal returns immediately and holds nothing, while a delay is what fills queues and ties up callers. The claim passes if the control path's bound does not move, if the affected path still answers inside a bound with a degraded or explicit failure answer rather than hanging, and if recovery is driven by something that is not the failed component, with the time it took recorded. Configuration that asserts resilience is not evidence of it.

code

pseudocode · 19 lines
pseudocode
experiment "containment of the pricing dependency":
    load = steady 12000 requests per second      // the on-sale rate, not idle

    control  = requests that never touch pricing
    affected = requests that always touch pricing

    baseline:  record p99(control), p99(affected), outcomes for 5 minutes

    inject:    pricing answers after 30s instead of 50ms   // slow, not dead
    during:    record the same series for 10 minutes

    remove:    pricing answers normally again
    after:     record until both paths return to baseline; note the elapsed time

    pass when:
        p99(control) during  <= p99(control) baseline bound
        p99(affected) during <= its degraded bound, with an explicit outcome
        offered load during  is not greater than baseline
        recovery happened with no manual step

go deeper

for a junior

Understand that resilience is not the absence of failures. It is a claim about what happens when one occurs, and claims of that kind are checked by causing the failure on purpose.

for a middle

Describe the shape of the experiment: real load, an injected failure, and a comparison between requests that depend on the broken component and requests that do not.

for a senior

Run it properly - slowness rather than death, a control path, three measurement windows, recovery timed and unassisted - and read a rise in the control path as evidence of a shared resource nobody drew on the diagram.

for a principal

Decide how often this runs and who owns the result, so containment is a standing measurement rather than a story told after an incident, and make the control-path bound part of what a service commits to.

## What the resilience claim actually says Resilience is one of the two means that keep responsiveness true, and its claim is narrower and more checkable than "the system is robust". It says three things: 1. A failure is **contained** inside the component where it happened, so the response-time bound for requests that do not depend on that component does not move. 2. The **affected** path still answers within a bound - degraded, or with an explicit failure - rather than holding the caller until the caller's own deadline expires. Responsiveness under failure is part of the property, not a separate concern. 3. **Recovery is delegated** to something that is not itself failing, so the system returns to its bound without a person in the loop, and the time that took is a number somebody can quote. Notice what is not in the claim: that nothing fails. A year without an incident is an absence of events, not a property. ## Designing the experiment 1. **Hold real load.** Run at, or near, the on-sale arrival rate the system claims to serve. Containment problems are usually contention problems - a shared pool, a shared queue, a shared connection limit - and none of them is visible when there is spare capacity everywhere. 2. **Choose the dependency and the two paths.** Name the component to be broken, then pick an **affected path** that always touches it and a **control path** that never does. Without the control, the experiment cannot distinguish containment from luck. 3. **Inject slowness, not death.** A component that refuses connections instantly teaches callers immediately and holds nothing. A component that answers after thirty seconds instead of fifty milliseconds is the case that accumulates waiting work and leaks into unrelated paths. 4. **Measure through three windows.** Record a high percentile and the outcome mix for both paths before the injection, during it, and after it is removed. 5. **Remove the injection and keep watching.** Recovery is part of the claim, and systems that recover only after a restart, or only after a queue is drained by hand, have not shown the property. ## Reading the result | Observation | What it means | |---|---| | Control path's percentile stays inside its bound | the failure was contained: this is the core evidence | | Control path's percentile rises with the injection | something is shared between the paths - the failure is not contained, whatever the design says | | Affected path answers inside a bound, degraded or explicitly failed | responsiveness under failure holds on the path that could not succeed | | Affected path hangs until callers' deadlines expire | the failure is being transmitted to callers rather than handled | | Total offered load rises during the injection | the system is answering a failure with more traffic than it was already carrying | | Bound returns after removal, with no manual step | recovery is delegated, and the time to recover is measurable | ## Failures that prove nothing - **A configuration review.** Settings describe an intention. Until the condition they cover has occurred under load, they are untested code paths. - **An experiment with no load.** Everything is contained when nothing is contended. - **A hard stop only.** The fast-refusal case is the easy one; teams that only ever test it are surprised by the first slow dependency they meet in production. - **Watching only the broken thing.** The failing dependency's own error rate is the least interesting series on the screen. The interesting one is the path that has nothing to do with it. - **A single run.** Containment that depends on which requests happened to be in flight is not containment; repeat the injection and look at the spread across runs. ## What recovery evidence looks like The property says recovery is **delegated**, which means something outside the failed component notices and acts: a supervising part of the system replaces or restarts it, or traffic is directed to an equivalent instance while it is unhealthy. Two things are worth writing down at the end of the experiment - how long the system took to return to its bound after the dependency recovered, and whether it did so on its own. An operator typing a command is a valid operational procedure and an invalid answer to this question. ## Why interviewers ask it this way Asking "what is resilience" gets a definition. Asking for the experiment gets the difference between a candidate who has run one and a candidate who has read about one. The tell is the control path: people who have actually done this reach for it immediately, because they have been fooled once by a system that stayed healthy during an injection for reasons having nothing to do with containment.

  • Why is a slow dependency a harder test than one that is completely down?
    A refusal or an unreachable address returns immediately, so callers learn at once and hold nothing. A slow answer keeps callers committed for the whole delay, which fills queues, consumes shared capacity and is the route by which a local failure reaches paths that never touched it.
  • The control path's percentile rose during the injection. What does that tell you?
    That something is shared between the two paths - a pool, a queue, a connection limit, a downstream store - and the failure is not contained however the architecture diagram is drawn. The control path is what turns that from an argument into a measurement.

saying these in an interview costs you the question

  • Offering configuration settings as evidence of resilience
  • Injecting only a hard outage, never a slow dependency
  • Running the experiment with no load applied
  • Measuring only the failing path and no control path
  • Calling it resilient when recovery needed a person