skip to content

Before you fail a dependency in a resilience test, what must you write down first, and why?

level: juniorimportance: must knowfreq 58%

answer

  1. Decide the verdict before the fault
  2. Not crashing is not passing
  3. Per feature, not per system
  4. Response, screen, side effects, signal, recovery
  5. Assert what must never happen

basics

~20 s

Write the expected degraded behaviour first: for each affected feature, what the caller gets, what the interface shows, which side effects stay allowed, and what must never happen. Without that written oracle, anything short of a crash looks like a pass.

solid answer

~50 s

Write the expected degraded behaviour for every affected feature before the fault is injected, and treat that statement as the test's oracle. Per dependency-and-feature pair, state the response the caller gets, what the interface shows, which side effects remain permitted, what must never happen, what operators are told, and how the feature returns to normal once the dependency recovers. Writing it first forces a product decision — refuse, or serve something reduced — to be made deliberately instead of being inherited from whatever a client library does on a timeout, and it converts straight into assertions you can automate. Written afterwards, it is only a description of what you happened to observe. Assert the absence of forbidden side effects too, not just the status: the dangerous failures are the ones that quietly do more than they should.

code

pseudocode · 12 lines
pseudocode
# expectation row: entitlement service unavailable x collaborative playlist edit
test "playlist edit while the entitlement service is unavailable":
    fail_dependency("entitlement-service", mode = "timeout")

    response = api.rename_playlist(playlist_id = 9142, actor = "collaborator")

    assert response.status == REFUSED
    assert response.code == "ENTITLEMENT_UNAVAILABLE"
    assert playlist(9142).name == "Late Night Drive"        # unchanged
    assert audit_log.contains(actor = "collaborator", outcome = "REFUSED")
    assert events_emitted_for(9142) == []                   # no side effect
    assert degraded_signal("entitlement") == FIRING

go deeper

for a junior

Be ready to say that a resilience test needs an expected result like any other test, and that the expected result is written before the fault. Name a few things it covers: the response, what the user sees, and what the system must not do.

for a middle

Explain how the statement is built per dependency-and-feature pair and how each clause becomes an assertion, including asserting that a forbidden side effect did not happen. Be ready to say why writing it afterwards produces a description rather than a requirement.

for a senior

Show that you drive the decision out of the code and into a reviewed requirement, cover the recovery leg as well as the outage, and can tell a defect in the degraded path from a defect in the fault injection. Expect to be pushed on how you catch a permissive default.

for a principal

Own the question of who decides refuse-versus-reduce across dozens of services, how those decisions are recorded so they survive team turnover, and how you keep the resilience expectations attached to requirements rather than to a drill document nobody reads twice.

## A fault is an action; it is not a verdict Every test has three parts: a setup, an action, and an **oracle** — the thing that decides whether what happened was right. Resilience testing keeps the first two and routinely loses the third. The action is a deliberate failure of a dependency, a node or a link, and failures are loud: logs fill with stack traces, retries stack up, dashboards turn red. In that noise it is easy to accept "the service stayed up" as the verdict. "The service stayed up" is not a specification of anything. It is satisfied equally by a system that refused the write cleanly and by one that silently did the wrong thing. The fix is cheap and almost entirely non-technical: **before the fault is injected, write down what the system is supposed to do while the dependency is unavailable.** That written statement is the oracle. Everything else — the injection mechanism, the duration, the automation — is plumbing around it. ## What the statement has to contain Write it per *feature*, not per system. "The service degrades gracefully" is unfalsifiable; a system has as many degraded behaviours as it has features touching the broken dependency. For each dependency-and-feature pair, state: - **The response the caller gets** — a refusal with a specific code, or a success built from a reduced data set. - **What the interface shows** — an empty state, a value explicitly labelled as stale, a control that is disabled rather than one that fails on click. - **Which side effects remain permitted** — may it still write? still charge? still emit an event downstream? - **What must never happen** — this is the clause people skip, and it is the one that catches the dangerous defects: no partial charge, no widened rights, no silent success reported to the caller. - **What operators are told** — the specific signal that fires, so that "degraded" is observable from outside. - **How it comes back** — does full behaviour return by itself when the dependency recovers, within what window, and does anything queued during the outage need draining or backfilling? ## Why the order matters Three reasons, and all three are about the order rather than the content. First, it is an oracle problem. Written after the run, the statement is a *description of what you observed*, dressed up as a requirement. Human judgement is extremely good at rationalising an outcome once it has been seen. Second, it forces a decision to be made by someone entitled to make it. Whether a feature fails closed (refuse and say so) or serves something reduced is a product and risk decision. Written first, someone answers it. Left unwritten, it is answered by accident — by whatever a client library happens to do on timeout, or by whichever branch a developer wrote at three in the afternoon. Third, a written expectation converts mechanically into assertions, which means the case is repeatable and can sit in the regression pack instead of being re-argued at each drill. ## A worked example A music-streaming service supports collaborative playlists. Edit rights come from a separate entitlement service. The team's 340-case regression pack contained 11 cases for "entitlement service unavailable", and every one of them asserted the same thing: the playlist page still renders and the service returns something that is not a server error. In production the entitlement service timed out for 26 minutes. The playlist service's client fell back to a permissive default, and any signed-in collaborator was treated as the playlist owner. Followers with view-only rights renamed and deleted other people's playlists — a permission escalation that all 11 resilience cases passed with flying colours, because nothing crashed and the page rendered. Had the expectation been written first, that row would have read: *reads succeed from cached membership; every write returns a refusal with code `ENTITLEMENT_UNAVAILABLE`; no rights are ever inferred when membership is unknown; the refusal is audited; the degraded signal fires.* That is four assertions and a signal check, and the fail-open default would have failed on its first run rather than in production. ## Turning it into assertions Build a small table: rows are dependency × feature, columns are the six items above. Each row becomes one test case that asserts, in order, the response, the persisted state (including **the absence** of the forbidden side effect — the record was not written, the event was not emitted, the audit line exists), the operator signal, and the recovery. Asserting absence is the habit that separates a resilience test from a smoke check, because escalation, double-charge and double-send bugs are all bugs of *extra* behaviour, and no status-code assertion will ever see them. ## Common ways it goes wrong - Treating "degraded" as one system-wide mode instead of a per-feature contract. - Stating the expectation in terms of an internal mechanism — a breaker opening, a cache being consulted — rather than observable behaviour. A test coupled to the mechanism passes whenever the mechanism runs, including when the mechanism is doing the wrong thing. - Asserting only the outage and never the return, so recovery is exercised for the first time during a real incident. - Letting whoever runs the fault also decide, on the day, what a pass looks like.

  • Who should own the decision that a feature refuses outright rather than serving a reduced result?
    The product or risk owner, with input from whoever owns security for that path. Engineers propose the options and the cost of each, but anything that grants rights, moves money or records consent should default to refusing, and the decision belongs next to the requirement so a test case can cite it. Left to the code, it is decided by a library default nobody reviewed.
  • The dependency comes back. What do you assert then?
    That full behaviour returns within a stated window without a restart or a manual step; that anything queued or retried during the outage is applied exactly once, not zero or twice; that no value cached during degradation outlives the recovery; and that the degraded-mode signal clears. Recovery assertions are skipped far more often than outage assertions, which is why recovery is usually first exercised during a real incident.
  • How do you stop these expectations from rotting as the system changes?
    Keep them with the functional requirement rather than in a drill document, one row per dependency-and-feature pair, and give each row an identifier the regression case cites. Then a new dependency with no rows, or a row with no case, is visible as a gap rather than discovered during an outage. Review the table whenever a new outbound call is added.

A fire drill that only proves the alarm was loud tells you nothing about whether people reached the right exit.

saying these in an interview costs you the question

  • Treats no crash as a passing resilience test
  • Writes the expected behaviour after seeing the run
  • Describes degraded mode as one system-wide state
  • Asserts the status code but never the forbidden side effect
  • Lets a client library's timeout default define the requirement
  • States the expectation as an internal mechanism, not observable behaviour

context