skip to content

A frontend suite simulates only a generic 500 for every endpoint it stubs. As the lead, how would you decide which additional failure and latency conditions are worth encoding as tests, and where each belongs?

level: principalimportance: should knowfreq 32%

answer

  1. one test per rendered outcome
  2. enumerate branches, not status codes
  3. cheapest level that exercises it
  4. latency needs a policy, not a spread
  5. observability closes the rest

basics

~20 s

Simulate a failure when the product renders a distinct response to it. Enumerate the branches the code actually takes, cover each once at the cheapest level that exercises it, and leave broad fault-injection matrices out of the functional suite.

solid answer

~50 s

The selection rule I use is: **one test per distinct user-visible outcome, not one per possible failure**. Start from the branches the client code really has — an auth failure that redirects, a validation failure that surfaces field errors, a not-found that shows an empty or missing state, a retryable server failure, a transport failure with no response at all, and a slow response that must show a pending state. Each of those earns one test, at the cheapest level that genuinely exercises it: most belong in component tests with an intercepted network, because they are pure rendering decisions. Reserve end-to-end fault injection for the failures that cross boundaries the component test fakes away — session expiry mid-flow, a failure during a real navigation or upload. Everything beyond that is a combinatorial matrix that costs maintenance and finds little, and genuinely broad fault behaviour is better learned from production observability than from more stubs.

go deeper

for a junior

Focus on the idea that a failure is worth simulating when the screen shows something different because of it, and that one test per distinct visible outcome is the unit to aim for.

for a middle

Be ready to enumerate the branches your client code really has — auth, validation, not-found, transient, transport, slow — and to place each test at the level that exercises it without faking away the subject.

for a senior

Argue the cost side: browser-level failure tests are far more expensive to keep alive, failure fixtures drift unless they share a source with success fixtures, and suite-wide retry settings quietly turn every failure test into a slow one.

for a principal

Own the policy and its feedback loop — which failure classes every screen must handle, where fault injection lives, what the default network posture of the suite is, and how production error telemetry decides what gets covered next.

## Start from branches, not from a catalogue of failures The tempting move is to enumerate every status code an API can emit and stub each one. That produces a suite that grows with the API surface rather than with the product's behaviour, and most of the tests assert the same generic error banner. The productive inversion is to read the client and the designs and list the **distinct outcomes** the user can end up in. Typically that is a short list: - authentication lost — the app redirects or prompts to sign in again; - authorization refused — a "you don't have access" state, which is not the same screen; - input rejected — field-level messages next to the offending inputs; - resource absent — an empty or not-found state, often deliberately friendlier than an error; - transient server failure — a retryable error banner, possibly after automatic retries; - no response at all — the transport failed, so the client's thrown-error branch must render something; - slow response — the pending affordance must appear and then clear. If two failures land on the same rendered outcome through the same code path, they are one test, and picking either status is enough. If a failure has no branch at all, that is a finding about the product, not a missing test. ## Choose the cheapest level that exercises the branch Most of that list is a rendering decision made from a value the client already computed, so a component test with the network intercepted is the right home: fast, deterministic, and it puts the assertion next to the UI it protects. Move a case up a level only when the component test would have to fake away the very thing being tested — a session that expires between two real page navigations, an upload interrupted after bytes were sent, a failure that must survive a full reload with persisted state. Those genuinely need the wider harness, and they are the minority. Be explicit that a browser-level test is roughly an order of magnitude more expensive to run and to keep alive than a component test, so each one you add should buy a failure mode nothing cheaper can reach. A duplicated 500 test at three levels is pure cost. ## Latency deserves a policy, not a matrix For delay, two questions decide it. **Does a pending affordance exist?** If a screen has a skeleton, a disabled button or an optimistic update, that state needs one deterministic test with a controlled response. **Does something time out?** If the client bounds a request and shows a distinct timed-out state, test that boundary with a virtual clock rather than a real wait. What does *not* belong in the functional suite is a spread of injected delays intended to characterize behaviour under a slow network — that is a different discipline with different tooling and different owners, and encoding it as stubbed delays makes the suite slow without making it more informative. ## Keep the cost visible Three operational rules keep this from decaying: 1. **Failure fixtures share a source with success fixtures.** If an error envelope is hand-written per test, every API change leaves a dozen stubs describing a response the server no longer sends — tests that pass while the app breaks. Deriving them from the same schema or factory as the happy-path fixtures is what stops that drift. 2. **A default network posture for the suite.** Unhandled requests should fail loudly rather than silently hitting a real host, and automatic retries should be off unless a test is about retrying — otherwise every deliberate failure quietly costs seconds of backoff. 3. **Delete tests when the branch goes.** A failure test whose branch was removed is a fossil that still costs review time on every change. ## Where the remaining confidence comes from Accept that a stub proves how the UI reacts to a *described* failure, never that the failure is described correctly. That gap is closed elsewhere: by contract checks that keep fixtures honest against the real API, and by production telemetry showing which failures actually occur and at what rate. I would rather have six well-chosen failure tests plus real error monitoring than forty stubs asserting the same banner — and I'd let observed production failure rates feed back into which cases the suite covers next, so the test list tracks reality instead of imagination.

  • Two endpoints answer 403 and 404 and both render the same generic banner. Do you write two tests?
    No — one is enough while the outcome and the code path are shared, and I'd note the duplication as a possible product question rather than a test gap. If a designer later gives "no access" its own screen, the branch has split and a second test earns its place. Tests should track distinct outcomes, not the size of the API's status vocabulary.
  • How do you stop failure fixtures from drifting away from what the API really returns?
    Derive them from the same source as the success fixtures — the schema or a shared factory — so an error envelope change breaks the stubs rather than silently invalidating them. Stubs are assertions about the server's contract, and without a shared source they decay into folklore that keeps tests green while the application breaks.
  • What would make you add an end-to-end failure test rather than a component-level one?
    When the component test would have to fake away exactly what is under test: a session expiring between real navigations, an interrupted upload, a failure that has to survive a reload with persisted state. Those cross boundaries the component harness replaces with a stub. Everything that is a pure rendering decision stays at the cheap level.

saying these in an interview costs you the question

  • Wants a test for every status code the API can return.
  • Duplicates the same 500 test at component and end-to-end level.
  • Injects a spread of random delays to characterize slow networks.
  • Hand-writes error envelopes separately in every test file.
  • Treats passing failure stubs as proof the real failures are handled.

context