skip to content

A data client retries a failed request three times with exponential backoff, and the test that asserts the final error UI takes seven seconds and fails intermittently. How do you test the retry and backoff behaviour without waiting real time?

level: seniorimportance: should knowfreq 45%

answer

  1. virtual clock instead of elapsed time
  2. scripted sequence, not one fixed stub
  3. count the attempts, don't time them
  4. not everything deserves a retry
  5. make the schedule a parameter

basics

~20 s

Replace real waiting with a virtual clock the test advances, and make the mock answer a scripted sequence — fail, fail, succeed. Assert the number of attempts and the UI at each stage, and shrink the backoff schedule through configuration rather than sleeping.

solid answer

~50 s

Two things have to come under the test's control: the schedule and the responses. For the schedule, install fake timers so the backoff delays are advanced explicitly instead of elapsing — the test then costs milliseconds and no longer depends on machine speed. For the responses, script the handler as a sequence keyed off a counter: the first two calls fail, the third succeeds, so a single test can assert both the retrying and the recovery. Then assert on observable facts — the request count the handler recorded, and what is rendered after each advance: spinner during the wait, list after the successful attempt, error UI once the attempts are exhausted. Two more cases matter as much as the happy retry: that a 4xx which will never succeed is *not* retried, and that a non-idempotent request is not blindly replayed. And keep the backoff parameters injectable so a test can pass a tiny schedule instead of your production one.

code

javascript · 27 lines
javascript
import { render, screen } from '@testing-library/react'
import { http, HttpResponse } from 'msw'
import { server } from './test-server'
import { Report } from './Report'

test('retries twice, then renders the data', async () => {
  vi.useFakeTimers()
  let attempts = 0

  server.use(
    http.get('/api/report', () => {
      attempts += 1
      return attempts < 3
        ? HttpResponse.json({ message: 'upstream down' }, { status: 503 })
        : HttpResponse.json({ rows: [{ id: 1 }] })
    })
  )

  render(<Report />)

  await vi.advanceTimersByTimeAsync(1000)  // first backoff
  await vi.advanceTimersByTimeAsync(2000)  // second backoff

  expect(attempts).toBe(3)
  vi.useRealTimers()
  expect(await screen.findByRole('table')).toBeInTheDocument()
})

go deeper

for a junior

Know that a test should never sit through real backoff delays, and that a mock can answer differently on each call so a retry has something new to receive.

for a middle

Explain the two levers — a virtual clock the test advances and a counter-driven handler that scripts the response sequence — and assert the attempt count rather than the elapsed time.

for a senior

Demonstrate production judgment: prove that permanent client errors are not retried, that unsafe writes are not blindly replayed, and that retry configuration is off by default across the suite so it does not poison unrelated tests.

for a principal

Own retry policy as a design concern — where backoff parameters live, which failure classes qualify, how idempotency is guaranteed for writes — so that the tests express a policy every team shares rather than each client's private habits.

## What the test is actually about Retry logic has three separable behaviours, and a good test pins each: **how many attempts** are made, **which failures qualify** for a retry, and **what the user sees** in between and at the end. The seven-second runtime in the question is a symptom of testing all of that through the real clock, which also imports every source of timing flakiness the CI machine has. ## Control the clock Install the runner's fake timers before rendering, and advance them by the amount the schedule expects. Real elapsed time drops to nothing and, more importantly, the test becomes an assertion about the schedule rather than an accident of it: advancing 999ms when the first backoff is 1000ms should leave the second attempt un-made, and advancing past it should trigger exactly one more. That is a far sharper test than "eventually it worked". Two hazards come with a frozen clock. Retrying queries that poll for an element can stall, because their own polling is timer-driven — use the async timer-advancing helper (`vi.advanceTimersByTimeAsync` / `jest.advanceTimersByTimeAsync`) so pending promise jobs get a chance to run between ticks, and configure the retrying wait to cooperate with fake timers. And simulated user input schedules its own timers between keystrokes, so a user-event instance must be told how to advance the fake clock (`userEvent.setup({ advanceTimers: vi.advanceTimersByTime })`) or interactions appear to hang. Always restore real timers afterwards. ## Script the responses A single fixed stub cannot express "fails, then succeeds". Give the handler a counter in test scope: ```js let attempts = 0 server.use(http.get('/api/report', () => { attempts += 1 return attempts < 3 ? HttpResponse.json({ message: 'upstream down' }, { status: 503 }) : HttpResponse.json({ rows: [] }) })) ``` That counter is also your assertion target: `expect(attempts).toBe(3)` proves the retry ran the right number of times, and — in the negative tests — that it did not run at all. Counting attempts is far more robust than trying to infer retries from timing. ## The cases worth writing - **Recovery.** Fails twice, succeeds on the third attempt: the user ends up seeing data and never sees an error. This is the case that justifies having retries at all. - **Exhaustion.** Fails every time: after the last attempt the error UI appears, the spinner is gone, and the attempt count equals the configured maximum plus the original call. Getting the off-by-one right in the assertion is part of the test's value. - **No retry on a permanent failure.** A 400 or a 404 will not become a 200 by being asked again, and retrying a 401 can lock an account. Assert the count is 1 and the error surfaced immediately. - **No blind replay of unsafe writes.** A retried POST can create a duplicate. If the client retries writes at all, the test should show the safeguard — typically an idempotency key held constant across attempts — rather than asserting that three identical creates went out. - **Backoff growth.** If the schedule is meant to be exponential, advance to just before each boundary and assert no new attempt, then past it and assert exactly one. Avoid asserting wall-clock durations. ## Make the schedule injectable The deepest fix is design, not test tooling: production backoff constants should be parameters of the client, not literals buried in it. A test that constructs the client with `{ retries: 2, baseDelayMs: 10 }` runs fast even without fake timers and reads clearly. Fake timers then remain available for the cases where you genuinely want to assert the schedule shape. Where the retry lives in a data-fetching library, that library's own configuration usually offers the same lever, and the test should turn retries **off** for every unrelated test — otherwise a deliberately failing stub in some other test file silently costs three attempts and several seconds of backoff. ## What not to do Do not sleep for the real backoff. Do not assert on internal retry counters reached through the client's private state. And do not leave production retry settings enabled globally across the suite: it is one of the most common causes of a slow, mysteriously flaky frontend test run, because every stubbed failure anywhere becomes a multi-second timing event.

  • Which failures should your test prove are *not* retried?
    Ones that cannot succeed by repetition or are unsafe to repeat: client errors such as 400, 404 and 422, authentication failures where repeated attempts can lock an account, and any request already rejected on validation. Assert the attempt count is exactly one and that the error surfaced immediately, since a client that retries these wastes time and can do real damage.
  • Why is asserting a recorded attempt count better than inferring retries from how long the test took?
    Because the count is a direct observation of the behaviour, while duration is a proxy that varies with the machine. Timing assertions are the flakiness you were trying to remove: they pass on a fast runner and fail on a loaded one. A counter incremented inside the handler is exact, fast to assert, and reads clearly in a failure message.
  • What happens to an unrelated test suite if production retry settings are left on globally?
    Every deliberately failing stub anywhere becomes a multi-second timing event, and any test asserting an error state waits out the whole backoff schedule before the UI appears. It is a leading cause of slow, mysteriously flaky frontend suites. Disable retries in the shared test configuration and enable them only in the tests whose subject is retrying.

saying these in an interview costs you the question

  • Sleeps for the real backoff duration in the test.
  • Uses one fixed stub, so the retry can never succeed.
  • Retries 4xx responses and calls it resilience.
  • Asserts elapsed milliseconds instead of the attempt count.
  • Leaves production retry settings enabled for the whole suite.

context