skip to content

A colleague tests asynchronous code by starting the work, sleeping for 100 milliseconds, and then asserting the result. Why is that approach problematic, and what should the test wait on instead?

level: juniorimportance: must knowfreq 70%

answer

  1. sleep = guess; fails both ways
  2. wait on a signal, not a duration
  3. quiescence = empty queue AND no active worker
  4. generous timeout, never a tuning knob
  5. healthy test never reaches the timeout

basics

~20 s

A fixed sleep is a guess about duration. Too short and the test fails on a loaded machine; too long and the suite crawls. Wait on a real completion signal instead — a future, a latch, a callback, or a quiescence check — with a generous timeout.

solid answer

~50 s

Sleeping hard-codes a guess, so the test fails in two directions at once. On a slow or contended CI runner the assertion runs before the work finishes and the test fails for no real reason; if you pad the sleep to be safe, every run pays that cost and a suite of hundreds loses minutes. The failure message is useless too: it says "not finished in 100 ms", not what went wrong. Wait on something the system actually signals: the future/promise the operation returns, a latch or semaphore released by the completion callback, an item published to a queue the test reads, or **quiescence** — drain until the work queue is empty and no worker is active. Give that wait a generous timeout (seconds), because a correct test returns immediately and only a broken one ever waits the full timeout. Best of all, remove waiting entirely by injecting a virtual clock or a scheduler the test itself drives.

code

text · 15 lines
text
BAD:
  submit(work)
  sleep(100ms)
  assert store.contains("x")

BETTER (completion handle):
  future = submit(work)
  future.await(timeout = 5s)      // returns as soon as work ends
  assert store.contains("x")

BETTER (no handle available):
  submit(work)
  scheduler.awaitQuiescent(timeout = 5s)
     // blocks until: queue.isEmpty() AND activeTasks == 0
  assert store.contains("x")

go deeper

for a junior

State the two-sided failure clearly (flaky when slow, wasteful when padded) and name at least one real thing to wait on: the returned future, a latch released by the callback, or a poll-until-condition helper.

for a middle

Add quiescence as the answer for fire-and-forget work, define it as empty queue plus no active task, and explain why the timeout should be generous rather than tuned.

for a senior

Frame it as removing wall-clock time from the test's correctness condition; mention virtual clocks and injected schedulers as the way to make the wait disappear entirely, and mention the diagnostic value of failing with state dumps.

for a principal

Talk about it as a suite-wide invariant — no raw sleeps in tests, enforced by lint — and about the cost curve: sleeps convert into CI minutes and, via flaky reruns, into lost trust in the suite.

## Why a fixed sleep is a bad wait `startWork(); sleep(100); assert(result)` bets that the work always finishes inside 100 ms on every machine that will ever run it. The bet fails in both directions. **Too short.** CI runners are shared, cold and CPU-throttled; a pause, a noisy neighbour or a container quota can stretch a 5 ms operation into 300 ms. The assertion then runs on an unfinished system and reports a failure that has nothing to do with the code. This is the classic recipe for a flaky test, and flaky tests get re-run, then muted, then deleted — so the sleep eventually destroys the coverage it was meant to provide. **Too long.** The obvious defence is to raise the sleep. Now every run pays the worst case even though the work is normally instantaneous. Three hundred such tests at 500 ms each is two and a half minutes of pure waiting per run, on every branch, forever. Sleep-based waits are the main reason async suites are slow. **Bad diagnostics.** A sleep-based failure says only "the expected state was not reached in time". It cannot distinguish "still running", "deadlocked", "threw and swallowed the exception", or "never started". ## What to wait on instead Wait on a *fact*, not on a *duration*: 1. **A completion handle.** If the operation returns a future/promise, block on it (with a timeout) or register a continuation the test observes. This is exact: the wait ends the instant the work ends. 2. **A synchronisation object the production callback touches.** A countdown latch, a semaphore, or a bounded queue the callback publishes to. The test blocks on it; the callback releases it. Zero polling, zero guessing. 3. **Quiescence (the drain check).** When the API gives you no handle — fire-and-forget work submitted to an executor — expose a way to ask the runtime "is everything settled?": pending queue empty **and** no task currently executing. Test helpers named `awaitIdle`, `drain`, or `runUntilQuiescent` do this. The composite condition matters: an empty queue alone is not idle, because a task may be mid-flight and about to enqueue more. 4. **Condition polling with a deadline.** As a fallback, poll the observable state every few milliseconds until it holds or a deadline passes. This is still nondeterministic in principle, but it converts "always wait X" into "wait only as long as needed, fail after a long ceiling". ## How to choose the timeout Set the timeout to the largest value that still keeps a genuine hang from stalling the build — typically 1–10 seconds. It is not a tuning knob for making the test pass: in a healthy run the wait returns in microseconds, so a big timeout costs nothing. If you find yourself increasing timeouts to fix failures, the test is timing-dependent and needs a real signal or a virtual clock, not a bigger number. ## When a sleep is still acceptable Rarely, and never as the sole synchronisation. Legitimate uses: asserting that something has *not* happened yet (a negative check has no event to wait for — but bound it and treat it as inherently weak), or deliberately exercising a timing path in a stress run. Even then, prefer advancing a virtual clock over sleeping on the real one. ## The deeper principle Deterministic async testing means every wait in a test corresponds to an event the system emits, not to wall-clock time. Once every wait is signal-based, test duration becomes proportional to actual work and failures point at causes instead of at timers.

  • If you must poll for a condition instead of waiting on a signal, what makes a poll loop safe?
    Poll a cheap, side-effect-free predicate on a short interval (a few milliseconds) against a long absolute deadline, and on expiry fail with a dump of the observed state rather than a bare timeout message. The deadline must be absolute, computed once, so retries cannot extend it indefinitely. Polling is still weaker than a signal because it can miss transient states, so use it only when the system exposes no completion event.
  • Is there any case where a sleep in a test is legitimate?
    Mainly for negative assertions — proving something has not happened yet, such as a debounced call not firing early — because there is no event to wait for. Even then it is a weak, slow check and a virtual clock is better: advance time by the debounce interval minus one tick and assert nothing fired. Sleeping to 'let things settle' before an assertion is never legitimate; that is a hidden race.

Waiting for a delivery by standing outside for exactly ten minutes versus waiting for the doorbell: the doorbell is exact and free, the ten minutes is both too long and sometimes too short.

saying these in an interview costs you the question

  • "Just increase the sleep until it passes" — treats the timeout as a tuning knob and hides a real race.
  • Thinking a long timeout slows the suite: a signal-based wait with a 10 s timeout costs nothing in a healthy run.
  • Claiming the empty work queue alone proves the system is idle, ignoring the task currently executing (which may enqueue more).
  • Adding a retry annotation instead of removing the timing dependency.
  • Treating sleep-based passes as proof of correctness — the test only proved 'finished within a guess'.

context