skip to content

A component suite's later tests hang only when one earlier test runs first; how can a faked clock cause that?

level: seniorimportance: should knowfreq 42%

answer

  1. passes alone, fails in the suite
  2. shared state crossing a boundary
  3. the restore never ran on failure
  4. unmount before restoring
  5. leftover callbacks blame the wrong test

basics

~20 s

The clock is process-wide state. A test that installs it and never restores it leaves later tests with time frozen, so their debounce, poll and timeout waits never fire - and its leftover callbacks can fire into a torn-down tree.

solid answer

~50 s

Faking time replaces shared host entry points for the whole worker, not just for the test that asked. If a test installs a clock and does not restore it - typically because it failed or returned early before its cleanup line - every test after it in that worker runs with time frozen. Anything those tests wait on that a timer gates simply never happens, so they hang until the runner's own limit stops them, and the symptom moves as soon as file order changes. Leftover callbacks are the second half: a timer still queued can fire after a tree is unmounted and write into something that no longer exists. The fix is structural - install and restore in shared setup and teardown so the restore runs on failure too, unmount before restoring, and fail a test that leaves callbacks pending.

go deeper

for a junior

Remember that faking time affects the whole worker, not one test. Whatever a test replaces, it has to give back, and the giving back belongs in a teardown hook.

for a middle

Explain the mechanism behind an order-dependent hang: a frozen clock inherited by the next test means nothing a timer gates ever fires, so waits stop instead of failing.

for a senior

Show the triage: reproduce by running the pair, bisect the preceding files, check install and restore symmetry, unmount before restoring, and assert that no callbacks remain registered.

for a principal

Rule the class out by construction. One shared time-control helper, symmetric hooks, a teardown check for leftovers, and a policy that order-dependent failures are never answered with a retry.

## Why this is an order-dependent failure, not a flaky test The signature is unmistakable once you have seen it: every test passes when its file runs alone, the suite fails when everything runs together, and which test fails depends on the order the runner chose. That is not non-determinism in the product. It is **shared state crossing a test boundary**, and a faked clock is one of the largest pieces of shared state a frontend test can touch, because it replaces host entry points that every other test also depends on. The usual way it escapes is not carelessness about writing the restore, but where the restore was written. A restore at the end of a test body is skipped whenever the body throws, returns early, or is stopped by the runner's timeout - so the first failing test poisons its worker and the suite starts reporting failures that have nothing to do with the change under review. ## Two leaks, not one - **Frozen time inherited by the next test.** Anything a timer gates - a debounce, a poll tick, a deadline, a retry gap, a settle helper that waits with a timeout - never happens, because nobody is advancing the clock any more. The test does not fail on an assertion; it stops. - **Callbacks left pending.** A timer or interval registered before teardown still holds a reference to its closure. Fired later, it writes into an unmounted tree and produces an update-after-teardown error attributed to whichever test was running. - **An interval that was never cleared.** With real time restored it keeps firing for the rest of the run, so its noise appears everywhere and belongs nowhere. - **A pinned current-time reading.** Output derived from the clock silently stays at the frozen instant, which does not hang anything - it makes a later assertion on a relative timestamp wrong in a way that looks like a product bug. - **An unreleased rejection.** Failing work left pending surfaces as an unhandled rejection later and is reported against an innocent test. | Leak | Symptom in a later test | Guard | |---|---|---| | clock left installed | waits never complete; the test times out | install and restore in shared hooks | | timer left pending | update after teardown, wrong test blamed | unmount first, then restore, then assert nothing is pending | | interval never cleared | repeating noise across the whole run | fail the test if callbacks remain registered | | pinned time reading | a relative-time assertion is quietly wrong | restore the reading with the clock, in the same hook | ## Triage that finds it quickly 1. **Reproduce with order.** Run the failing test alone; it passes. Run it after the suspect file; it fails. You now have a deterministic reproduction, which is the whole game. 2. **Bisect by pairs.** Run halves of the suite before the victim until one file reproduces it. Runners that can print or fix the order make this a few minutes' work. 3. **Look for asymmetry.** In the suspect file, check whether every install of shared state has a matching restore in a hook rather than in a body. 4. **Disable the fake.** Remove the clock from the suspect file; if the victim passes, the diagnosis is settled. 5. **Assert on leftovers.** Add a teardown assertion that no callbacks remain registered, and you convert every future instance of this class from a mystery hang into a failure at the source. ## Teardown that actually runs - Put the install in shared setup and the restore in shared teardown, so the restore executes on the failure path as well as the success path. - **Unmount the tree before restoring the clock.** Unmounting usually cancels the component's own timers; doing it in the other order can leave them registered against a restored real clock. - Make the pairing structural rather than remembered: one shared helper that installs, hands the test a way to advance, and restores - so an individual test cannot forget half of it. - Give the runner per-file isolation where it is available, but do not rely on it. Isolating module state does not undo a replacement of host entry points performed inside the same worker. ## What this is not It is not a reason to raise the victim's timeout: the wait is not slow, it is never going to finish. It is not a candidate for a CI retry either - the failure is deterministic given an order, so a retry that happens to run the file alone hides a real leak and lets it spread to more tests. And it is not the victim's bug, which is why fixing the victim is the most common wasted afternoon in this area. Fix the test that took shared state and did not give it back.

  • Which leftovers besides the clock itself cause the same order-dependent failures?
    Any shared handle a test replaces or registers: an interval still scheduled, a module-level cache left warm, a subscription never disposed, a queued callback holding a reference to an unmounted tree, a replaced network or storage entry point. Symmetric install and restore in shared hooks removes the whole class.
  • Is retrying the hanging test in CI an acceptable response here?
    No. The failure is deterministic given an order, so a retry that reruns the file in isolation hides a real leak and lets it spread to more tests. Reproduce it by running the pair, fix the teardown, and keep retries for genuinely non-deterministic infrastructure.
  • Why unmount the component before restoring the clock rather than after?
    Because unmounting is what cancels the component's own timers. Restore first and those cancellations run against the real clock while callbacks already registered on the fake are stranded - or the reverse, leaving live timers pointed at a tree that is about to disappear. Unmount, then restore, then check nothing is pending.

saying these in an interview costs you the question

  • Raises the hanging test's timeout instead of finding the frozen clock
  • Restores the clock at the end of the test body, so a failure skips it
  • Assumes a faked clock is scoped to the test that installed it
  • Marks the order-dependent test as flaky and retries it in CI
  • Treats callbacks left pending at teardown as harmless noise
  • Debugs the victim test rather than the test that took shared state