skip to content

A unit test for a subscription renewal job passes by day but fails on the nightly run. How do you diagnose it?

level: seniorimportance: should knowfreq 57%

answer

  1. The verdict changes, the code did not
  2. Find the input nobody wrote down
  3. Pin the instant and reproduce locally
  4. Month ends, shifts, offsets, midnight
  5. Sequence asserted, never promised

basics

~20 s

Suspect a hidden ambient input. A test that passes at one time of day and fails at another is reading the wall clock or relying on an unspecified ordering. The fix is to pass the instant in and make the expected order explicit, not to loosen the assertion.

solid answer

~50 s

The verdict changing with the hour means the test has an input nobody wrote down. Two causes account for most of them. The first is the wall clock: the code or the fixture computes a date from 'now', so the case straddles a midnight, a month end, a daylight-saving shift or a time-zone offset only during the window the nightly run occupies. The second is an ordering assumption — the test asserts on a sequence that the production code never promised, and the iteration order happens to differ when the data or the run differs. Diagnose by pinning the clock to the failing instant and re-running locally; if it reproduces, the clock is the input. The repair is to make the input explicit: pass the instant as a parameter and assert order-insensitively. Adding a delay or a retry hides the defect and makes the level untrustworthy.

code

pseudocode · 11 lines
pseudocode
// ambient: the case under test changes with the hour
function isDue(subscription):
    return subscription.termEnd <= Clock.now()

// explicit: the test states the instant it means
function isDue(subscription, at):
    return subscription.termEnd <= at

test renewal_is_due_at_the_first_instant_after_the_term_ends:
    subscription = Subscription(termEnd = INSTANT_2026_03_31T23_59_59Z)
    assert isDue(subscription, at = INSTANT_2026_04_01T00_00_00Z)

go deeper

for a junior

Recall the rule that a test must not read the current time or depend on which order a collection happens to produce. If you need a date, write the exact one you mean into the test.

for a middle

Explain the mechanics: which date arithmetic breaks at month ends, daylight-saving shifts and zone offsets, and how passing the instant in turns those into cases you can write deliberately. Know why widening an assertion is not a fix.

for a senior

Show the diagnosis path end to end — read the failing run's instant, reproduce with a pinned clock, decide whether the code or the expectation was wrong, then remove the ambient input and add the boundary cases the old design could not reach.

for a principal

Own the policy angle: make ambient clock, locale and zone access a reviewable design smell, decide where a time source is injected across services, and resist the organisational habit of answering instability with retries, which quietly converts a fast suite into an unreliable one.

## The symptom is the diagnosis A test whose verdict depends on when it runs has an input that is not in the test. Nothing else can produce that pattern. So the diagnostic question is not 'why is this test unstable' but 'which unwritten input does it read', and there are only a few candidates at this level: the wall clock, an iteration order nobody promised, a locale or time-zone setting inherited from the machine, and leftover state. On a renewal job scheduled inside a six-hour nightly window, the first two dominate. ## Cause one: the clock as an ambient parameter Renewal logic is dense with date arithmetic — term end, grace period, proration, the 'due today' predicate. If either the production code or the fixture calls a global 'now', the case being tested silently changes every run. The failures cluster in recognisable places: - **midnight straddle** — the fixture computes 'today' at one instant and the code computes it milliseconds later, on the other side of a date boundary; a job starting at 23:58 and finishing after midnight reproduces this reliably and a 10:00 developer run never does; - **month and year ends** — 'one month after 31 January' has no obvious answer, and whatever the code picks, a test built on a relative date only exercises that branch on certain days; - **daylight-saving shifts** — a nominal day is 23 or 25 hours twice a year, so a test asserting a 24-hour interval is wrong twice a year and right the rest; - **time-zone offset** — the runner's zone differs from the developer machine's, and a date derived from an instant lands on a different calendar day. The confirming experiment is cheap: pin the clock to the instant printed in the failing run's log and execute the test locally. If it reproduces deterministically, you are done diagnosing — and you now have a permanent regression case. ## Cause two: the ordering assumption The second flavour is a test that asserts on a sequence the code never promised. The renewal job collects due subscriptions, processes them, and the test asserts on the resulting list positionally: ```pseudocode assert result[0].id == 4181 assert result[1].id == 4192 ``` Nothing in the production contract says which order those come back in; the arrangement happened to be stable in the developer's fixture and differs when the set is built differently or iterated by a hash-derived order. This is a false assertion, not an unlucky one: it can be red while the behaviour is entirely correct, and — worse — green while the behaviour is wrong, because it is checking a property nobody maintains. Assert on the set, or sort by an explicit key first, or, if order is genuinely part of the promise, make the production code state it and then assert on it deliberately. ## Repairing it properly The governing principle is that a test at this level may have no ambient inputs. Concretely: ```pseudocode // before: reads a global clock, so the case changes hourly function isDue(subscription): return subscription.termEnd <= Clock.now() // after: the instant is an argument, so the test states the case function isDue(subscription, at): return subscription.termEnd <= at test renewal_is_due_at_the_first_instant_after_the_term_ends: subscription = Subscription(termEnd = INSTANT_2026_03_31T23_59_59Z) assert isDue(subscription, at = INSTANT_2026_04_01T00_00_00Z) ``` Where passing an instant into every call is impractical, the equivalent move is to give the behaviour a time source it receives rather than reaches for, and hand it a fixed value in the test. Either way, the calendar becomes something the test *chooses*, which turns the previously untestable cases — month end, the shift day, the second either side of the boundary — into ordinary named cases you can write on purpose. Locale and zone deserve the same treatment: state them explicitly rather than inheriting the machine's. ## What not to do Three responses look like fixes and are not. **Widening the assertion** to a range that swallows the discrepancy destroys the oracle — the check that decides whether the behaviour is wrong — and the test now passes for the wrong reason. **Adding a delay or a retry** buys nothing at this level, because there is no asynchrony here to wait for; it converts a five-millisecond test into a slow one and leaves the ambient input in place. **Skipping the test on the nightly schedule** deletes coverage exactly where the failure was found. One more diagnostic caveat: before concluding the code is at fault, check whether the code is right and the *test's* expectation is the thing that assumed a fixed calendar. Time-dependent failures are as often over-specified assertions as they are real defects, and the distinction decides whether you ship a fix or a test change.

  • Once the instant is an argument, which extra cases would you deliberately add?
    The ones the ambient clock made unreachable: the second either side of a term boundary, the last day of a 31-day month rolling into a 30-day one, 29 February, a nominal day that is 23 or 25 hours long because of a shift, and a calendar date derived under two different offsets. Each is a named case now, and together they cover the branches that previously only ran on certain days.
  • The test asserts on the first two elements of a returned list. When is that legitimate?
    Only when ordering is part of the behaviour's stated promise — a sorted result, a queue, a paginated sequence with a documented key. If it is, assert on it explicitly and make the production contract say so. If it is not, positional assertions check a property nobody maintains: they can go red on correct behaviour and green on broken behaviour. Compare as a set, or sort by an explicit key first.
  • Why is adding a short delay to the test the wrong response here?
    Because nothing is being waited on. The failure comes from an unspecified input, not from work still in flight, so a delay leaves the cause untouched while making a millisecond test slow. Multiplied across a suite it erodes the speed that gives the level its value, and it teaches the team that instability is handled by tolerance rather than by removing the ambient dependency.

saying these in an interview costs you the question

  • It is just flaky; re-run the build
  • Add a small delay so the timing settles
  • Widen the assertion until it stops failing
  • Disable the test on the nightly schedule
  • Reading the current clock inside logic is unavoidable
  • Iteration order is stable in practice, so asserting positions is fine

context