Why is sleeping in a test a poor way to prove a job groups records by when they happened?
answer
- which clock the test depends on
- wall clock versus the record's own moment
- hand records over one at a time
- the test advances the claim itself
- no sleep, same result every run
basics
~20 sSleeping proves only that the machine's wall clock advanced. Instead hand the job records stamped with chosen moments and advance from the test the job's own claim that nothing older will arrive, so groups close on command.
solid answer
~50 sA job that groups by the **event moment** — when the thing happened in the world, rather than when the system received the record or when a worker reached it — closes a group using a running assertion that no record older than a stated moment will still arrive. Call that assertion the **completeness claim**. A sleep touches neither the moments nor the claim: it advances the machine's wall clock and hopes the job reacts, which makes the test slow, flaky on a loaded build machine, and silent about the thing in question. The replacement is a *driven source* — an input the test hands records to one at a time and whose completeness claim the test advances itself. Records go in with the stamps and arrival order you chose, and you move the claim exactly where a group should close. Runtimes expose different amounts of this, so check yours before designing the suite.
go deeper
Recall that three different moments are attached to a record and that grouping uses the one the record carries, not the machine's clock. A test that sleeps is testing the clock.
Explain the mechanism: the group closes because the completeness claim moved past its end, so a test must feed the records and move the claim itself rather than wait for either to happen.
Show you have converted a flaky timed suite: what the driven feed controls, which runtimes expose the claim directly, and the honest limit that a deterministic single-process run leaves whole defect classes untouched.
The trade-off is suite design: a fast deterministic core that proves logic, plus a deliberately small number of slow end-to-end runs that prove wiring. Say where you draw the line and what each tier is allowed to claim.
## The three moments, and which one the test is about Every record in a job of this kind has several moments attached to it: the **event moment**, when the thing actually happened in the world; the **arrival moment**, when the system first received the record; and the **processing moment**, when a worker actually reached it. Grouping logic anyone cares about is written against the first. To decide that a group of event moments is finished, a job carries a running assertion that no record older than a stated moment will still arrive — the **completeness claim**. When the claim passes the end of a group, the group can be treated as settled. That logic is the subject of the test: which record lands in which group, and what the job does at the instant the claim crosses a boundary. Not one part of it is a property of the machine's clock. ## What a sleep actually couples you to A test that starts the job, sleeps, then asserts has quietly made several bets: - **That the wall clock is the trigger.** For a job grouping by event moments it usually is not. The group closes because the claim moved, and the claim moved because of records or because someone moved it. - **That today's machine is fast enough.** The same suite on a loaded build agent produces a different answer, so the test reports the agent's load as a defect in the job. - **That the wait is long enough but not too long.** Both errors are silent: too short and the assertion sees a partial result, too long and the suite costs minutes for nothing. - **That the interesting cases are reachable.** They are not. You cannot wait out a record dated yesterday, a two-hour group, or a source that is running six hours behind. - **That a failure will be legible.** `expected 7, got 3` is equally consistent with a wrong aggregate, a wrong boundary, a slow machine and a record still in flight. The usual repair — lengthen the sleep until the red goes away — makes the suite slower and the bet larger without moving it. ## The driven source The replacement is a **driven source**: an input the test hands records to one at a time, whose completeness claim the test advances itself, so nothing waits on the wall clock. With one in place the test owns four things it could not own before: 1. **Membership** — exactly which records exist, and no others. 2. **Arrival order** — the order the job sees them in, deliberately different from event-moment order where that is the point. 3. **Claim position** — where the job believes completeness has reached, at every step of the feed. 4. **Observation points** — assertions run between feeds rather than after a guess. That converts a timing test into an ordinary deterministic one: same input, same sequence of advances, same output, in milliseconds, on any machine. ## How much of this a runtime hands you | What the test needs | Where a runtime offers it directly | Where it does not | |---|---|---| | Hand records over one at a time | A source built for tests that the test feeds | Write into whatever real input the job reads, from the test | | Move the claim without waiting | An explicit way to set it | Feed a record whose event moment is far enough ahead that the claim follows the data | | Run the whole job in one process | An in-process run — every step inside the test's own process, one machine, no network | A small real deployment the suite starts, costing seconds and reintroducing timing | | See each emission as it happens | A collecting destination the test can read | A destination the test polls and reads after each advance | Runtimes in this class differ substantially here: some expose all four, some none, and the difference is worth discovering before the suite is designed rather than after. Where the claim cannot be set directly, driving it through the data is equivalent — what matters is that the *test* decides when it moves. ## Pinning the machine clock is a different fix A suite may also pin the machine's clock so that anything genuinely wall-clock-driven becomes reproducible. That is ordinary test craft and is owned elsewhere, and it is not a substitute here. A job grouping by event moments does not consult the machine clock to decide a group, so freezing it changes nothing about the property under test: the records still have to be supplied and the claim still has to move. ## What determinism does not buy A driven, sleep-free test proves the time logic and nothing else. It normally runs in one process, over a handful of records, from one input, so whole classes of defect — what moving records between workers, losing a worker, or recovering does to the same logic — are not present to be caught, and that is its own subject. State the claim precisely: these records, in this order, with the claim moved here, produce this output. That is far more than a sleep proved, and it is not everything.
- What if the runtime gives a test no way to set the completeness claim directly?Drive it through the data instead. Where a runtime derives the claim from the event moments it has seen, feeding a record whose moment is far enough ahead moves it; under a model that runs continuous work as a fast series of small finite runs, you control it by controlling which records land in each small run. Either way the test, not the wall clock, decides.
- Does pinning the machine clock remove the need for a driven source?No. A pinned clock makes genuinely wall-clock-driven behaviour reproducible, but grouping by event moments never consults the machine clock in the first place. You still have to supply the records and move the claim, so the pin changes nothing about the property under test.
- When is a slow, waiting test still worth keeping?As a small number of end-to-end runs against a real deployment, where the point is the wiring rather than the logic: that a real input actually gets an event moment assigned, and that output reaches its destination. Keep them out of the fast suite; they are slow and fail for reasons unrelated to the code under test.
A rehearsal for a midnight scene. The director calls out that it is now midnight and watches what the cast does; nobody sits in the theatre until midnight to find out. The behaviour at the boundary is the subject, and the waiting adds nothing to it.
saying these in an interview costs you the question
- Just lengthen the sleep until the test stops failing
- It passed on my machine, so the grouping is proven
- The event moment is whatever the clock says when the worker processes the record
- Time-dependent logic cannot be tested, only watched in production
- A sleep is fine because the build machine is fast