Which records, in what arrival order, would you hand a job to prove its grouping tolerates out-of-order input?
answer
- arrival order is yours to choose
- disorder before the claim passes
- one record after its group settled
- advance the claim between feeds
- assert at each advance, not only at the end
basics
~20 sFeed records whose event moments deliberately do not ascend, and move the job's completeness claim in steps you choose: one advance that leaves the group open, one that closes it, and one record fed afterwards.
solid answer
~50 sCraft the feed around the three cases the code must separate. First, a record whose **event moment** — when the thing happened, as opposed to when it arrived — is earlier than one already fed but whose group is still open: it must be counted. Second, a record on the boundary itself, which pins whether the interval's end is inclusive. Third, a record handed over after the job's running claim that nothing older will arrive has already passed its group's end: late relative to its own group. Feed them one at a time, in the arrival order you chose, and advance the claim explicitly *between* records rather than once at the end — the mid-feed advance is the only thing that separates "still open" from "closed". What the job then does with the late one varies, so assert the policy your job is actually on.
code
python · 15 lines# The test hands over one record at a time and moves the completeness
# claim itself. No sleeps, and nothing depends on the machine's clock.
feed(event_moment="10:00:30", key="a", amount=1)
feed(event_moment="10:00:10", key="a", amount=1) # earlier than the last: disorder, still in time
advance_claim_to("10:00:59") # group [10:00, 10:01) is not complete yet
seen = collect_emissions() # some runtimes emit an early, revisable value here
feed(event_moment="10:01:20", key="a", amount=1) # belongs to the next group
advance_claim_to("10:01:30") # [10:00, 10:01) may settle; [10:01, 10:02) may not
seen += collect_emissions()
feed(event_moment="10:00:55", key="a", amount=1) # late: its own group's claim has already passed
seen += collect_emissions() # drop, divert or revise - assert the policy chosengo deeper
Remember that arrival order and the order things happened are different, and that a test fixture sorted by event moment cannot demonstrate anything about disorder.
Explain which records you craft and why: one out of order but in time, one on the boundary, one after its group settled, plus where you move the claim between them.
Demonstrate that the tolerance came from measured lateness on real input, that the late-record policy is asserted by name rather than inherited, and that the far-future stamp has its own test.
The judgment is how much disorder the platform promises to absorb: a wider tolerance buys correctness and costs latency and retained records. Set that as a standing decision, not per pipeline.
## Two different things hide in "out of order" A record is out of order when its **event moment** — when the thing happened in the world — is earlier than one the job has already seen. That single phrase covers two cases with different correct behaviours: - **Out of order but still in time.** The record's group has not yet been declared complete, because the job's running claim that nothing older will arrive has not passed the group's end. The record simply belongs in the group and must be counted. - **Late relative to its own group's claim.** The claim has already passed that group's end, so the group was settled before the record showed up. Now a policy decides its fate, and the policy is a choice. A test that only proves the first has proved the easy half. Keep the second sense of "late" distinct from a job that is behind its input overall, which is a different subject with unrelated remedies. ## The records worth crafting 1. **A plain in-order record.** The control: without it, a failure elsewhere is hard to localise. 2. **A record earlier than one already fed, inside a still-open group.** Assert it lands in that group and changes its value. 3. **A record exactly on a boundary.** Whether the interval's end is inclusive is a genuine defect source and no reader should have to guess; the test should pin it. 4. **A record fed after the claim passed its group's end.** This is the one that distinguishes a job with a considered policy from one that never thought about it. 5. **A record with an event moment far in the future.** It is what a corrupt or mis-parsed stamp looks like in the wild, and where the claim follows the data it drags the claim past everything and settles every open group at once. Both consequences deserve a named test. ## Where the advances go The arrival order is yours, and so is the claim's position at each point. Advance it **between** feeds: once to a moment short of the first group's end, so you can show the group is still open; then past the first group's end but short of the next group's, so exactly one group settles. An advance only at the end of the feed collapses every case into one and proves almost nothing, because you can no longer tell whether a record was counted for the right reason. ## What a correct job may do with the late record | Policy | What the test asserts | What it costs | |---|---|---| | Drop it | The settled value is unchanged, and any counter of discarded records has risen | The published number is quietly short, and only the counter says so | | Divert it aside | It appears on a separate output; the settled value is unchanged | Somebody has to consume and reconcile that output | | Revise the group | The group is emitted again carrying a new value | Every consumer downstream has to accept a correction rather than an append | Runtimes differ in which of these they offer and which applies when nobody chose, so treat the policy as part of the job's contract and assert it by name in the test. ## Traps in the crafting - **Sorting the input first.** A fixture built by sorting records by event moment cannot fail the test it was written for. - **Handing over the whole collection at once.** Then the runtime, not the test, picks the interleaving; under a model that runs continuous work as a fast series of small finite runs, one batch containing everything makes the disorder invisible. - **Crafting less disorder than production has.** Three seconds of crafted lateness says nothing about an input whose tail routinely runs minutes behind. - **Assuming the record arrives with an event moment.** The pipeline assigns it, usually from a field; a pipeline that never assigned one is silently grouping by arrival, and a fixture that always sets the field will never reveal that. - **Asserting only the final state.** The intermediate observations are where the boundary behaviour shows. ## How much disorder is enough Measure rather than imagine. The useful number is the observed gap between the event moment and the arrival moment on real input — its typical value and its tail — and the crafted records should straddle whatever tolerance the job intends to honour: one comfortably inside it, one just outside. That makes the test a statement about a decision someone made, instead of a statement about what the fixture's author happened to picture.
- Why feed the records one at a time instead of handing over the whole list?Because the variable under test is what the job had seen, and where the claim sat, at the moment each record arrived. A whole collection lets the runtime choose that interleaving, and under a model that processes continuous work as a series of small finite runs it collapses the scenario into one run where everything is present together.
- How do you decide how much disorder to craft?From measurement, not intuition: take the observed gap between event moment and arrival moment on real input, look at its tail, and craft one record inside the tolerance the job intends to honour and one beyond it. A fixture built on imagined disorder tests the author's imagination.
- What does the far-future record catch that the others do not?A stamp parsed from the wrong field or the wrong unit. Where the claim is derived from the moments seen, one such record pushes it past every open group at once, so every subsequent real record is late and quietly discarded. The test proves whether the job rejects or clamps such a value.
saying these in an interview costs you the question
- Sort the fixture by event moment before feeding it in
- Out-of-order arrival is the source's problem, not the job's
- A record arriving after its group settled is always discarded
- Feed a far-future record first so the claim is out of the way
- If the totals match at the end, ordering was handled correctly