A continuous job keeps receiving records stamped earlier than ones it already handled — why is that normal, and what causes it?
answer
- arrival order is not occurrence order
- a transport property, not a fault
- buffering, retries, merged parallel inputs
- uniform delay creates no spread
- out of order is not yet late
basics
~20 sOut-of-order arrival is a property of the transport, not a fault: buffered devices, retries and parallel inputs merged together all deliver an earlier-stamped record after a later one. It costs nothing by itself until a time-based grouping has to be treated as finished.
solid answer
~50 sEvery record carries two moments that need not agree: **occurrence time**, the moment the thing happened, written into the payload by whatever produced it, and the moment the record actually reached the job. Nothing along the path preserves the first order into the second. A phone buffers readings while it has no signal and flushes them on reconnect; a failed send is retried and lands after everything produced meanwhile; a job reading several parallel inputs interleaves them in whatever order its readers happen to pull. So the job routinely handles a record stamped `12:04:10` after one stamped `12:04:30`. That is *disorder*, and by itself it is free — the record still belongs in its own group. It only becomes a loss once the job has already stopped expecting anything that old, which is a separate decision made elsewhere.
go deeper
Recall that arrival order and occurrence order are two different sequences, and be able to name two ordinary causes — a device that buffered while offline, and a retried send.
Explain why a delay applied uniformly creates no disorder while a variable one does, and why merging several parallel inputs interleaves stamped moments even when each input is ordered.
Show that you treat disorder as a measured property of one specific source rather than a constant, and that you separate it from the job's own backlog when diagnosing.
Frame disorder as part of the source contract: which producers may go offline, for how long, and what that implies for every published number derived from their records.
## Two moments, two orders A record in a continuous pipeline carries more than one moment. The first is **occurrence time**: the moment the thing being recorded actually happened, written into the payload by whatever produced the record. The second is the moment the record reached a system that could observe it — either **receipt time**, when the system that first accepted it wrote it down, or the wall clock of the worker that eventually handled it. Nothing between those moments preserves ordering. Arrival order is the order of the transport; occurrence order is the order of the world. They are two different sequences over the same records, and expecting them to coincide is the mistake the whole subject exists to correct. So a job reading a durable, re-readable input — an input whose records can be read again from an earlier position, each carrying a source timestamp — routinely handles a record stamped `12:04:10` after one stamped `12:04:30`. That is **disorder**, and it is the ordinary condition of every real source, not a symptom of a broken one. ## Where the disorder comes from - **Producers that buffer.** A phone, a vehicle, a meter or a sensor with intermittent connectivity records locally and flushes when it can. The gap may be seconds or days, and this is the single largest contributor to a long tail. - **Retries.** A send that fails and is retried arrives after everything produced in the meantime. - **Many producers, many paths.** Records from different producers travel different routes, through different regions, with different queueing behaviour along the way. - **Merging parallel inputs.** Even when each parallel input of the source is perfectly ordered within itself, a job reading several of them interleaves them in whatever order its readers happen to pull, so the merged sequence is unordered by construction. - **Batching at the producer.** A producer that accumulates records and ships them every thirty seconds delivers a block whose oldest member is thirty seconds behind its newest. - **Producers that disagree about what time it is.** Two producers stamping from their own clocks contribute apparent disorder; the accuracy of machine clocks themselves is a separate subject from what the pipeline does about the resulting spread. ## A uniform delay is not disorder The distinction most candidates miss is that a delay applied equally to everything creates no disorder at all. Only the *spread* matters. | pattern | what the job sees | effect on a time-based grouping | |---|---|---| | every record delayed by ten minutes | arrival order equals occurrence order | none — results are simply ten minutes old | | most records within a second, some by ten minutes | an earlier-stamped record arrives after a later-stamped one | a ten-minute group cannot be treated as finished promptly | | a producer offline for a week, then reconnecting | records arrive a week after their stamped moment | no wait anyone would accept covers them | This is why "the pipeline is ten minutes behind" and "the source runs ten minutes out of order" are different statements with different remedies. The first is about how fast the job consumes; the second is about how the source delivers, and a job that is perfectly caught up still sees the second. ## Out of order is not yet late These two words are routinely used interchangeably and mean different things: 1. **Out of order** — the record arrived after a record with a later stamped moment, but before the job stopped expecting its own time range. It still lands in its own group. Nothing is lost. 2. **Late** — the job had already stopped expecting anything that old by the time the record arrived, so its group was already eligible to be treated as finished. Every late record arrived out of order; most out-of-order records are never late. Keeping the two apart is the difference between a candidate who understands the subject and one who thinks every delayed record is a data-loss incident. ## What the engine's execution mode changes Engines of this class do not observe disorder at the same granularity, so describe the mode before describing the effect: - A runtime that executes continuous work as a **rapid succession of small finite jobs** only reconsiders time at the boundary of each small job. Disorder entirely inside one of those intervals is invisible to it, because everything in the interval is handled together. - A **record-at-a-time** runtime sees each record as it arrives and can move its notion of time between any two records, so the same source's disorder is visible at much finer resolution. - A **single pass over a finished bounded input** sees no disorder in the operational sense at all: every record is already present when the read starts, so putting them in order is a sorting cost, not a waiting problem. The practical consequence for a junior answer is simply this: disorder is a property of the source and its transport, it is present in every real stream, and what a job does about it is a deliberate choice rather than something the runtime handles invisibly.
- Does a record arriving out of order automatically become a problem?No. Arriving after a record with a later stamped moment is ordinary: the record still belongs in, and lands in, its own time group. It becomes a problem only if the job had already stopped expecting anything that old — a separate condition, reached by a minority of out-of-order records on most sources.
- Why does merging several parallel inputs create disorder even when each input is perfectly ordered internally?Because the job pulls from them independently. One input may be read ahead while another lags, so a record stamped earlier on the lagging input surfaces after a later-stamped record from the one running ahead. The per-input order is intact; the merged sequence has no order at all.
- Would processing the input faster reduce the disorder the job observes?No. Disorder is created upstream, by producers and transport, and is already fixed by the time the records are available to read. Processing faster reduces backlog — how far behind the job is — which is a different quantity entirely and can be zero while disorder is large.
saying these in an interview costs you the question
- Claims out-of-order records mean a broken producer or corrupt data
- Assumes records reach the job sorted by the moment they occurred
- Says every delayed record is late and will therefore be dropped
- Treats a uniform transport delay as disorder that needs a wait
- Attributes all disorder to unsynchronised machine clocks and stops there
- Thinks reading faster or adding workers would remove out-of-order arrival