skip to content

A trip driven Tuesday is uploaded Friday: in a telematics feature pipeline, what distinguishes its event time from its ingestion time?

level: juniorimportance: must knowfreq 62%

answer

  1. when it happened versus when it arrived
  2. devices buffer, uplinks drop
  3. aggregates are defined over event time
  4. knowledge time answers what was known
  5. closed windows keep changing

basics

~20 s

Event time is when the driving happened; ingestion time is when the trip summary reached the feature pipeline and became readable. Devices buffer and retry, so the two differ by hours or days and records arrive out of order.

solid answer

~40 s

Every trip record carries at least two timestamps. Event time is a property of the world - the drive started and ended on Tuesday - and it never changes. Ingestion time is a property of the pipeline: Friday, when the upload landed and the value became visible to anything reading the store. The gap comes from real causes, such as a device with no connectivity, batching to save power, or a failed upload that retried later. The two answer different questions. `How much did this driver drive in March?` is an event-time question. `What did the pricing path hold at 14:05 on 3 April?` is an ingestion-time question. Conflating them puts Tuesday's driving in Friday's window and makes it impossible to reconstruct what was known when.

go deeper

for a junior

Be able to name the two timestamps and give one honest reason they differ, such as a device buffering with no connectivity. Say which one describes the world and which one describes the pipeline.

for a middle

Explain what out-of-order arrival does to a window that has already closed, and why an aggregate over a fixed event window can legitimately hold two values a week apart.

for a senior

Show the operational consequence: a connectivity outage that reads as low mileage, then a backlog burst that reads as high mileage, and the pricing decisions taken in between on values that were never wrong at the time.

for a principal

Frame the choice of how long to wait for stragglers as a business trade-off - pricing on fresher but less complete behaviour, against pricing later on a settled account - and say who owns the number.

## Two clocks on one record A trip summary from an in-car device carries a timestamp for the drive itself and picks up a second one when it lands. They are easy to confuse because a single record holds both, and in a system with a fast, reliable uplink they are close enough that nothing breaks for a while. - **Event time** - when the trip happened. Fixed by the world, immutable, and the only clock that answers questions about behaviour. - **Ingestion time** (also called arrival or processing time) - when the record reached the feature pipeline and became visible to a reader. Fixed by the infrastructure, and it is the clock that decides what a system standing at some moment could have known. ## Why they diverge, routinely - The device sits in an underground car park or a rural gap with no uplink and buffers whatever it collects. - It batches uploads deliberately, to save power and data allowance, rather than streaming each trip. - An upload fails and is retried on the next successful connection, hours or days later. - The phone acting as the uplink is off, out of storage, or has the application suspended. - A replaced or re-provisioned device flushes a backlog in one burst. Late arrival is therefore the normal case, not a defect report. A pipeline designed as though every trip appears the moment it ends will be wrong every week. ## Which clock answers which question | the question | the clock that answers it | |---|---| | How many night-time kilometres did this driver cover in March? | event time | | What did the feature store hold for this driver on 3 April? | ingestion time | | Is the pipeline falling behind? | the gap between the two | | Which trips are still missing from a window that already closed? | event time, discovered through later ingestion | | Which aggregate version did this quote actually read? | ingestion time | ## Out-of-order arrival and the window that keeps moving Because records keep arriving for periods that have already ended, an aggregate over a closed event window is not a single number - it is a number that changes for a while after the window closes. A worked sequence: 1. Tuesday: the driver takes a long motorway trip. Event time is Tuesday. 2. Wednesday: a quote is priced. The 30-day aggregate does not contain Tuesday's trip, because nothing has uploaded it. 3. Friday: the device uploads the backlog. Ingestion time is Friday. 4. Friday night: the aggregate for that same 30-day event window is recomputed and now includes the trip. 5. Next Monday: a second quote for the same driver reads a materially different value, over a window that overlaps almost entirely with the first. Both quotes are correct about what was known at the time. Anything that wants to explain the first quote later must be able to recover the pre-Friday value. ## What breaks when the two are conflated - **Windows built on ingestion time misattribute driving.** Tuesday's kilometres land in Friday's bucket, and every boundary row of every window is wrong. - **Outages become behaviour.** A three-day connectivity gap reads as a quiet week for those drivers, followed by a record week when the backlog lands - a pattern in the data with no counterpart in the world. - **History cannot be reconstructed.** With only an event-time key, there is no way to ask what a value was on a given date, because the stored value has been recomputed since. - **Comparisons across drivers are unfair.** Drivers with poor connectivity look like low-mileage drivers, which is a pricing consequence rather than a data-quality footnote. ## What a well-formed feature row carries Four timestamps, not one: the **start and end of the event window** the value summarises, and the **interval of knowledge time** during which the store held that value. The first pair answers `which driving is this?`; the second answers `when did we believe it?`. Keep both and any past moment can be reconstructed exactly; keep either alone and it cannot.

  • Which clock bounds the 30-day aggregate, and which one selects the version an as-of join reads?
    The aggregate's window is event time: it says which driving is summarised, so a trip driven inside the window belongs there however late it uploads. The version selection is knowledge time: it says which value the store held at the quote instant. Both bounds apply to the same training row, and dropping either one lets a value in that no reader could have seen at that moment.
  • Why can the same aggregate, over the same event window, hold two different values a week apart?
    Because records keep arriving for a period after it ends. When a buffered trip uploads, the pipeline recomputes the window it belongs to and the value moves. Nothing is wrong: the first value was a complete account of what had arrived, and the second is a complete account of what has arrived now. The difference is knowledge, not driving.

saying these in an interview costs you the question

  • Treats the upload timestamp as the time the driving happened.
  • Assumes every trip for a day has arrived once that day ends.
  • Calls out-of-order arrival a device defect rather than the normal case.
  • Uses the pipeline's processing clock to define a 30-day driving window.
  • Believes one timestamp per row is enough to reconstruct past state.