skip to content

A feature history keeps one value per 30-day event window, recomputed as late trips arrive - why can a join bounded at the quote's event time still leak?

level: seniorimportance: should knowfreq 42%

answer

  1. two bounds, only one enforced
  2. the window ended, the value did not
  3. recomputation is its own clock
  4. hold features behind a settled point
  5. measure the upload-delay tail first

basics

~20 s

An event-time bound only limits which trips the value summarises; it says nothing about when that value was computed. A window recomputed after late uploads carries figures the quoting path never held, so the row still reads the future.

solid answer

~50 s

There are two independent time bounds and this design enforces only one. Bounding on event time keeps post-quote driving out of the window - necessary, and not sufficient, because the stored value for a window that ended before the quote has been recomputed since, absorbing trips that uploaded after the quote was priced. The row reads a number that did not exist at `quoted_at`. Two fixes exist. Version the value in knowledge time, so the join can select the version the store actually held. Or, where only event-time history is kept, introduce a **settled-history watermark**: measure how late uploads arrive, pick a lag that covers nearly all of them, and let a row read only windows that ended at least that long before its own quote instant - by which point the value had stopped moving.

code

pseudocode · 12 lines
pseudocode
LATE_ARRIVAL_BOUND = 7 days

function training_feature(quote):
    cutoff = quote.quoted_at - LATE_ARRIVAL_BOUND

    row = latest_window_ending_at_or_before(
              featureHistory, quote.driver_id, cutoff)

    if row is null:
        return NO_SETTLED_HISTORY

    return row.value

go deeper

for a junior

Hold on to the distinction: knowing when the driving happened is not the same as knowing when the number describing it was produced. Only the second says what a past reader could have seen.

for a middle

Walk the four-step timeline - window closes, quote priced, late trip lands, window recomputed - and point at the step where the training row picks up a value that did not exist at the quote.

for a senior

Compare the two fixes on their real costs: versioned history is exact and pays in storage and query complexity, a watermark is cheap and pays in staleness. Derive the lag from the measured arrival distribution rather than asserting it.

for a principal

Decide which guarantee the business needs. Explaining an individual past price to a regulator or a customer demands versioned history; training data that merely has to generalise can accept a measured watermark, and that call has a cost attached.

## Two bounds, not one Point-in-time correctness needs two independent constraints on the same training row, and they are easy to mistake for each other because both are expressed in timestamps. | bound | what it constrains | what it prevents | |---|---|---| | event time | which trips the aggregate summarises | driving that happened after the quote entering the feature | | knowledge time | which computation of that aggregate is read | a later recomputation of an earlier window entering the feature | A history keyed only by event window has no way to express the second one. There is exactly one value per window per driver, and it is whatever the last recomputation produced. ## How the leak survives an event-time bound 1. A 30-day window closes on 31 March. The pipeline computes `harsh_braking_per_100km` = 1.4. 2. A quote is priced on 3 April. The pricing path reads 1.4 and sets a premium on it. 3. On 5 April a buffered trip from 28 March uploads. It belongs to that window by event time, so the pipeline recomputes: 2.9. 4. In September the training job assembles the row for that 3 April quote. It applies its event-time bound - the window ends 31 March, comfortably before 3 April - and reads 2.9. Every timestamp in step 4 checks out and the row is still wrong. The value it carries was created two days after the decision it is supposed to explain. Where those late-arriving trips correlate with what eventually happened - a burst of unusual driving before an incident, a device that went quiet during a difficult period - the model learns from the correlation, offline error drops, and the pricing path can never reproduce it. ## Fix one: version by knowledge time Store each computation as its own version with a validity interval, and the join selects the version whose interval contains the quote instant. This is exact: a 3 April quote reads 1.4 forever. It costs a row per change and an extra predicate on every offline read, and it is the only option that reproduces the past perfectly. ## Fix two: a settled-history watermark Where history is event-time only, hold features back until they have stopped moving: - Measure the upload-delay distribution from the records themselves - the gap between event time and ingestion time. Suppose 95% of trips arrive within 2 days and 99% within 7. - Choose a lag that covers the tail you care about; take 7 days. - The **watermark** at any instant is that instant minus 7 days: the point before which the history is treated as final. - A row quoted at `quoted_at` may read only a window that ended at or before `quoted_at - 7 days`. By construction that window had already absorbed nearly all of its stragglers before the quote was priced, so the stored value is close to what the pricing path read. The arithmetic runs one way only. The newest usable window ends **before** the quote, never at it and certainly never after it. A window ending after the quote instant is the future outright, and one ending exactly at it is still absorbing uploads. ## What the watermark costs, and what it does not fix - **Staleness by construction.** Every score is computed on driving as of at least a week ago. A driver who changed behaviour last Tuesday is priced as though they had not. - **A tail you chose to drop.** The 1% of trips arriving after 7 days never make it into a value anyone reads at decision time, and lengthening the lag to catch them makes every feature staler. - **It is approximate, not exact.** It removes the class of leakage that comes from recomputation, because the value had settled; it does not prove that a particular quote read that particular number. - **It does nothing about an unbounded lookback.** A dormant driver's year-old window is perfectly settled and still the wrong value to present as current. ## Choosing between them Versioning is exact and costs storage and query complexity. The watermark is cheap and costs freshness. A platform that must explain individual past decisions needs versioning; one that only needs training data honest enough to generalise can often live with a watermark, provided the lag is measured from real arrival data rather than guessed, and re-measured when the fleet or the upload path changes.

  • How do you choose the watermark lag rather than guessing at it?
    Measure it from the data the pipeline already carries: for every record, the gap between event time and ingestion time. That distribution gives the fraction of trips captured at each candidate lag, and the lag is the point where the curve flattens and further waiting buys almost nothing. Re-measure it when the device fleet or upload path changes, because the tail is a property of the hardware and the network, not a constant.
  • If the watermark makes every feature a week old, is that a bug?
    No, it is the price, and it is paid consistently. The model is trained on driving as of a week before each quote and prices on driving as of a week before each new quote, so the two agree. It becomes a bug only if one side of the system honours the lag and the other does not, or if the business genuinely needs to react to behaviour inside that week - in which case the answer is versioned history, not a shorter lag.

saying these in an interview costs you the question

  • Believes an event-time bound alone makes the join point-in-time correct.
  • Says a window that has closed can no longer change.
  • Sets the watermark lag by intuition instead of the arrival distribution.
  • Puts the usable window's end after the quote instant.
  • Claims a watermark also fixes reading an ancient value as current.