How should a playlist ranker's training set handle actions that arrive after the training snapshot was cut?
answer
- unmatured is not negative
- wait one window before closing
- horizon sits before the cut
- the freshest hours never train
- append a version, never mutate
basics
~20 sClose labels only for impressions whose attribution window ended before the cut; anything younger is unmatured and must be excluded, never written as a negative. Corrections that arrive later go into a new snapshot version rather than editing the old one.
solid answer
~40 sTwo things go wrong at the cut. First, an impression whose action has not arrived yet looks identical to one that was genuinely ignored, so a label job that writes `absent = negative` mislabels exactly the slow and late-uploading population - offline listening, deep playlist slots, deliberate saves. Second, if you re-run the job later to pick those actions up, the snapshot changes underneath a model that already trained on it, and two runs on the same named dataset stop agreeing. The discipline is a **maturity horizon**: labels close only for impressions served before `cut - window`, everything newer is excluded as unknown, and late arrivals beyond the horizon land in the **next** immutable snapshot version. The price is that the freshest hours never train, and that price is the honest one.
go deeper
Know that an action can arrive after the data for a training run has been fixed, and that treating a missing action as a rejection is the mistake this creates.
Explain the maturity horizon: labels close one attribution window before the cut, newer impressions are excluded as unknown, and the horizon differs per action type because the windows do.
Show the operational discipline - immutable versioned snapshots, corrections appended in the next version, and monitored late-arrival fractions sliced by client - plus the diagnosis when a tight horizon shows up as a segment quality gap.
The judgment is how much of the freshest data you are willing to leave untrained in exchange for labels that mean what they say, and whether shortening the window is the cheaper concession for this product.
## Two clocks, and the gap between them A label job runs against a snapshot with a cut: a timestamp after which nothing is included. But the actions it is labelling live on a different clock. A listener acts at their own pace, a device uploads when it reconnects, and the event pipeline delivers when it delivers. At the instant the cut is taken, some impressions have their final outcome recorded, some have an outcome that exists but has not landed, and some have an outcome that has not happened yet. The label job cannot tell these three apart from the absence of a row. ## The default that is wrong The naive rule is: every rendered impression with no matching action is a negative. Applied at the cut, it produces a mislabelling that is **systematic, not random**, and concentrated in the newest slice of data: - impressions served in the final hour before the cut, whose window has barely opened; - listeners on devices that upload in bursts after reconnecting; - the slow actions - a save the next morning - that the attribution window was widened to catch in the first place. The model that results under-predicts for exactly those tracks and those listeners. In production it presents as a quality gap in one segment, which sends people looking at features and the model, when the cause is a boundary in the label job. ## The maturity horizon The fix is a second boundary inside the snapshot. If the attribution window for an action type is `W` and the snapshot cut is `T`, then labels may only be closed for impressions served at or before `T - W`. Everything served in `(T - W, T]` is **unmatured**: it is carried as unknown and excluded from training, not written as a negative. | impression served at | window state at the cut | what the label job emits | |---|---|---| | before `T - W` with a matched action | closed | a positive of that action type | | before `T - W` with no matched action | closed | a negative | | after `T - W` | still open | nothing - excluded as unmatured | The horizon must be per action type, because the windows are: skips mature in seconds, saves in days. A snapshot can legitimately carry closed skip labels much closer to the cut than closed save labels. ## Corrections, and why the snapshot never changes Some actions still land after the horizon - a device that was offline for a week, a pipeline backlog that drained late. The temptation is to go back and fix the rows. Do not: 1. A model has already trained on that snapshot, and its evaluation numbers refer to the data as it was. Mutating it breaks the ability to reproduce or explain that run. 2. Two training jobs reading the same snapshot name at different times would then see different data, which is a class of bug that is nearly impossible to diagnose from a metric. 3. The corrections themselves are evidence: their volume and their distribution over clients tell you whether the horizon is set correctly. Instead, snapshots are immutable and versioned, corrections accumulate into the next version, and every training run records the exact snapshot version it consumed. Content-addressing the snapshot makes the guarantee mechanical rather than procedural. ## What to measure, and what it costs The horizon is a number, so it should be derived and monitored rather than assumed: - **Late-arrival fraction**: what share of actions land after `T - W + W`, i.e. after the horizon that was used. If it is rising, the horizon is stale. - **Late arrivals by client and platform**: a single client holding events for hours is an instrumentation fact, not a behavioural one. - **Unmatured volume**: how much of the freshest data is being excluded each run, which is the direct cost of the horizon. And the cost is real and unavoidable: the most recent stretch of data is never in any training set at the moment it would be most useful. A team that finds this intolerable has one honest option - shorten the window, knowing precisely which slow actions that discards - and one dishonest one, which is to move the cut forward and let absence stand in for silence. The second option produces a fresher dataset that is wrong in a direction nobody will notice until a segment starts underperforming.
- If the horizon is set too tight, what does the damage look like in the model?It looks like a segment problem, not a data problem. The ranker under-predicts for slow-converting tracks and for populations whose devices upload late, because their positives were written into the training set as negatives. Nothing in the pipeline errors, and the offline metric is computed on the same corrupted labels, so it agrees.
- How do you keep a snapshot reproducible when corrections keep arriving?Make snapshots immutable and content-addressed, let corrections accumulate into the next version, and have every training run record which version it consumed. Then a past run can be re-executed exactly, and the delta between two versions is itself a measurement of how late your actions really arrive.
saying these in an interview costs you the question
- Writes an impression with no action yet as a negative label.
- Re-runs the label job over an existing snapshot in place.
- Assumes every action reaches the store before the cut.
- Sets the training cut to now so the data stays fresh.
- Treats late-arriving plays as data loss rather than an expected tail.