skip to content

In a playlist ranker's label job, what does a short attribution window on saves quietly drop?

level: middleimportance: should knowfreq 48%

answer

  1. the window is a causal assumption
  2. short is not neutral, it selects
  3. drops slow actions, keeps fast ones
  4. measured in event time
  5. longer window, later closable labels

basics

~20 s

It drops the slow saves - the track someone returns to that evening - and the loss is not random: it removes long playlists, background listening and deliberate second thoughts, leaving a positive class made of instant reactions.

solid answer

~50 s

The attribution window is the span after an impression in which an action still counts as caused by it. Shortening it does not just reduce volume, it **selects**: a 30-minute window keeps the save made during the first listen and drops the one made the next morning, so slow-burn tracks, deep playlist positions and listeners who queue music for later fall out of the positive class. The ranker then learns to optimise immediate reactions. Lengthening the window recovers those positives but buys mis-attribution - the listener may have found the track through search or a friend - and it pushes back the moment labels can be closed, since every impression needs its full window to elapse first. The window is measured in **event time**, not in when the label job saw the record, or a device that uploads late falls outside it for no behavioural reason.

code

pseudocode · 13 lines
pseudocode
for each play in playEvents:
    impression = lookup(impressionStore, play.requestId, play.trackId)

    if impression is null:
        emit unattributed(play, reason = "no impression for request")
        continue

    delay = play.eventTime - impression.servedAt

    if delay < 0 or delay > WINDOW[play.actionType]:
        emit unattributed(play, reason = "outside window", delay = delay)
    else:
        emit label(impression, positive, actionType = play.actionType)

go deeper

for a junior

Know that an action only counts as a label if it happens within a chosen span after the impression, and that the span is a decision someone made rather than a property of the data.

for a middle

Explain why shortening the window biases rather than merely shrinks the positive class, and why the window is measured between event timestamps rather than against when records were processed.

for a senior

Show you would derive the number from the delay distribution per action type, inspect the excluded tail for a distinct population, and version the window alongside the label definition it changes.

for a principal

The open call is how much label completeness a fresher training set is worth. A longer window buys slow-intent positives and costs both mis-attribution and staleness, and only the product can say which side it wants to be wrong on.

## What the window is, and what it is deciding An attribution window is the interval after an impression within which a later action is treated as having been caused by that impression. It is not a technical detail of the join; it is a causal assumption written as a number. Choosing 30 minutes for a save asserts that a save made 31 minutes later had some other cause. Nobody believes that literally - the number is a tolerable approximation, and the job is to know which way it is wrong. ## A short window removes a specific kind of action, not a random sample The crucial property is that the delay between seeing a track and acting on it is **correlated with the kind of listening**. Cutting the window therefore removes a biased slice: - **Deliberate saves** happen after the track has been heard once and thought about, often hours later. - **Deep-slot impressions** are reached late in a long playlist, so the remaining window at the moment of exposure is already short. - **Queued or offline listening** produces actions that are made, and sometimes uploaded, long after the impression. - **Shared devices and background sessions** produce sparse, delayed interactions of every kind. What survives a short window is the immediate reaction. Train on that and the ranker learns a preference for tracks that provoke an instant response - a real preference, but a narrower one than the product asked for, and one nothing in the offline metric will flag, because the metric is computed on the same truncated labels. ## What a long window buys and what it costs | window | positives captured | mis-attribution | delay before labels close | |---|---|---|---| | minutes | instant reactions only | very low | negligible | | hours | most of the delay distribution's mass | moderate | hours | | days | nearly all, including second thoughts | high - other discovery paths dominate the tail | days | Beyond the bulk of the delay distribution you are mostly buying noise: the further out you go, the larger the share of saves that the listener reached through search, a friend's recommendation or another surface entirely. Those rows teach the ranker that a track it happened to show was responsible for something it did not cause. The second cost is structural rather than statistical. An impression's label is not final until its window has elapsed, so a longer window pushes the frontier of closable labels further into the past. The training set is then either fresher and partly wrong, or complete and staler. That is a real trade, and it is paid by whoever owns the refresh schedule, not by the label job. ## Event time, not processing time The window must be measured between the serve timestamp and the action's own event time, both as the events themselves report them. Measuring against the time the record reached the store punishes exactly the population you least want to drop: a device that was offline uploads its actions in a burst hours later, and every one of them falls outside a processing-time window despite the listener having acted within seconds. The result is a training set that under-represents mobile and offline listening for a reason that has nothing to do with music. ## Choosing the number instead of guessing it 1. Plot the delay from impression to save on historical data, out to several days. 2. Pick the quantile that captures the bulk of the mass - the point where the curve flattens. 3. Inspect the excluded tail directly: are those tracks, listeners or slots systematically different from the included ones? If yes, the cut is doing damage a volume count will not show. 4. Set a different window per action type. A skip resolves in seconds; a completed play in minutes; a save in hours to days. One window for all three is one assumption applied to three behaviours. 5. Record the window in the label spec and version it, because changing it changes every downstream number, including the baselines a future comparison uses. ## The failure mode to name in an interview The worst version of this is a window chosen to fit the batch schedule - one hour because the job runs hourly. The number then encodes an operational convenience as a causal claim, and every model trained afterwards inherits it. Pick the window from the behaviour and let the schedule accommodate it, or change the schedule deliberately and write down what the shorter window is known to drop.

  • How would you choose the window length for saves rather than guess it?
    Plot the impression-to-save delay distribution and pick the quantile where it flattens, then look at the excluded tail to check it is not a distinct population. Use a different window per action type, since a skip resolves in seconds and a save in days, and write the chosen numbers into the versioned label spec.
  • Does a longer window always produce better labels?
    No. Past the bulk of the delay distribution the added rows are increasingly saves the listener reached by some other route, so you are attributing effects the impression did not cause. You also push back the point at which labels can be closed, making the training set staler for whoever schedules the refresh.

saying these in an interview costs you the question

  • Picks the window to match the batch schedule rather than the behaviour.
  • Thinks a shorter window loses volume but not a particular kind of action.
  • Measures the window against when the label job saw the event.
  • Assumes every action inside the window was caused by the impression.
  • Uses one window for skips, completed plays and saves alike.