skip to content

What must a playlist ranking service log at serve time so a later play can be joined to the ranking that produced it?

level: middleimportance: must knowfreq 62%

answer

  1. one record per serving request
  2. log what was shown, never re-score
  3. position and propensity die at serve time
  4. join on request id, not timestamps
  5. stamp model and feature versions

basics

~20 s

One impression record per serving request: a request id the client echoes on every later event, each slot's track id, position and rendered flag, the serve-time score and any randomisation propensity, plus the model and feature-spec versions.

solid answer

~50 s

The training row is `(what was shown, what happened next)`, and only the serving path knows the first half. So the ranker writes an impression record as it responds: a **request id** minted at serve time, the ordered slots with track id and position, which slots the client actually rendered, the score each slot got, the selection propensity if any randomisation was applied, and the model and feature-definition versions in force. The client then echoes the request id on every play, skip and save. The label job joins on `(requestId, trackId)` inside the attribution window - not on user id and a nearby timestamp, which mis-attributes whenever a listener has two surfaces open. Rebuilding the ranking later by re-scoring the request is not a substitute: features have moved and the model may have been promoted, so you would label a ranking that was never shown.

code

json · 13 lines
json
{
  "requestId": "req-8f21c0a3",
  "listenerId": "lst-4411",
  "surface": "generated-playlist",
  "servedAt": "2026-09-18T09:14:02.331Z",
  "modelVersion": "ranker-2026-09-11",
  "featureSpecVersion": "3.4.0",
  "slots": [
    { "position": 1, "trackId": "trk-9931", "score": 0.81, "propensity": 0.92, "rendered": true },
    { "position": 2, "trackId": "trk-1027", "score": 0.77, "propensity": 0.90, "rendered": true },
    { "position": 3, "trackId": "trk-5560", "score": 0.74, "propensity": 0.88, "rendered": false }
  ]
}

go deeper

for a junior

Know that the training row needs both halves - what was shown and what happened - and that only the serving response knows the first half. An id minted at serve time is what ties them together.

for a middle

List the fields and say why each one exists: request id for the join, position and rendered flag for what was really presented, propensity for randomisation, model and feature versions for attribution.

for a senior

Demonstrate the operational half: a measured join rate sliced by client version, unjoined actions kept in a side stream with a reason, and the impression schema treated as a contract shared with client teams.

for a principal

The trade is volume against fidelity. Impressions dwarf actions, so someone must decide what to keep, at what sampling rate, and for how long - and record the rate, because the exposure-to-action ratio is what the next ranker learns.

## The join you are trying to make hours later A training row for a ranker is a pair: the item as it was presented, and what the listener then did. The second half arrives minutes or days later, from a different system, often from a device that was offline in between. The first half exists for a few milliseconds inside a serving response and then is gone. Everything in this question follows from that asymmetry - **the serving path is the only place some of these facts ever exist**, so it either writes them down or they are lost. The artifact that carries them is the impression record: one per serving request, written on or derived directly from the response, not reconstructed afterwards. ## What only the serving path knows - **A request id**, minted when the ranking is produced and returned to the client, which the client echoes on every subsequent event for that surface. - **The ordered slots**: track id and position, exactly as sent. - **Which slots were rendered or entered the viewport**, reported back by the client. Without this the label job cannot tell an ignored track from one nobody ever reached. - **The score** the ranker assigned each slot, which pins the ranking to a model's actual behaviour at that moment. - **The selection propensity** wherever the service randomised anything - the probability with which that item landed in that slot. It is a serve-time quantity and is unrecoverable once the response is gone. - **The model version and the feature-definition version** in force for this request. - **The serve timestamp in event time**, which anchors the attribution window. ## Why re-scoring the request later is not a substitute A tempting shortcut is to store only the request inputs and rebuild the ranking on demand. It does not work, for three independent reasons: 1. **The features have moved.** The online store has been overwritten many times since; recomputing gives the listener's state now, not at serve time. 2. **The model may have changed.** A refresh may have been promoted between the request and the replay, so the replayed ranking comes from a model that never served that listener. 3. **Randomisation is not repeatable.** Any exploration or tie-breaking that fed the response cannot be reproduced after the fact. The result would be a plausible ranking that was never on anyone's screen, joined to actions that were caused by a different one. That is worse than missing data, because it is silently wrong. ## What each missing field costs | field left out | what breaks in the label job | |---|---| | request id | joins fall back to user and timestamp proximity, mis-attributing whenever two surfaces are open | | slot position | you cannot separate a track shown deep in the list from one shown at the top and ignored | | rendered flag | slots the listener never reached enter the training set as negatives | | model version | a change in label quality can never be attributed to a specific promotion | | feature-spec version | the training job may read a definition that differs from the one serving wrote | | propensity | any serve-time randomisation becomes uninterpretable, permanently | ## Making the join robust in operation The join itself is mechanical; the failure modes are operational, and they are silent by default. Three habits catch nearly all of them: 1. **Count the unjoined.** Every play that finds no impression is emitted to a side stream with a reason, never dropped. A join rate that is not measured is a join rate that is falling. 2. **Slice the join rate by client version and platform.** A client release that stops echoing the request id shows up here first, as a clean step change in one slice while the totals barely move. 3. **Treat the impression record as a contract.** Its schema is owned jointly with the client teams, and a change to the event that carries the request id is a breaking change to the training set even though no model code was touched. One structural point is worth stating plainly: the impression stream is usually far larger than the action stream, since most slots are never acted on. Some services keep every impression, others down-sample unacted impressions and record the sampling rate so the label job can weight them back. Either is defensible; what is not defensible is sampling silently, because the ratio between exposures and actions is exactly what the next ranker is being taught.

  • Why can the impression record not be rebuilt by re-running the ranker over yesterday's requests?
    Because three things have moved since: the online features the request read, the model version serving at that moment, and any randomisation that shaped the response. A replay produces a ranking that was never shown, and joining real actions to it manufactures training rows that describe a system that never existed.
  • One client release stops echoing the request id. What does the label job see?
    Plays from that client fail the join and vanish from the positives, while its impressions keep arriving, so that population looks almost entirely negative. Totals move only slightly, which is why the guard is a join rate sliced by client version rather than a global count.

saying these in an interview costs you the question

  • Rebuilds the ranking later by re-scoring with today's model.
  • Joins plays to rankings using user id and timestamp proximity.
  • Writes an impression record only for tracks that were played.
  • Assumes the client rendered every slot the service returned.
  • Leaves the serving model version out of the impression record.