In a music playlist ranking service, why is a completed play not proof the listener wanted that track?
answer
- proxy signal, not ground truth
- the service chose what was shown
- silence has several causes
- skip, complete, save disagree
- the label spec defines the truth
basics
~20 sA completed play records what the listener tolerated in a context the service chose: autoplay, background listening and the slot given to the track all produce plays. It is a proxy label, not a verdict.
solid answer
~50 sA generated playlist chooses the tracks, their order and when playback starts, so every action is conditioned on that choice. A play to completion can mean the listener loved the track, left the room, or never reached for their phone; an early skip can mean rejection or wrong moment. Silence is worse: a track with no interaction may have been ignored, or may have sat in a slot the listener never reached. The four signals a playlist service can see - an early skip, a completed play, a save, a down-vote - disagree with each other, and the ones closest to real intent are the rarest. The training label is therefore not discovered in the logs, it is **defined**: a written label spec names which event at which threshold counts positive, what counts negative, and what stays unknown. Whatever that spec says is what the next ranker learns to maximise.
go deeper
Be able to say that implicit actions are proxies: a completed play is not a verified like, and a track with no play is not a verified dislike. Name two ways a play happens without intent.
Explain how the four playlist signals disagree, and why intent and volume trade against each other - saves are the clearest and the scarcest. Describe what a written label spec pins down.
Show that you treat the label definition as an owned, versioned artifact, tested against the client's event schema and chosen for the outcome the business names rather than for whichever event is easiest to collect.
The angle is what the chosen label turns the product into. Optimising completed plays buys inoffensive background listening; optimising saves buys a sparse, power-user signal. Say which one the service is willing to pay for.
## The action is evidence about the system, not only about the listener In a generated playlist the service picked the tracks, picked their order, and in most clients started playback without being asked again. Every action that follows is conditioned on all three. A play to completion is a joint outcome of the track, the slot it occupied, the listening context - a commute, a focus session, a party - and whether anyone was paying attention at all. Treating that event as a positive training label is a decision you are making on the listener's behalf, and it is the most consequential decision in the whole retraining loop, because the ranker's objective, the offline metric and the reported win all inherit whatever the definition encodes. Three properties separate an implicit action from a rating someone typed in: - **It is conditioned on exposure.** You only ever learn about tracks the ranker chose to show. Nothing in the logs says anything about the rest of the catalogue. - **It is ambiguous.** The same event is produced by enthusiasm, indifference and absence, and the log cannot tell them apart. - **It is defined by instrumentation.** A completed play is whatever the client decided to emit - at 30 seconds, at 90 percent of duration, or at the track boundary. Change the client and you change the label without touching the model. ## The four signals, and how each one misleads | signal | what it suggests | how it misleads | |---|---|---| | skip in the first seconds | rejection of this track, here | fires when the listener is sampling the playlist, or restarting a track | | play to completion | tolerance | autoplay, background listening, a device in a pocket | | save to library | deliberate intent | rare, and concentrated in a small set of engaged listeners | | down-vote or hide | explicit rejection | rarest of all, and used by an unusual minority | The pattern is consistent: the closer a signal sits to genuine intent, the less of it you have. A label built on saves is clean and tiny and describes power users; a label built on completed plays is plentiful and describes what nobody bothered to stop. ## Absence is the harder half A track that produced no action at all sits in one of at least three states: 1. It was rendered on screen and the listener passed over it. 2. It was returned by the ranker but sat below the last slot the listener ever reached. 3. It was rendered and started, and the session ended for a reason that had nothing to do with the track. The label job cannot distinguish these from the play stream alone. If it labels all three as negative, the ranker learns that the tail of every playlist is bad - which is mostly a statement about how far people scroll, not about the music. Keeping the third state as **unknown** rather than negative costs you training rows and buys you a label that means what it says. ## Writing a label spec you can defend 1. Name the outcome the product actually wants: repeat listening, library growth, session length. 2. Choose the action closest to it and write the exact event and threshold - not 'a play' but 'playback progressed past the stated fraction of track duration without a skip'. 3. State separately what counts as negative and what is left unknown, including the unreached-slot case. 4. Version the spec, and stamp the version on the training snapshot the job produces. 5. Keep the other signals as their own columns rather than folding them into one label, so a later change of objective does not require re-deriving history. ## What it costs to skip this With no written spec the definition becomes whatever the client happens to emit, which means it changes on the client's release schedule rather than yours. Two teams then compute different positive rates from the same logs and spend a week reconciling them. Worse, the unstated choice quietly steers the product: a ranker trained to maximise completed plays learns to prefer inoffensive, low-variance background music, because that is what survives inattention. Nothing in the metric will tell you this is happening - the completion rate goes up, exactly as designed. The label was the lever, and it was pulled by default.
- Which non-actions would you record as a negative label, and which as unknown?A track rendered on screen and passed over is a usable weak negative. A track returned by the ranker but below the last slot the listener ever reached is unknown, because nothing was ever presented to reject. A session that ended mid-track is unknown too. The distinction only exists if the impression record says which slots were actually rendered.
- The skip, completed-play and save signals disagree on one track. How do you settle on a training label?You do not average them. Pick the action closest to the outcome the product is optimising, write it into the label spec with its exact threshold, and keep the other signals as separate columns or auxiliary targets. Disagreement is information about intent strength, and collapsing it early throws that away permanently.
A kitchen that counts every cleared plate as a five-star review will drift toward whatever is hardest to leave uneaten. The plate records tolerance; it was never asked about preference.
saying these in an interview costs you the question
- Treats a completed play as a verified preference for the track.
- Assumes a track with no interaction was rejected by the listener.
- Believes more logged events make the label less biased.
- Says a save and a completed play carry the same intent.
- Waits for the model to average out label noise it could remove at the source.