skip to content

questions

5

In a music playlist ranking service, why is a completed play not proof the listener wanted that track?

level: juniorimportance: must knowfreq 72%

answer

  1. proxy signal, not ground truth
  2. the service chose what was shown
  3. silence has several causes
  4. skip, complete, save disagree
  5. the label spec defines the truth

basics

~20 s

A completed play records what the listener tolerated in a context the service chose: autoplay, background listening and the slot given to the track all produce plays. It is a proxy label, not a verdict.

solid answer

~50 s

A generated playlist chooses the tracks, their order and when playback starts, so every action is conditioned on that choice. A play to completion can mean the listener loved the track, left the room, or never reached for their phone; an early skip can mean rejection or wrong moment. Silence is worse: a track with no interaction may have been ignored, or may have sat in a slot the listener never reached. The four signals a playlist service can see - an early skip, a completed play, a save, a down-vote - disagree with each other, and the ones closest to real intent are the rarest. The training label is therefore not discovered in the logs, it is **defined**: a written label spec names which event at which threshold counts positive, what counts negative, and what stays unknown. Whatever that spec says is what the next ranker learns to maximise.

go deeper

for a junior

Be able to say that implicit actions are proxies: a completed play is not a verified like, and a track with no play is not a verified dislike. Name two ways a play happens without intent.

for a middle

Explain how the four playlist signals disagree, and why intent and volume trade against each other - saves are the clearest and the scarcest. Describe what a written label spec pins down.

for a senior

Show that you treat the label definition as an owned, versioned artifact, tested against the client's event schema and chosen for the outcome the business names rather than for whichever event is easiest to collect.

for a principal

The angle is what the chosen label turns the product into. Optimising completed plays buys inoffensive background listening; optimising saves buys a sparse, power-user signal. Say which one the service is willing to pay for.

## The action is evidence about the system, not only about the listener In a generated playlist the service picked the tracks, picked their order, and in most clients started playback without being asked again. Every action that follows is conditioned on all three. A play to completion is a joint outcome of the track, the slot it occupied, the listening context - a commute, a focus session, a party - and whether anyone was paying attention at all. Treating that event as a positive training label is a decision you are making on the listener's behalf, and it is the most consequential decision in the whole retraining loop, because the ranker's objective, the offline metric and the reported win all inherit whatever the definition encodes. Three properties separate an implicit action from a rating someone typed in: - **It is conditioned on exposure.** You only ever learn about tracks the ranker chose to show. Nothing in the logs says anything about the rest of the catalogue. - **It is ambiguous.** The same event is produced by enthusiasm, indifference and absence, and the log cannot tell them apart. - **It is defined by instrumentation.** A completed play is whatever the client decided to emit - at 30 seconds, at 90 percent of duration, or at the track boundary. Change the client and you change the label without touching the model. ## The four signals, and how each one misleads | signal | what it suggests | how it misleads | |---|---|---| | skip in the first seconds | rejection of this track, here | fires when the listener is sampling the playlist, or restarting a track | | play to completion | tolerance | autoplay, background listening, a device in a pocket | | save to library | deliberate intent | rare, and concentrated in a small set of engaged listeners | | down-vote or hide | explicit rejection | rarest of all, and used by an unusual minority | The pattern is consistent: the closer a signal sits to genuine intent, the less of it you have. A label built on saves is clean and tiny and describes power users; a label built on completed plays is plentiful and describes what nobody bothered to stop. ## Absence is the harder half A track that produced no action at all sits in one of at least three states: 1. It was rendered on screen and the listener passed over it. 2. It was returned by the ranker but sat below the last slot the listener ever reached. 3. It was rendered and started, and the session ended for a reason that had nothing to do with the track. The label job cannot distinguish these from the play stream alone. If it labels all three as negative, the ranker learns that the tail of every playlist is bad - which is mostly a statement about how far people scroll, not about the music. Keeping the third state as **unknown** rather than negative costs you training rows and buys you a label that means what it says. ## Writing a label spec you can defend 1. Name the outcome the product actually wants: repeat listening, library growth, session length. 2. Choose the action closest to it and write the exact event and threshold - not 'a play' but 'playback progressed past the stated fraction of track duration without a skip'. 3. State separately what counts as negative and what is left unknown, including the unreached-slot case. 4. Version the spec, and stamp the version on the training snapshot the job produces. 5. Keep the other signals as their own columns rather than folding them into one label, so a later change of objective does not require re-deriving history. ## What it costs to skip this With no written spec the definition becomes whatever the client happens to emit, which means it changes on the client's release schedule rather than yours. Two teams then compute different positive rates from the same logs and spend a week reconciling them. Worse, the unstated choice quietly steers the product: a ranker trained to maximise completed plays learns to prefer inoffensive, low-variance background music, because that is what survives inattention. Nothing in the metric will tell you this is happening - the completion rate goes up, exactly as designed. The label was the lever, and it was pulled by default.

  • Which non-actions would you record as a negative label, and which as unknown?
    A track rendered on screen and passed over is a usable weak negative. A track returned by the ranker but below the last slot the listener ever reached is unknown, because nothing was ever presented to reject. A session that ended mid-track is unknown too. The distinction only exists if the impression record says which slots were actually rendered.
  • The skip, completed-play and save signals disagree on one track. How do you settle on a training label?
    You do not average them. Pick the action closest to the outcome the product is optimising, write it into the label spec with its exact threshold, and keep the other signals as separate columns or auxiliary targets. Disagreement is information about intent strength, and collapsing it early throws that away permanently.

A kitchen that counts every cleared plate as a five-star review will drift toward whatever is hardest to leave uneaten. The plate records tolerance; it was never asked about preference.

saying these in an interview costs you the question

  • Treats a completed play as a verified preference for the track.
  • Assumes a track with no interaction was rejected by the listener.
  • Believes more logged events make the label less biased.
  • Says a save and a completed play carry the same intent.
  • Waits for the model to average out label noise it could remove at the source.
open as a page

What must a playlist ranking service log at serve time so a later play can be joined to the ranking that produced it?

level: middleimportance: must knowfreq 62%

basics

~20 s

One impression record per serving request: a request id the client echoes on every later event, each slot's track id, position and rendered flag, the serve-time score and any randomisation propensity, plus the model and feature-spec versions.

open as a page

In a playlist ranker's label job, what does a short attribution window on saves quietly drop?

level: middleimportance: should knowfreq 48%

basics

~20 s

It drops the slow saves - the track someone returns to that evening - and the loss is not random: it removes long playlists, background listening and deliberate second thoughts, leaving a positive class made of instant reactions.

open as a page

How should a playlist ranker's training set handle actions that arrive after the training snapshot was cut?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Close labels only for impressions whose attribution window ended before the cut; anything younger is unmatured and must be excluded, never written as a negative. Corrections that arrive later go into a new snapshot version rather than editing the old one.

open as a page

In a playlist ranker's training set, the positive rate halves overnight - how do you find the cause?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Read the shape first: a step change aligned to a release is instrumentation, not behaviour. Then walk the three places a positive dies - the event was not emitted, the join failed, or the label rule rejected it - slicing every count by client version.

open as a page