skip to content

In a playlist ranker's training set, the positive rate halves overnight - how do you find the cause?

level: seniorimportance: nice to knowfreq 30%

answer

  1. a step change is not behaviour
  2. slice by client version first
  3. numerator or denominator
  4. watch the join rate, not just counts
  5. late is not the same as missing

basics

~20 s

Read the shape first: a step change aligned to a release is instrumentation, not behaviour. Then walk the three places a positive dies - the event was not emitted, the join failed, or the label rule rejected it - slicing every count by client version.

solid answer

~40 s

A training positive is the product of three things: the client emitted an event, the event joined to an impression, and the label rule accepted it. Halving means one of the three broke, and the shape says which kind. Behaviour moves gradually; a 50% step overnight, aligned to a client release boundary, is almost always instrumentation - a renamed event, a new completion threshold, background playback no longer reported. Check the **denominator** too: an impression flood from a newly instrumented surface halves the rate with positives untouched. Then check the **join rate** by client version, because a client that stops echoing the request id loses its plays silently while its impressions keep flowing. Finally rule out lag: if the stream is late rather than missing, yesterday backfills today and the drop heals itself.

go deeper

for a junior

Know that the number of positives in a training set depends on what the client emits, so an abrupt change there usually means the logging changed rather than the listeners.

for a middle

Walk the three places a positive dies - emission, join, label rule - and check the denominator as well as the numerator, since an exposure flood halves the rate with positives untouched.

for a senior

Distinguish late from missing using event time against arrival time, localise the change by slicing on client version, and say what you would do with the affected training window once the cause is known.

for a principal

The standing question is who owns the label contract. The definition lives in someone else's release cycle, and deciding how that dependency is governed and tested matters more than any one incident.

## Read the shape before you look for a cause The first useful fact is not in the logs, it is in the curve. Listener behaviour changes over days and weeks and never aligns exactly to midnight. A clean 50% step, starting at a deploy boundary and holding flat afterwards, is the signature of a definition change somewhere in the pipeline. Treating it as a taste shift and retraining on it is the worst available response, because it teaches the next ranker a measurement artefact and then hides the artefact behind a new baseline. ## The three places a positive can die Walk them in order; each has a cheap check. 1. **The event was never emitted.** The client changed what it reports: a completion event now fires at a different fraction of the track, background playback stopped being reported, a save moved to a new event name. Check: raw event counts by type, sliced by client version and platform. 2. **The event was emitted but did not join.** The request id stopped being echoed, its format changed, or the impression stream is late or partially missing. Check: join rate - matched actions over total actions - sliced the same way, plus the volume of the unattributed side stream and its reasons. 3. **The event joined but the label rule rejected it.** A window was shortened, an action type was dropped from the positive definition, or a maturity horizon moved. Check: the label spec's version history, and the count of matched actions falling outside the window. ## Numerator or denominator A rate has two halves and only one of them is usually examined. Both produce the same headline: | what actually moved | positives | impressions | what it means | |---|---|---|---| | positives halved | down 50% | flat | emission, join or rule broke | | impressions doubled | flat | up 100% | a new surface began logging exposures | | both moved | down | up | a release changed emission on both streams | The second row is the one that gets missed. A newly instrumented surface, or a client that starts reporting slots it previously never rendered, inflates the exposure stream and halves the positive rate while nothing at all happened to listeners. It also changes what the next ranker learns, because the exposure-to-action ratio is part of the training signal. ## The case that heals itself Before concluding anything, separate **missing** from **late**. If the drop is concentrated in the most recent hours and the affected window keeps filling in as the day goes on, the pipeline is lagging, not losing. Comparing counts by event time against counts by arrival time separates the two immediately, and a lag alarm on the action stream makes the distinction automatic rather than a judgment call each time. ## Once you know, decide what to do with the affected window Suppose the cause is real: a client release changed what a completed play means, and the new semantics are staying. There is no way to average two definitions into one training set, so you choose: - **Re-derive the old definition** from finer-grained events, if the client still emits enough to reconstruct it. This is the best outcome and is often possible when the underlying progress events survived the change. - **Cut the training window at the release boundary** and accept less history, which is honest and costs volume. - **Keep both and carry the label-definition version as a feature of the row**, so a model can at least tell the regimes apart. This is a last resort, because it hard-codes an instrumentation accident into the model. Whichever you pick, the label-definition version belongs on the snapshot, so that the discontinuity is discoverable next year rather than being rediscovered by whoever sees the next step. ## The two guards that would have caught this same day Both are cheap and both belong to whoever owns the label job: - a **join rate and positive rate sliced by client version**, alarmed on a step rather than on a level, which localises the change to a release within minutes; - a **schema contract test** on the events the label job depends on, run against the client's payloads, so a renamed or re-thresholded event fails before it reaches the training set rather than after. The underlying lesson is the one the leaf keeps returning to: the label means whatever the logging says it means, and the logging is owned by somebody else.

  • The change is real and the new client semantics are staying. What do you do with the training set?
    Never blend two definitions in one window. Re-derive the old definition from finer-grained events if they survived, otherwise cut the training window at the release boundary and accept less history. Either way, stamp the label-definition version on the snapshot so the discontinuity stays discoverable.
  • What guard would have caught this the same day?
    A positive rate and join rate broken out by client version, alarmed on a step rather than a level, would have localised it to one release within minutes. Pair it with a schema contract test on the events the label job consumes, so a renamed or re-thresholded event fails before it reaches a training set.

saying these in an interview costs you the question

  • Retrains immediately on the new data to let the model adapt.
  • Blames a shift in listener taste for an overnight step change.
  • Checks the positive count and never the impression count.
  • Assumes the event schema cannot change without the model team hearing.
  • Reads an unjoined play as a play that never happened.