skip to content

How do implicit feedback signals like plays and skips differ from explicit star ratings?

level: juniorimportance: must knowfreq 78%

answer

  1. stated opinion versus observed behaviour
  2. abundant, cheap, always arriving
  3. no dislike is ever recorded
  4. counts mean confidence, not degree
  5. autoplay and shared accounts pollute it

basics

~20 s

Explicit ratings are stated preferences on a scale and can express dislike. Implicit signals only record actions, so they are abundant and cheap but noisy, one-class, and count-valued: no interaction is not a stated dislike.

solid answer

~50 s

An explicit rating is a deliberate judgement on a scale, so it carries both direction and degree, and a one-star means the user actively disliked the item. An implicit signal is a side effect of using the product — a play, a skip, a save, time spent — so it only ever records that something happened. Three consequences matter. It is **one-class**: positives are observed, negatives are never stated. It is **count-valued**: eight replays of a podcast episode signal stronger confidence than one open, not eight times the preference. And it is **noisy**: autoplay, a shared account, a mis-tap and background listening all produce interactions nobody chose. In exchange you get orders of magnitude more data, from every user rather than the small minority who bother to rate, and it keeps arriving without asking anyone anything.

go deeper

for a junior

Be ready to name concrete implicit signals — plays, skips, saves, dwell, purchases — and to say the one thing that matters: nothing in the log records a dislike, so absence of an interaction is ambiguous.

for a middle

Explain the mechanics: preference is binary, the count is confidence, and unobserved cells have to be handled deliberately rather than dropped. Be able to say why squared error on counts is the wrong objective here.

for a senior

Show you know where the noise comes from in a real product — autoplay, shared accounts, background sessions — and how you would clean or downweight it before it becomes training labels.

for a principal

Own the tradeoff at the product level: whether to spend product surface asking for explicit signal at all, given that it reaches a small biased slice of users, versus investing in better instrumentation of behaviour everyone already produces.

## Two kinds of feedback A recommender learns from a user-item interaction matrix. What fills the cells decides how you may model it. **Explicit feedback** is a rating the user was asked for and chose to give: a star score, a thumbs up or down, a numeric review. It has three properties that make modelling easy. It is *signed* — a one-star is real evidence of dislike, not just an absence. It is *graded* — four stars is more than three by a stated amount. And it is *intentional* — the user meant to express the value that got recorded. **Implicit feedback** is the log of what people did: a track played, a track skipped, a save, a share, seconds of dwell, a purchase. Consider a music-streaming product with no star ratings anywhere: every training signal has to be squeezed out of plays, skips and saves. ## What implicit data gives you *Volume.* Almost nobody rates; everybody uses. Implicit logs cover the whole active user base, not the small self-selected slice that fills in stars, so the sample is far less biased toward opinionated users. *Freshness.* Behaviour arrives continuously and reflects current taste, whereas ratings age. *No user cost.* Collecting it asks nothing of the person. ## What implicit data takes away **It is one-class.** The log contains positives only. Nothing in it says "this listener dislikes this track". A cell with no play is unobserved, and unobserved mixes two completely different situations: the listener was never shown the track, or the listener saw it and passed. That ambiguity is the single defining problem of implicit feedback and it dictates how negatives have to be manufactured during training. **Counts are confidence, not degree.** In explicit data the value *is* the preference. In implicit data the value is how much evidence you have. A podcast episode replayed eight times gives you high confidence the listener likes it; an episode opened once gives you weak evidence of the same binary preference. Modelling a count as if it were a rating on a stretched scale is a classic mistake — it makes a track someone loops while cleaning outrank a track they adore but hear once a week. **It is noisy in a way ratings are not.** Autoplay keeps playing after the listener leaves the room. A household shares one account. A mis-tap opens something nobody wanted. A track finishes because the user was driving and could not reach the phone. Every one of these lands in the log as a positive. Explicit ratings have their own noise, but at least someone deliberately produced each value. **Negative-looking events are weak.** A skip is tempting to read as a dislike, and it does carry some signal, but people skip because they already know the track, because they are in the wrong mood, because they are hunting for one specific song. A skip after two seconds means something different from a skip at 80% through. Treating skips as hard negatives will suppress items the listener actually loves. ## How this changes the modelling The usual formulation splits the signal in two. A binary preference — did any interaction happen at all — and a confidence attached to it, derived from the interaction count. Unobserved cells are not dropped and not labelled dislikes; they are handled either as very low-confidence negatives across the whole matrix, or by sampling a manageable number of them as negatives per observed positive. Evaluation changes too. There is no held-out rating to predict, so squared error on ratings is meaningless. What you evaluate is ranking: given the items this user interacted with in a later window, how high does the model place them among everything else. ## The interview framing Say explicitly: implicit feedback trades *precision of the signal* for *coverage of the population*. You get vastly more data about vastly more users, at the cost of never being told what anyone disliked, and of every count being contaminated by mechanics of the product rather than taste. The rest of implicit-feedback modelling is engineering around those two costs.

  • Is a skip a reliable negative signal?
    Only weakly. People skip tracks they already know, tracks that fit the wrong mood, or anything while hunting for one specific song. Skip position helps — a skip two seconds in is far more informative than one at 80% through — but even an early skip is soft evidence, so it belongs as a low-confidence negative or a separate feature, never as a stated dislike on par with a one-star rating.
  • Why do teams usually prefer implicit data despite it being noisier?
    Coverage. Only a small, self-selected minority ever rates anything, so an explicit matrix is both tiny and biased toward opinionated users and polarising items. Implicit logs cover every active user and every session, and they refresh continuously. Noise that is roughly random across millions of interactions hurts less than a systematically unrepresentative sample of a few thousand ratings.
  • Can you have both signal types in one product?
    Yes, and it is common: plays and saves flow constantly while a small set of explicit thumbs arrives from engaged users. The usual treatment is to keep them as separate signals with different weights rather than mashing them into one scale — an explicit thumbs-down is the only trustworthy negative you have, so it is worth far more per event than any play count.

An explicit rating is a customer filling in a survey. An implicit signal is a security camera recording which aisles they walked down — far more footage, but nobody ever said what they thought.

saying these in an interview costs you the question

  • Maps play counts onto a one-to-five rating scale
  • Calls a skip a confirmed dislike
  • Says no interaction means the user rejected the item
  • Describes implicit data as simply sparser explicit data
  • Ignores autoplay, shared accounts and mis-taps as noise sources

context