A newly published podcast has no listens and so no behavioural vector — what makes it retrievable by the funnel?
answer
- no interactions, no row
- features exist before behaviour does
- same space or a separate source
- the merge needs a reserved quota
- missing is not the same as zero
basics
~20 sA representation built from what exists at publish time — declared category, description text, host and network links, transcript topics. It is either projected into the retrieval space alongside behavioural items or run as a separate candidate source with its own quota in the merge.
solid answer
~50 sThe behavioural retrieval space is built from interactions, so an item with zero interactions has no row in it and cannot be returned by any nearest-neighbour lookup, however good the ranker downstream. The funnel needs a second path. One arrangement projects publish-time content features into the same vector space, so the cold item is retrievable by the existing lookup and is replaced by its behavioural representation once interactions accumulate. The other runs content retrieval as a separate source whose results are merged into the candidate set. Either way the merge needs an explicit quota for the content source: merged on raw score against behavioural candidates, cold items lose and never reach the shortlist. And whatever the scoring stage does with an item whose engagement features are null has to be a stated choice, not an accident.
go deeper
Recall that the scoring stage only ever sees candidates retrieval handed it, so an item missing from every retrieval source cannot appear in a feed however good it is.
Explain where a cold item's representation comes from — publish-time metadata, relations, transcript topics — and the two ways that path is wired into the funnel.
Show the operational detail: the reserved quota in the merge, missing engagement features marked as missing rather than zero, and a measured contribution rate for the content source.
Weigh running one shared retrieval space against a separate source: one is simpler to reason about, the other is attributable, disableable and independently tunable when something goes wrong.
## Why the behavioural index has no row for it Retrieval in a recommendation funnel is a lookup over a space learned from interactions: co-listening, completion, subscription. A show published an hour ago contributed nothing to that space, so there is no vector to look up and no neighbour list that contains it. This is a **serving** fact before it is a modelling one — no amount of ranking quality downstream helps an item that the candidate stage never emits, because the scoring stage only ever sees the shortlist retrieval hands it. So the item side of cold start is answered at the retrieval stage, with material that exists the moment the episode publishes: - the declared category and any tags the publisher supplied; - title and description text; - relations — the show it belongs to, the host, the network, a series it continues; - topics extracted from the transcript, which for a podcast is unusually rich and available within minutes of publication; - the publish timestamp itself, which separates "new" from "old and ignored". ## Two arrangements that make a cold item retrievable | arrangement | what it needs | what it buys | what it costs | |---|---|---|---| | project content features into the same retrieval space | a mapping from publish-time features into the vector space the behavioural items live in | one lookup path, one shortlist, no merge policy to tune | the two representations must be comparable, and you need a rule for when an item swaps to its behavioural vector | | run content retrieval as a separate candidate source | a second index over content features, plus a merge step | independently tunable, easy to disable, easy to attribute in logs | the merge has to reserve room for it, or its candidates are outscored and vanish | Both are common; platforms genuinely differ on which they run, and a design answer that names the trade-off rather than a winner is the stronger one. The separate-source arrangement is easier to operate — you can see exactly how many candidates it contributed and turn it off without touching the main lookup — while the shared-space arrangement avoids a merge policy altogether. ## The quota that keeps the merge honest When candidates from several sources are merged into one shortlist, they are ordered by some comparable quantity — a retrieval score, a fused rank. A cold item's content score is not systematically higher than a warm item's behavioural score, and there are vastly more warm items. Merged on score alone, content candidates are squeezed out of a shortlist of a few hundred, and the whole cold path quietly contributes nothing while looking healthy in its own metrics. The fix is structural rather than statistical: **reserve a fixed number of shortlist positions for the content source**, so the merge always carries some cold candidates forward. Whether those candidates then win a place on the final slate is the next stage's business; the quota only guarantees they are considered. ## What the scoring stage sees A cold item arrives at the second stage with its engagement features — play rate, completion rate, subscriber count — all null or zero. Zero is the dangerous encoding, because a zero completion rate is a factual claim that the item performs terribly, and the model will duly rank it last. Two defensible arrangements: 1. Mark the absence explicitly, so "unknown" is distinguishable from "bad", and let the model have been trained with that marker present. 2. Score cold candidates with a reduced feature set, or with a cheap prior over content similarity, and let the slate-assembly stage decide how many of them it wants. ## What to measure - The content source's **contribution rate**: candidates it supplied that survived into the shortlist, per request. A rate near zero means the quota is missing or too small. - The share of the catalogue that is retrievable at all — items with no representation in any source are invisible no matter what the rest of the funnel does. - The gap between when an episode publishes and when the funnel can first return it as a candidate. ## Failure modes - Treating "no interactions" as "no features", when publish-time metadata and a transcript are a substantial description of the item. - Merging the content source without a quota, so it is present in the architecture diagram and absent from every shortlist. - Encoding missing engagement as zero, which tells the scorer the item is proven bad rather than unproven. - Assuming the item becomes retrievable on its own once someone plays it — nobody can play what is never shown.
- When should a cold item stop using its content representation?When it has enough interactions for a behavioural representation to be more informative than the content one — a stated minimum, not a calendar date. Run the swap as an explicit transition with a logged before-and-after, because it changes which neighbours the item is retrieved beside, and a show whose content description and real audience disagree will visibly move. Some arrangements blend the two across the transition rather than switching outright.
- A new episode of an established, popular show is also cold — is it the same problem?Weaker, and worth exploiting. The episode has no interactions but inherits strong relations: the show's audience, its subscribers, its historical completion rate. A useful retrieval path uses show-level behaviour as the item's prior and reserves true cold-start handling for a show with no history at all. Conflating the two spends scarce exploration on items that already have a defensible audience estimate.
saying these in an interview costs you the question
- Thinks a new item is retrievable at publish with no content path at all
- Treats zero interactions as zero features to work with
- Merges the content source with no quota, so its candidates never survive
- Encodes an unknown completion rate as zero rather than as missing
- Expects the ranker to rescue an item the retrieval stage never emitted