A feed retrained monthly only on its own logged impressions narrows users from twelve interest categories to three — why?
answer
- the model chooses its own training data
- no impression means no label at all
- missing not at random, not missing at random
- offline score rewards agreeing with the incumbent
- track interest entropy across months
basics
~20 sThe training log only contains items the previous model chose to show. Categories it stopped surfacing collect no engagement, so the next model sees even less evidence for them and shows them less again. The loop compounds every retrain.
solid answer
~50 sThis is a closed feedback loop with exposure bias. Engagement labels exist only for items that were served, and what was served was decided by the incumbent model — so the training set is a sample the model itself selected, not a sample of user interest. A category the model under-ranks gets few impressions, therefore few clicks, therefore looks even weaker in the next month's data, therefore gets ranked lower still. At the user level that ratchet narrows twelve interests to three; at the item level it starves most of the catalog of exposure. The trap is that offline metrics keep improving, because they are computed on the same logged data: a model that faithfully reproduces the incumbent's choices scores well against a log the incumbent generated. You detect it with longitudinal instruments — per-user interest entropy across cohorts, catalog coverage and impression concentration over successive retrains — not with the top-k score of a single model.
go deeper
Remember the core fact: the log only records what the system chose to show, so an item with no clicks may simply never have appeared. Be able to say why that makes the next model narrower.
Explain the ratchet step by step over successive retrains, and state precisely why the labels are missing not at random rather than merely sparse.
Show the diagnosis: which longitudinal instruments you would put in place, over what window, and why a single model's offline score cannot reveal the problem because it is scored on data the incumbent produced.
Own the argument that beyond-accuracy metrics belong in the release gate rather than a monthly report, and be ready to defend the cost of maintaining a measurement population the production model did not shape.
## What a closed loop is A recommender does not observe the world; it observes the consequences of its own actions. For each request it chooses a slate, logs the impressions, and logs whatever the user did with them. When the next model is trained on that log, every training row descends from a decision the previous model made. The system is training on data it generated. That circularity is the closed feedback loop, and the specific distortion it produces is **exposure bias**: outcomes are observed only for items that were exposed, and exposure was not random — it was the incumbent's ranking. Say it statistically: the missing labels are missing *not* at random. An item with no clicks may be uninteresting, or it may simply never have been shown. Nothing in the log distinguishes those two cases, and a model fit to the log will happily read the second as the first. ## The ratchet, month by month Follow one user with twelve genuine interest categories on a feed retrained monthly. - **Month 0.** The model is uncertain, spreads impressions over most categories, and gets engagement signal on all of them. - **Month 1.** Three categories happened to convert best — partly real preference, partly noise, partly which items were available that month. They get more slots. The other nine get fewer. - **Month 2.** Nine categories now have thin, stale evidence. Not "negative evidence" — *no* evidence. Their estimated scores drift down or stay flat while the top three accumulate fresh positives. Slot share shifts further. - **Month 3 onward.** The nine are effectively unreachable. The user cannot express interest in something they are never shown, and the log records that silence as disinterest. The same ratchet runs on the item side. An item that is under-ranked receives no impressions, accumulates no engagement statistics, and stays under-ranked — the effective catalog shrinks toward whatever the first few models happened to favour, and new items face a cold start that the loop keeps cold. Two things make this worse than it sounds. First, the narrowing is *self-confirming*: each cycle's data genuinely supports the narrowed model, so nothing in training looks wrong. Second, it is slow. Over one retrain the shift is within noise; over a quarter or a year it is a different product. ## Why the metrics do not scream The reason teams miss this for months is that the offline evaluation is computed on the same biased log. If the held-out data consists of interactions with items the incumbent chose to show, then a candidate model is being scored on how well it agrees with the incumbent's choices. A model that reproduces the incumbent's narrow behaviour scores well; a model that would have surfaced one of the nine abandoned categories is penalised, because the log contains no interaction with those items to credit it for. The evaluation inherits the bias it should be exposing. Online metrics can be similarly reassuring in the short term. Click-through rate on a narrowed feed often *rises*: the three surviving categories really are the user's strongest, and short-horizon engagement is exactly what tightening around them optimises. The costs — boredom, session shortening, churn, suppliers with no exposure leaving — land on a horizon longer than a typical experiment window. ## How to detect it Detection is longitudinal and system-level, never a single model's score. - **Interest breadth over time.** For a fixed cohort of users, track the entropy (or distinct-category count) of the categories they are *shown* and the categories they *engage with*, month over month. A monotone decline in shown-category entropy that the engagement side follows is the signature. - **Catalog coverage and impression concentration.** Distinct items served over the catalog, plus a Gini coefficient or top-1%-share over the impression distribution, tracked across successive retrains. A loop shows up as a steadily thinner effective catalog. - **New-item exposure.** What share of impressions goes to items added in the last N days? A closed loop drives this toward zero. - **A measurement population that the current model did not shape.** Any cohort whose data was not produced by the incumbent's ranking gives you an uncontaminated read on whether the narrowing reflects preference or exposure. Without one, every number you have was written by the thing you are auditing. ## What to do about it, at the level of the objective The engineering corrections are their own subject, but two responses belong to how you frame the problem here. First, promote the beyond-accuracy metrics from a monthly report to a **release gate**: a retrain that improves top-k relevance while dropping catalog coverage or user interest entropy past a threshold does not ship. Second, make model selection consider more than one horizon — evaluate a candidate not only on next-click agreement with the log but on what it would do to breadth if it ran for six months. The interview point to land: the model is not just fitting data, it is *choosing* the data the next model will fit. Any argument about a recommender's quality that ignores that is arguing about one frame of a film.
- Why doesn't the offline top-k score fall as the feed narrows?Because the held-out interactions were themselves produced under the incumbent ranker. A candidate model is effectively scored on how well it agrees with what the incumbent chose to show, and there is no logged interaction with the abandoned categories for it to be credited on. The evaluation is drawn from the same distorted distribution, so it validates the narrowing rather than flagging it.
- How would you tell a genuine preference shift from a loop-induced one?Compare the cohort against a population whose slates were not shaped by the current model. If breadth holds up there and collapses in the treated population, the narrowing is exposure, not preference. Also check whether the drop began at a model release boundary and whether engagement per shown item stayed constant — a real preference shift usually shows up as declining engagement with the abandoned categories before their exposure falls.
- What early warning would you put on a dashboard before user churn shows up?Share of impressions going to items published in the last thirty days, distinct items served per week over catalog size, and per-user shown-category entropy for a fixed cohort. All three move months before retention does, and all three are cheap to compute from the impression log you already keep.
A bookshop that reorders only titles that sold last week eventually stocks four books. Sales per title look excellent, and nothing in the sales data can tell you about the customers who stopped coming in.
saying these in an interview costs you the question
- Assumes no clicks on a category means users dislike it
- Trusts a rising offline top-k score as proof nothing is wrong
- Thinks a bigger model or more data fixes the loop
- Diagnoses it from one retrain rather than a trend over months
- Confuses the narrowing with ordinary overfitting to the training set