A marketplace trains one ranker on clicks pooled from its browse feed and its typed search results. What goes wrong?
answer
- same schema, different label meaning
- interest against match-to-stated-intent
- abandonment means opposite things
- traffic share silently sets the objective
- separate evaluation before separate models
basics
~20 sThe two surfaces label different things. A feed click signals interest; a search click signals a match to a stated intent, and an abandoned query is a failure on search but an ordinary session on the feed. Pooled, the higher-traffic surface's objective wins.
solid answer
~50 sBoth logs look identical — impression, position, click — but the **meaning of the label differs**. On the browse feed a click means "this caught my eye", so popularity, imagery and freshness are genuinely predictive. On the search page a click means "this matches what I typed", and a query abandoned with no click is a search failure rather than a normal browsing session. Pool the two and the surface with more traffic dominates the objective: usually the feed, so the ranker learns engagement-shaped signals and starts returning popular-but-off-intent results for precise queries. Adding a `surface` feature is not enough, because the problem is in the label semantics, not the inputs. The fix is per-surface objectives — separate models or separate heads — and, at minimum, **separate evaluation sets**, since the feed is judged on engagement while the query surface is judged against relevance judgments.
go deeper
Remember that a click on a browse feed and a click on a search result are not the same signal, even though the two logs have identical columns.
Explain how pooled training lets traffic share decide the objective, and why a surface feature changes the inputs without changing the target.
Describe the production symptom — head queries fine, precise queries degraded, aggregate click-through steady — and split the evaluation before proposing to split the model.
Frame it as a data-against-bias trade: pooling buys examples for a sparse surface and pays in objective mismatch, and say what would make you accept that exchange.
## Identical schemas, different meanings The logs from a browse feed and a typed search page have the same columns: request id, item id, slot position, whether it was clicked, what happened afterwards. That similarity is what makes pooling them so tempting, and it hides the real difference, which is in what the click **means**. | | click on the browse feed | click on a search result | |---|---|---| | what the shopper was doing | browsing without a stated goal | pursuing a goal they typed out | | what a click asserts | "this interested me" | "this matches what I asked for" | | what no click asserts | an ordinary session; nothing is wrong | possibly a failed query | | what predicts it | imagery, popularity, freshness, personal affinity | attribute match, exactness, then attractiveness | A model trained on pooled data learns a single function from features to click probability. Because the two surfaces disagree about what a click means, that single function is a weighted average of two different objectives, and the weights are set by traffic share — an accident of product design, not a decision anyone made. ## Which objective wins, and what it looks like in production On most marketplaces the browse feed generates far more impressions than search. Pooling therefore tilts the model toward engagement-shaped signals. The symptom is specific and recognisable: - **Head queries look fine.** Popular items genuinely are the right answer for `running shoes`, so the engagement bias is invisible. - **Tail and precise queries degrade.** A query naming an exact attribute returns the popular near-miss above the exact match, because popularity is a strong click predictor in the pooled training data and exactness is not. - **Offline relevance numbers drift down while overall click-through holds**, because the feed's traffic carries the aggregate metric. That last point is why the failure survives so long: the aggregate number is fine. It is only visible when the surfaces are measured apart. ## Why a surface feature does not fix it The reflex is to add `surface = feed | search` as an input and let the model learn the difference. That helps with *feature interactions* — it lets popularity matter more on one surface than the other — but it cannot repair the label. The training objective is still "predict a click", and on search the thing you actually want to predict is "is this item a correct answer to this query". Those are different targets, and no input feature turns one into the other. In particular: - **Abandonment is inverted.** A feed session with no clicks is normal. A query with no clicks is a signal the surface failed, and a pooled objective has no way to treat it as such. - **Presentation bias differs.** A grid of large images and a dense list of text rows produce different position-bias curves, so even a correctly-specified position correction fitted on pooled data is wrong for both. - **The candidate distribution differs.** The search ranker only ever sees on-intent candidates, the feed ranker sees a broad pool, so the pooled model is trained on a mixture it will never face at serving time on either surface. ## What to do instead In increasing order of cost: 1. **Separate the evaluation sets first.** This is cheap and it is what makes the problem visible: judge the feed on engagement and session metrics, and the query surface against graded relevance judgments on a sampled query set. If you do nothing else, do this — you cannot manage a regression you cannot see. 2. **Separate the objective.** Either train two models, or train one shared representation with a per-surface head and a per-surface loss, so each surface's label semantics stay intact. 3. **Keep the shared parts shared.** Item and query representations, feature pipelines and the retrieval tier can genuinely be shared; it is the *objective* and the *evaluation* that must not be. ## The counter-argument worth stating Sharing is not always wrong. When one surface has very little traffic, borrowing the other's data is the difference between a trained model and no model, and a shared model with a surface feature can beat a starved dedicated one. The honest framing is a bias-against-data trade: pooling buys examples and pays in objective mismatch, and the exchange rate depends on how different the two surfaces' label semantics really are. What is never defensible is pooling **by default**, without measuring the surfaces apart, because then the trade is being made silently and its cost lands on exactly the queries where the shopper was clearest about what they wanted.
- Why isn't adding a surface feature enough?A feature changes the inputs, not the target. Pooled training still optimises "predict a click", while the search surface needs "is this a correct answer to this query". It also cannot fix inverted abandonment semantics or the fact that the two surfaces have different position-bias curves and different candidate distributions.
- What is the cheapest step that makes this failure visible?Split the evaluation, not the model. Judge the feed on engagement and session outcomes, and the query surface against graded relevance judgments on a sampled set of queries stratified by head and tail. The pooled aggregate hides the regression because feed traffic dominates it; separate readouts expose it immediately.
- When is pooling the two logs actually the right call?When one surface is too sparse to train on alone. Borrowing the other's examples can beat a starved dedicated model, especially for shared item representations. The condition is that you measure the surfaces apart and accept the objective mismatch knowingly, rather than pooling by default and discovering the cost on tail queries.
saying these in an interview costs you the question
- Assuming a click means the same thing on both surfaces
- Believing a surface input feature repairs mismatched label semantics
- Reading a healthy pooled click-through rate as evidence search is fine
- Treating a query abandoned with no click as an ordinary session
- Fitting one position-bias correction across two different layouts