How does a content-based recommender score an item using only that user's own history?
answer
- describe the item, not the crowd
- user and item live in one space
- profile from the user's own likes
- average the liked feature vectors
- cosine between profile and candidate
basics
~20 sIt describes each item by its own features, builds a profile vector from the features of items that user liked, and ranks candidates by similarity between profile and item vector. No other user's data is involved.
solid answer
~50 sContent-based recommendation has two halves. First, every item becomes a feature vector: a job posting becomes its seniority level, its skill tags, its location, its employment type. Second, the user becomes a vector in the same space, usually the average of the feature vectors of the items they applied to or saved, optionally weighted by recency or rating. Scoring is then a similarity between the two vectors, most often cosine, `score = dot(u, i) / (norm(u) * norm(i))`, which normalises away the fact that some items simply carry more tags than others. An equivalent framing is a tiny per-user classifier trained on that user's own likes and dislikes over item features. The payoff is that an item posted this morning with zero applications is scorable the moment its features exist. The price is overspecialisation: it keeps proposing more of what the user already picked.
code
python · 25 linesimport math
# each item is described by its own skill tags; the user applied to A and B
items = {
"A": {"python": 1, "sql": 1},
"B": {"python": 1, "ml": 1},
"NEW": {"python": 1, "ml": 1, "sql": 1}, # posted today, zero applications
"OTHER": {"figma": 1, "css": 1},
}
liked = ["A", "B"]
profile = {}
for key in liked:
for tag, weight in items[key].items():
profile[tag] = profile.get(tag, 0) + weight / len(liked)
def cosine(a, b):
dot = sum(v * b.get(t, 0) for t, v in a.items())
na = math.sqrt(sum(v * v for v in a.values()))
nb = math.sqrt(sum(v * v for v in b.values()))
return dot / (na * nb) if na and nb else 0.0
print("profile:", profile)
for key in ("NEW", "OTHER"):
print(key, round(cosine(profile, items[key]), 3))go deeper
Be ready to state the two halves out loud: items become feature vectors, the user becomes the average of the vectors of items they liked, and the score is the similarity between them. Know that this is what lets a brand-new item be ranked.
Explain why cosine rather than a raw dot product, how you weight liked items by recency and signal strength, and the equivalent framing as a small per-user classifier over item features.
Show where the approach breaks in production: feature coverage gaps, profiles that go stale, overspecialisation choking discovery, and the point at which you start blending in a collaborative signal instead.
Own the tradeoff between an explainable feature-driven system and a stronger but opaque behavioural one, including what feature pipelines and editorial tagging you are committing the organisation to maintain forever.
## What "content-based" means A recommender is content-based when the score for a (user, item) pair is computed from a **description of the item itself** plus **that one user's own history**, and from nothing else. No other user's behaviour enters the calculation. The contrast is collaborative filtering, which ignores what an item is *about* and learns purely from the pattern of who interacted with what. That single design choice is why content-based methods are the standard answer to new-item cold start. A job board that publishes a posting at 09:00 with zero applications has, from a collaborative point of view, nothing at all to work with. From a content point of view it has everything it needs: the title, the seniority, the skill tags, the location. ## Half one: representing the item Every item is turned into a numeric feature vector over a fixed set of columns: - **Categorical attributes** — one column per skill tag, per category, per employment type; 1 if present, 0 otherwise. - **Ordinal or numeric attributes** — required years of experience, salary band midpoint, usually rescaled so one large-valued column does not dominate the similarity. - **Attributes derived from item text** — the free-text description has to be turned into numbers by some featurisation step before it can be used here. Once you have that vector, the recommender treats it exactly like any other block of columns. The quality ceiling of the whole approach is set here. If the only feature you record is a top-level category, every posting in that category looks identical to the model and it cannot rank within the category at all. ## Half two: representing the user The user profile lives in the *same* feature space as the items, which is what makes them comparable. The simplest construction is the centroid: average the feature vectors of the items the user engaged with positively. ``` profile = (1 / n) * sum(feature_vector(item) for item in liked_items) ``` Common refinements: - **Weight by strength of signal** — an application counts more than a page view; a 5-star rating counts more than a 3. - **Decay by recency** — a profile built from three-year-old behaviour describes a person who no longer exists. An exponential decay on the weights keeps it current. - **Use negative signal** — instead of averaging only the likes, fit a small discriminative model per user (a logistic regression over the item features, say) on liked versus skipped items. This can learn that the user wants Python *without* on-call, which a positives-only centroid cannot express. ## Scoring With both sides as vectors, the score is a similarity. Cosine dominates: ``` cos(u, i) = dot(u, i) / (norm(u) * norm(i)) ``` The normalisation matters. A raw dot product rewards items simply for having many non-zero features, so a posting that lists twenty skill tags outranks a precise match with three. Cosine measures the *angle* — how aligned the descriptions are — and is scale-free. For non-negative feature vectors it lands in `[0, 1]`; with signed features it ranges `[-1, 1]`. The alternative framing is per-user supervised learning: treat the user's history as a labelled dataset (item features as X, engaged/not as y), fit a small model, and score candidates with it. Same information, different bias-variance tradeoff — the classifier can learn feature interactions the centroid cannot, but it overfits fast on a user with eleven data points. ## What it buys and what it costs **Buys:** - **New items are scorable immediately.** No interaction history is required on the item side, which is the entire reason this leaf exists. On a news homepage where most inventory is hours old, this is not an edge case; it is the steady state. - **Explainability.** "Recommended because you applied to two other senior Python roles" falls straight out of which features drove the similarity. A latent-factor score has no such reading. - **Independence from other users.** A niche user with unusual taste is served as well as a mainstream one, because nobody has to resemble them. **Costs:** - **Overspecialisation.** The profile is built from what the user already chose, so recommendations stay in that neighbourhood. There is no mechanism by which an item unlike anything in the history can surface, and therefore little serendipity. - **A hard ceiling at feature quality.** Content similarity cannot capture "people who liked this also liked that" correlations that no recorded attribute encodes — production quality, writing style, how well a hiring manager actually treats people. - **It does not solve new-*user* cold start.** The profile still needs the user's own history. A subscriber who has clicked nothing has an empty profile, and content features do not help. - **Similarity is not preference.** Two postings can be near-identical in feature space and differ enormously in whether the user wants them. In practice this is why real systems are hybrids: content features carry an item until it has enough interactions for a collaborative signal to take over, and the two scores are blended with a weight that shifts as data accumulates.
- Why is a content-based recommendation easier to explain to a user than a latent-factor one?The features are human-readable, so you can point at the ones that drove the similarity: recommended because it is a senior Python role and you applied to two others. A latent factor is an anonymous coordinate fitted to interaction data; nothing names what dimension 7 means, so any explanation attached to it is reconstructed after the fact.
- What is overspecialisation, and why is it the characteristic weakness here?The profile is the average of what the user already chose, so the highest-scoring candidates are always near that average. Anything genuinely different scores low by construction, so the user never discovers a new area. It is structural, not a tuning bug: the only inputs are the user's own past choices, so nothing can pull the recommendations outside them.
- How would you weight items when averaging them into a profile?Weight by signal strength and by recency: an application outweighs a page view, and an exponential decay stops three-year-old behaviour from defining a current profile. If you have negative signal, prefer fitting a small per-user classifier on liked versus skipped items over a positives-only centroid, since that can express what the user avoids.
A content-based recommender is a librarian who has actually read the books and remembers what you borrowed. A collaborative one has read nothing and only knows which borrowers overlap with you.
saying these in an interview costs you the question
- Says content-based means finding users with similar taste
- Claims content-based needs no user history at all
- Uses a raw dot product and lets tag-heavy items win
- Assumes feature similarity guarantees the user will like it
- Thinks content features also fix new-user cold start