Why does a rating-prediction factorization model need global, user and item bias terms?
answer
- not all of a rating is about taste
- generous reviewer, universally liked film
- one scalar each, not a whole factor dimension
- penalise offsets built on two ratings
- the baseline your factors must beat
basics
~20 sBias terms absorb the part of a rating that is not about taste: the overall rating level, how generous the reviewer is, and how well-liked the film is. Factors then model only the leftover interaction.
solid answer
~50 sThe prediction is `r_hat = mu + b_u + b_i + dot(p_u, q_i)`. Much of the variance in a ratings table is not personal: a reviewer who gives every film a 4 shifts all their ratings up, and a film everyone likes shifts all of its ratings up, whoever is watching. With only the dot product, the model must burn latent factors encoding "this user is generous" and "this film is popular" — capacity spent on effects one scalar each explains perfectly. Adding `mu`, `b_u` and `b_i` strips those out first, so the factors carry genuine affinity. Biases are learned jointly with the factors and get their own L2 penalty, which matters because a user with two ratings would otherwise receive a wild offset. A bias-only baseline is strong on its own: if the factor model barely beats it, the factors are not earning their keep.
go deeper
Be ready to name the three bias terms — global mean, user offset, item offset — and give the everyday example of a reviewer who rates everything highly versus a film almost everyone likes.
Explain why an additive scalar is a better place for those effects than a latent dimension, and what regularizing the biases does to an item supported by only two ratings.
Demonstrate that you benchmark against the bias-only model, treat a small gap as a signal that the factor model is mis-specified, and know biases drift over time in a live catalogue.
Own the call on how much of the system should be popularity versus personalisation, and be ready to argue when a strong bias model plus business rules beats shipping and maintaining a factor model at all.
## Decomposing a rating When someone gives a film four stars, at least four things went into that number: 1. **The overall level of the rating scale as people actually use it.** Ratings cluster around 3.5-3.7 on a five-point scale in most catalogues, not around 3.0. That is the global mean `mu`. 2. **How this particular reviewer uses the scale.** Some reviewers rate almost everything 4 or 5; some reserve 5 for a handful of films a decade. That is the user bias `b_u`, the reviewer's typical offset from `mu`. 3. **How this particular film is received in general.** A widely-loved film pulls ratings up from nearly everyone. That is the item bias `b_i`. 4. **What is left over: whether *this* reviewer, specifically, is a match for *this* film.** That is the interaction, and it is the only part the latent factors should have to model. So the model is ``` r_hat(u, i) = mu + b_u + b_i + dot(p_u, q_i) ``` and the biases are learned parameters, not fixed preprocessing constants. ## Why not let the factors handle it They can, badly. A dot product *can* represent a per-user constant — reserve one latent dimension, set that coordinate of every item vector to 1, and put the user's offset in the matching coordinate of the user vector. But that spends a whole factor dimension on something one scalar per user encodes exactly, and it does so through a product of two learned quantities, which is a much harder thing to estimate than a single additive number. On a model with `k = 20`, giving up two dimensions to popularity and generosity is a 10% cut in the capacity available for real taste. Worse, without biases the factor vectors are contaminated. A generous reviewer's vector acquires a large positive component along the "everything is popular" direction, which then makes them look similar in factor space to every other generous reviewer — regardless of whether their tastes match at all. Biases keep the latent space about taste. ## A concrete pair of cases Take a reviewer who has rated 300 films and given all of them a 4. Their user bias is roughly `4 - mu`, a few tenths of a star, and their factor vector should shrink towards zero because their ratings carry no information about *which* films they prefer — there is no variation to explain. That is exactly what a bias term plus L2 regularization produces: the prediction for this user on any film becomes `mu + b_u + b_i`, that is, the general opinion of the film shifted by the reviewer's generosity. Sensible. Now take a film every reviewer rates highly. Its item bias is large and positive; its factor vector should be near zero, because it does not divide opinion. The model then predicts a high rating for everyone, which is right — universally liked films are the case where personalisation has nothing to add. Both are absorbed by the bias terms *before any factor is learned*, and that is the point. ## Regularize the biases too A film with two ratings, both 5, would get an item bias near `5 - mu` if left unpenalised, and the model would then confidently recommend it to everyone. The same L2 penalty applied to the factors is applied to `b_u` and `b_i`, which shrinks thinly-supported offsets towards zero — the shrunken estimate is close to a damped mean, where the item's observed mean is pulled towards the global mean in proportion to how few ratings support it. This is one of the highest-value details in the whole model, because the long tail of items with a handful of ratings is where an unregularized bias does its worst damage. ## The baseline you must beat `mu + b_u + b_i` alone, with no factors at all, is a genuinely strong predictor — it explains a large share of the variance in rating data. Two practical consequences. First, always report it: if the factor model's held-out error is barely better than the bias-only baseline, the factors are noise and you should question the rank, the regularization or the data. Second, it is your degraded-mode answer. A user or item too new to have a meaningful factor vector still has a bias (or falls back to `mu`), so the system always returns something defensible. ## A caveat worth voicing Biases are stated here as static parameters, but rating behaviour drifts — a reviewer's standards harden over time, and a film's reception changes after it becomes a classic or a punchline. Production models often let `b_u` and `b_i` vary with time, which typically buys more accuracy than adding factor dimensions does.
- Could you just subtract the user and item means as preprocessing instead of learning biases?You can, and it is a reasonable quick baseline, but it is worse in two ways. Raw means are unshrunk, so an item with two ratings gets an extreme offset, and sequential subtraction of user then item means is order-dependent. Learning `b_u` and `b_i` jointly with the factors under one L2 penalty solves both.
- Why should the bias terms be regularized at all — they are only one number each?Because their support varies enormously. A film with 40,000 ratings supports a precise offset; a film with two ratings does not, yet an unpenalised fit gives it just as extreme a value. The penalty pulls low-support offsets towards the global mean, which is exactly the damped-mean behaviour you want in the long tail.
- How do you tell whether the latent factors are adding anything over the biases?Fit the bias-only model and the full model on the same split and compare held-out error. If the gap is small, the factors are not finding structure — usually because the rank is too high for the amount of data, the penalty is mis-set, or the interactions in your catalogue are genuinely weak relative to popularity.
saying these in an interview costs you the question
- Says the dot product already covers per-user offsets
- Leaves bias terms out of the regularization penalty
- Uses raw unshrunk means for items with two ratings
- Never reports the bias-only baseline
- Claims popularity effects are what the factors are for