skip to content

When would you prefer NMF over latent semantic analysis for extracting topics?

level: middleimportance: should knowfreq 45%

answer

  1. one constraint decides readability
  2. signed versus non-negative entries
  3. parts you add, never subtract
  4. synonymy is what LSA buys
  5. who reads the output decides

basics

~10 s

Prefer non-negative matrix factorization when people must read the topics: non-negativity makes every topic an additive list of words. Latent semantic analysis produces signed components that capture synonymy well but rarely read as themes.

solid answer

~50 s

Both factorise a term-by-document matrix into a small number of latent dimensions, but they differ in one constraint that decides interpretability. Latent semantic analysis takes a truncated singular value decomposition, so its components are orthogonal and **signed** — a dimension can have a negative weight on `election`, and nobody can explain what a negative amount of a word means. NMF instead minimises reconstruction error subject to both factors being non-negative, so a topic is a purely additive weighted word list and a document is a non-negative blend of topics. On a news archive that difference is stark: NMF gives readable parts-based themes, LSA gives axes an analyst cannot label. What LSA buys is synonymy — `car` and `automobile` collapse onto a shared latent dimension, so documents match without sharing terms. If the output is a retrieval index, that matters more than readability; if it goes on a slide, choose NMF.

go deeper

for a junior

Know that both compress a term-by-document matrix into a few latent dimensions, and that NMF's weights are all non-negative while latent semantic analysis produces signed ones.

for a middle

Explain why the non-negativity constraint yields additive, readable topics, what objective NMF minimises, and what synonymy handling buys latent semantic analysis in retrieval.

for a senior

Bring the operational differences: NMF is not nested across ranks and depends on initialisation, LSA is deterministic, and scaling of the input matrix changes which NMF topics dominate.

for a principal

Frame the choice by consumer and lifecycle: human-facing themes demand non-negativity and stable naming, machine-consumed latent features do not, and that decision sets the evaluation you commit to.

## The shared idea Both methods start from a term-by-document matrix and compress it to `k` latent dimensions, on the bet that the tens of thousands of vocabulary columns are really a few dozen themes wearing different words. Both are linear and deterministic given the same inputs and rank. Where they differ is the constraint imposed on the factors, and that single difference drives almost everything an interviewer wants to hear. ## Latent semantic analysis LSA factorises the matrix with a truncated singular value decomposition and keeps the top `k` dimensions. Two properties follow: - **Orthogonality.** The latent dimensions are mutually orthogonal, ordered by how much of the matrix's structure each explains. That gives a clean, nested sequence: the rank-k solution is the rank-(k-1) solution plus one more dimension. Convenient, and it means there is exactly one answer, with no random restart to worry about. - **Signed entries.** Weights can be negative. A latent dimension might load +0.3 on `budget` and -0.2 on `striker`. Mathematically fine; as a theme, meaningless. Asking what a document containing a negative amount of `striker` looks like has no answer. The headline benefit of LSA is **synonymy**. Because latent dimensions are built from co-occurrence structure rather than exact strings, `car` and `automobile` — which never need to appear in the same document — end up loading on the same dimension because both co-occur with `engine`, `mileage`, `dealer`. A query about cars then retrieves automobile documents with zero term overlap, which exact keyword matching cannot do. The mirror-image weakness is **polysemy**: `bank` has one vector, so the river sense and the finance sense are averaged into a single confused direction. ## Non-negative matrix factorization NMF approximates the matrix `V` as `W H` where every entry of both factors is constrained to be at least zero. It minimises a reconstruction error — squared Frobenius error, or a Kullback-Leibler style divergence when the data are counts — subject to that constraint, typically with multiplicative update rules that preserve non-negativity by construction. The non-negativity constraint is doing the interpretive work. With no subtraction allowed, the only way to reconstruct a document is to *add* topics together, and the only way to build a topic is to *add* words together. The result is what the literature calls a parts-based representation: each of NMF's columns is a set of words that genuinely co-occur, and each document is a positive blend of them. On a news archive this is the difference between a topic you can label `municipal budgets` at a glance and an LSA axis you would need a paragraph to describe and still get wrong. The costs: - **Non-uniqueness.** The objective is not jointly convex in both factors, so the solution depends on initialisation, and different runs give different (usually similar-quality) topics. - **Not nested.** Fitting with `k = 20` does not contain the `k = 15` solution. Every rank is a separate fit, so a sweep is more expensive than LSA's. - **No probabilistic story.** Topic weights are non-negative numbers, not probabilities. They do not sum to one and there is no likelihood, so likelihood-based model comparison is unavailable — you evaluate by reconstruction error, by coherence of the top words, or by whether the downstream task improves. - **Scale sensitivity.** NMF is fitting a weighted least-squares-style objective on the raw entries, so how the input matrix is weighted changes which topics dominate. Frequent terms can drown the rest if the weighting does not damp them. ## Where LDA sits between them Asking how NMF relates to LDA is a standard follow-up. LDA is a generative probabilistic model: documents get *probability* mixtures over topics, topics get probability distributions over words, and priors control sparsity. NMF is deterministic matrix arithmetic with a non-negativity constraint. In practice their outputs often look similar, and NMF is usually faster and more stable on small corpora, while LDA gives calibrated proportions, a principled way to score held-out documents, and priors to tune. Under a particular divergence objective and normalisation, NMF and a probabilistic topic model are closely related — but in an interview, the useful statement is the practical one: NMF for speed and clean readable topics on modest corpora, LDA when you want a probabilistic mixture with held-out scoring. ## How to answer the choice Name the deciding question: *who reads the output?* If a human reads topics and names them, non-negativity is not a nicety, it is the requirement, and NMF (or LDA) wins. If the latent space is machine-consumed — a retrieval index, features for a downstream model — the signed LSA representation is fine and its determinism and nestedness are advantages. Saying nothing about the consumer and reciting both algorithms is the weak answer.

  • Why does NMF give different topics on different runs when LSA does not?
    LSA's decomposition is unique up to sign and ordering, so the same matrix and rank always yield the same subspace. NMF's objective is not jointly convex in its two factors; the updates converge to a local optimum that depends on the random initialisation. In practice you run it several times, keep the fit with the lowest reconstruction error or best coherence, and record the seed.
  • How would you pick between NMF and LDA once you have ruled out LSA?
    Ask whether you need probabilities. LDA gives per-document proportions that sum to one, a likelihood for held-out documents, and priors that tune sparsity — useful when the mixture itself is the deliverable. NMF is faster, deterministic given a seed, and often crisper on small or short-document corpora, but its weights are just non-negative numbers with no probabilistic reading.
  • What does LSA do badly that a reader of its results should know about?
    Polysemy. Every term gets exactly one vector, so a word with two senses is represented as an average of both, and documents from either sense are pulled toward that blurred direction. Latent dimensions are also unlabelled and signed, so what looks like a clean semantic axis often mixes several themes with opposing signs.

NMF builds a picture out of ingredients you can only pour in. LSA is allowed to pour some out again, so its themes can contain a negative amount of a word, which no reader can picture.

saying these in an interview costs you the question

  • Says NMF components are orthogonal like LSA's
  • Claims NMF weights are probabilities that sum to one
  • Thinks a negative component weight means the topic is absent
  • Assumes LSA and NMF give the same topics with different names
  • Ignores that NMF depends on its random initialisation

context