skip to content

How does latent Dirichlet allocation model a document as a mixture of topics?

level: middleimportance: must knowfreq 60%

answer

  1. documents get proportions, not labels
  2. two distributions are learned, not one
  3. topic drawn per word position
  4. Dirichlet prior over the mixture
  5. 0.6 billing, 0.3 login, 0.1 shipping

basics

~20 s

Latent Dirichlet allocation treats each document as a probability distribution over K topics, and each topic as a probability distribution over vocabulary words. Fitting infers both, so one support ticket comes out as 0.6 billing, 0.3 login, 0.1 shipping.

solid answer

~50 s

LDA is a generative story. Each document is assumed to be written by first drawing a mixture of `K` topics from a Dirichlet prior, then, for every word position, drawing a topic from that mixture and a word from that topic's distribution over the vocabulary. Fitting runs the story backwards: inference recovers, per document, its topic proportions, and per topic, its word probabilities. The practical consequence is soft membership. Over 50,000 customer-support tickets, one ticket comes back as 0.6 billing, 0.3 login, 0.1 shipping instead of a single label, which is what you want for documents that genuinely cover several things. Two concentration parameters control how peaked things are: a small document-topic concentration pushes each document onto few topics, a small topic-word concentration pushes each topic onto few words. `K` is your choice, fixed before fitting.

go deeper

for a junior

Be ready to say what comes out of a topic model: per-document proportions over topics, and per-topic word probabilities. Know that a document can belong to several topics at once.

for a middle

Explain the generative story end to end, including that the topic is drawn per word position, and say what the two Dirichlet concentration parameters do to sparsity. Know that K is fixed by you.

for a senior

Show judgment about when mixtures beat hard labels, how you score new documents against a frozen model, and what short documents or a wrong K do to the output quality.

for a principal

Own the framing question: whether an unsupervised mixture is the right instrument at all versus an agreed taxonomy with a supervised classifier, and what the business does with proportions it cannot name.

## What LDA is trying to explain Latent Dirichlet allocation is a *generative* probabilistic model of a document collection. Rather than starting from a similarity measure and grouping documents, it writes down a made-up story of how each document could have been produced, then asks: which hidden quantities make the documents I actually observe most probable? Those hidden quantities are the topics. Two objects are latent (unobserved) and both are recovered by fitting: - **Topic-word distributions.** A *topic* is not a name or a label. It is a probability distribution over the whole vocabulary. Topic 3 might put probability 0.04 on `invoice`, 0.03 on `charge`, 0.03 on `refund`, and almost nothing on the other 40,000 words. A human reading the top terms calls it billing; the model itself has no idea what it is called. - **Document-topic distributions.** Each document has a vector of `K` non-negative proportions that sum to 1: how much of this document came from each topic. ## The generative story For a chosen number of topics `K`: 1. For each topic k, draw a word distribution `phi_k` from a Dirichlet prior with concentration `beta`. 2. For each document d, draw a topic mixture `theta_d` from a Dirichlet prior with concentration `alpha`. 3. For each word position in document d: draw a topic index `z` from `theta_d`, then draw the actual word from `phi_z`. That is the whole model. Note step 3: the topic is drawn *per word*, not per document. This is exactly why a document ends up as a mixture — different word positions in the same document can come from different topics, and the document's proportions are just the tally of where its words came from. ## Why a Dirichlet prior A Dirichlet distribution is a distribution over probability vectors that sum to 1 — the natural prior when the thing you are drawing is itself a mixture. Its concentration parameter controls sparsity: - `alpha` well below 1: documents concentrate on a handful of topics (most documents are about one or two things). - `alpha` above 1: documents spread mass evenly across many topics. - `beta` well below 1: topics concentrate on a small set of words, which usually reads as crisper themes. So the priors are not decoration — they are the main non-`K` knobs you have, and they change how interpretable the output looks. ## Inference The posterior over the hidden variables is intractable in closed form, so it is approximated. The two standard families are **collapsed Gibbs sampling**, which repeatedly reassigns each word token to a topic given all the current assignments, and **variational inference**, which fits a simpler tractable distribution to the posterior by optimisation. Both are iterative, both start from a random initialisation, and neither is guaranteed to reach a global optimum. That is worth remembering: two runs on the same corpus give different topics, and their numbering is arbitrary anyway. ## Soft membership is the point The single most useful thing to say in an interview: LDA does not partition documents. A hard clustering forces every ticket into one bucket, which is wrong for a ticket that says *my card was declined so I could not log in and my order never shipped*. LDA returns proportions, so mixed documents stay mixed. Downstream you can still take the argmax topic if you need a single label, but you are throwing information away when you do, and you can see from the proportions whether that argmax was 0.9 (confident) or 0.34 (a coin flip between three themes). ## Assumptions and limits - **Word order is ignored.** LDA treats a document as an exchangeable collection of word occurrences; `bank charged me` and `me charged bank` are the same input. That is why it is cheap and why it cannot model phrases or negation. - **`K` is fixed in advance.** The model will not tell you how many topics exist; it will happily split one real theme into three or merge two into one if you mis-set `K`. Non-parametric extensions such as the hierarchical Dirichlet process infer the count instead, at a cost in complexity. - **Topics are not labels.** Naming a topic from its top words is a human step, and a stakeholder-facing one. - **Small or short documents are hard.** With a few dozen words per document there is little evidence to split among `K` topics, and mixtures come out noisy. ## Scoring a new document Once fitted, the topic-word distributions are frozen and a new unseen document can be scored by inferring only its own topic proportions against those fixed topics. That is what lets a fitted topic model act as a fixed feature extractor: every incoming ticket becomes a `K`-dimensional vector of proportions that a downstream classifier or dashboard can consume.

  • What do the two Dirichlet concentration parameters actually change in the output?
    The document-topic concentration controls how many topics a typical document uses: below 1 it pushes each document onto a few topics, above 1 it spreads mass across many. The topic-word concentration does the same for words within a topic — small values give sparse, sharper-reading topics dominated by a few terms. Neither changes how many topics exist; that is fixed by K.
  • How do you get topic proportions for a brand-new document without refitting the model?
    Freeze the topic-word distributions and run inference for that document alone, estimating only its mixture over the existing topics. This is cheap and keeps the topics stable, which is what you want in production. The catch is vocabulary: words the fitted model never saw are simply dropped, so a document about a genuinely new theme is forced into old topics rather than revealing the gap.
  • When is taking the single highest-probability topic per document defensible?
    When you need one label for routing or reporting and the mixtures are actually peaked — check the distribution of the top proportion first. If most documents have a top topic above roughly 0.7 the argmax loses little. If the top proportion clusters near 1/K, the model is telling you documents are genuinely mixed or K is wrong, and a hard label will be arbitrary.

A document is a playlist, not a genre. LDA reports what fraction of the playlist was drawn from each genre's song pool, rather than stamping the whole playlist as jazz.

saying these in an interview costs you the question

  • Says LDA assigns each document to exactly one topic
  • Confuses it with linear discriminant analysis, a supervised method
  • Thinks the model outputs human-readable topic names
  • Claims LDA uses word order or sentence structure
  • Believes the model discovers the number of topics itself

context