Every retrain of a production topic model returns different topics — how do you deliver something stakeholders can trust?
answer
- the instability is inherent, not a bug
- separate the taxonomy from the model
- freeze and score, refit on a schedule
- align topics, then diff the change
- discovery tool versus reporting tool
basics
~20 sTopic models are not identifiable: random initialisation, a chosen topic count and shifting text mean each refit gives different, differently numbered topics. Freeze one fitted model as the published taxonomy, refit on a deliberate schedule, and review changes with humans.
solid answer
~50 sFirst I name the cause: the instability is inherent, not a bug. Inference starts from a random initialisation and the objective has many near-equal solutions, so topic identities and numbering are arbitrary across runs, and the corpus drifts underneath as well. So I stop treating the model as a live service. I freeze one fitted model, have humans name its topics once, and publish that as the taxonomy; new documents are scored against the frozen topic-word distributions, which is stable by construction. Refitting becomes a scheduled, reviewed event: fit the candidate, align its topics to the incumbent by similarity of their word distributions, and show a diff of what merged, split or appeared. And if the business needs consistent labels against a settled taxonomy, a supervised classifier is the honest instrument — the topic model's real job was helping write that taxonomy.
go deeper
Know that a topic model's topic numbers are arbitrary and that two runs on the same documents can produce different topics, so results should not be quoted as if they were fixed categories.
Explain the causes: random initialisation with many near-equal solutions, sensitivity to the topic count and priors, and a corpus that changes underneath. Know that scoring against a frozen model is stable.
Describe the operating pattern end to end: freeze and score, refit on a cadence, align topics between runs and review a diff, and monitor for documents that fit no existing topic well.
Own the instrument choice and the expectations that go with it: topic models discover, supervised classifiers report, and committing an organisation to unstable categories in its dashboards is the decision you have to make or refuse.
## Diagnose before you fix Three distinct sources of movement get conflated, and separating them is most of the answer. 1. **Non-identifiability of the fit.** Inference for topic models starts from a random initialisation and converges to a local optimum among many of near-equal quality. Topic numbering is arbitrary — there is no reason run 2's topic 7 corresponds to run 1's topic 7 — and beyond relabelling, the actual partition of vocabulary can genuinely differ. Fixing the random seed makes a run reproducible, which is necessary hygiene, but it does not make the result *stable*; it just picks the same arbitrary answer every time, and it stops being the same answer the moment the data or the topic count changes. 2. **Corpus drift.** Real document streams change: new products, new incident types, seasonal vocabulary. Some of the topic movement is the model correctly tracking a world that moved. This is the movement you want to see, and it is why the answer is not simply 'never refit'. 3. **Sensitivity to settings.** The number of topics and the prior concentrations shift where boundaries fall. Two topics that were separate at one setting merge at another, so a change of configuration between runs will look to a stakeholder exactly like model instability. ## The operating pattern that works **Separate the taxonomy from the model.** The deliverable stakeholders care about is a stable set of named themes. Make that an artefact you own, versioned in one place, with human-written names and definitions. The topic model produced version 1; it does not get to silently rewrite it every night. **Freeze and score.** Fit once, have humans review and name the topics, then freeze the topic-word distributions. New documents are scored against those frozen topics, which is cheap and gives identical results for identical inputs forever. Day-to-day reporting is then perfectly stable, and the only moving part is the schedule on which you decide to revisit. **Make refitting a reviewed event.** On a cadence — quarterly is a reasonable default for most content streams — fit a candidate model on refreshed data and produce a *diff*, not a replacement. Align candidate topics to incumbent topics by similarity between their word distributions, then report: which topics matched cleanly, which incumbent topic split into two, which two merged, which candidate topic has no counterpart (a genuinely new theme), and which incumbent topic has no counterpart (a theme that died). Humans approve that diff, keep the old names where topics matched, and name only the new ones. That gives stakeholders continuity through change instead of a fresh set of unfamiliar labels. **Monitor for the thing a frozen model cannot see.** A frozen model forces every new document into old topics; a genuinely new theme is silently smeared across them rather than announced. So track the signal that reveals it: the share of documents whose top topic proportion is unusually low, and the rate of out-of-vocabulary or newly frequent terms. A rise in either is the trigger to refit early rather than wait for the calendar. **Consider that topics were the wrong deliverable.** This is the judgment an interviewer is really probing. Unsupervised topic models are a *discovery* instrument: excellent for the question 'what is in this pile of documents that we did not know about?'. They are a poor *reporting* instrument, because reporting demands a fixed set of categories with agreed definitions and consistent counts quarter over quarter. Once the organisation has settled on a taxonomy, the honest tool is a supervised classifier trained against that taxonomy: labels are stable by definition, accuracy is measurable per class, and adding a category is an explicit decision with a cost. The mature arc is topic model to write the taxonomy, human curation to fix it, supervised model to apply it, topic model again on residual or low-confidence documents to find what the taxonomy is missing. ## What to say about stability numbers If asked to *quantify* instability, describe fitting several models with different seeds and measuring how consistently the same topics reappear — for example, the fraction of a topic's top terms that recur in its best-matching topic from another run, averaged over topics. A model whose topics reappear across seeds is one you can defend; one whose topics scatter is telling you the corpus does not support that many themes, or that the documents are too short for the evidence required. Reporting that fragility up front, before a stakeholder discovers it by noticing last quarter's chart no longer reconciles, is the difference between a model people trust and one they quietly stop using. ## Anti-patterns to name - Refitting nightly and regenerating topic names from each run's top words, so the dashboard's categories silently change meaning. - Fixing the seed and declaring the problem solved, without ever asking whether the topics reproduce under a different seed. - Raising the topic count until labels 'settle', which fragments themes rather than stabilising them. - Averaging topic word distributions across runs without aligning them first, which blends unrelated topics into mush.
- How do you match a new run's topics to the previous run's topics?Compare topic word distributions pairwise — cosine similarity between the word-probability vectors, or overlap of the top terms — and solve the assignment to get the best one-to-one matching. Then read the leftovers: an incumbent topic matching two candidates has split, two incumbents matching one candidate have merged, and an unmatched candidate is a new theme. The unmatched cases are what humans should review.
- Does fixing the random seed solve the stability problem?No. It makes a single run reproducible, which you need for debugging and audit, but it does not make the solution stable: change the data, the topic count or the priors and the topics move again. Worse, it hides the diagnostic — running several seeds and seeing whether the same topics reappear is exactly how you learn whether the structure is real.
- When would you replace the topic model with a supervised classifier?Once the organisation has agreed a taxonomy and needs consistent counts over time. Supervised labels are stable by definition, per-class accuracy is measurable, and adding a category becomes an explicit decision. Keep the topic model running on low-confidence or residual documents, where its discovery strength still earns its place by surfacing themes the taxonomy has no slot for.
saying these in an interview costs you the question
- Says fixing the random seed makes the topics stable
- Regenerates topic names from every nightly refit
- Treats changing topics purely as a data-quality bug
- Averages topics across runs without aligning them first
- Never questions whether topics are the right deliverable