A topic model's coherence peaks at 18 topics while held-out perplexity keeps improving — how do you choose K?
answer
- the two metrics measure different goals
- more topics always fit held-out words better
- co-occurrence of top terms versus prediction
- read the top words at each candidate K
- downstream metric outranks both proxies
basics
~20 sThe two metrics answer different questions. Perplexity measures held-out predictive fit and usually keeps improving as topics multiply; coherence measures whether a topic's top words really co-occur. If humans read the topics, follow coherence and inspect the top words yourself.
solid answer
~50 sThey disagree because they are not measuring the same thing. Held-out perplexity is the exponentiated negative average log-likelihood of unseen words: more topics means more flexibility, so it keeps falling long past the point where topics stop being meaningful, and it correlates poorly with human judgments of topic quality. Coherence scores whether a topic's top terms actually co-occur in the corpus, which is much closer to what a reader means by a good topic, and it peaking at 18 is a real signal. So I'd anchor on the coherence peak, then do the decisive step: print the top terms for K = 14, 18 and 24 and read them. Look for duplicate topics, junk topics dominated by boilerplate, and themes you know exist that are missing. If the topics feed a downstream task, let that task's own metric arbitrate — that beats both proxies.
code
python · 23 linesimport math
docs = [
{"loan", "rate", "credit", "bank"},
{"loan", "credit", "score"},
{"rate", "bank", "credit"},
{"pizza", "loan"},
]
top_words = ["credit", "loan", "rate"] # a topic's top terms, most probable first
def doc_freq(w):
return sum(1 for d in docs if w in d)
def co_doc_freq(a, b):
return sum(1 for d in docs if a in d and b in d)
coherence = 0.0
for m in range(1, len(top_words)):
for l in range(m):
rarer, common = top_words[m], top_words[l]
coherence += math.log((co_doc_freq(rarer, common) + 1) / doc_freq(common))
print(round(coherence, 3)) # -0.405go deeper
Know that perplexity is a predictive-fit score where lower is better, that coherence approximates readability, and that the number of topics is a choice you make, not something the model finds.
Explain why perplexity keeps improving as topics are added while coherence peaks, and describe how a coherence score is built from co-occurrence counts of a topic's top words.
Show the practical selection loop: refit with several seeds, read the top terms at candidate K values, diagnose duplicate or boilerplate topics, and defer to a downstream task metric when one exists.
Own the position that no automatic criterion is the acceptance test here, set what the team will report to stakeholders, and be willing to say the corpus has no clean topic count rather than manufacture one.
## Why the two curves disagree This disagreement is the normal case, not an anomaly, and knowing why is the whole answer. **Held-out perplexity** is a predictive-fit measure. You hold out documents (or held-out words within documents), compute the model's average log-probability per held-out word, and report `perplexity = exp(-average log-likelihood per word)`. Lower is better. Adding topics adds parameters and flexibility, so the model can shape the word distribution of unseen text more finely; the curve therefore tends to keep descending, flattening slowly rather than turning up. It is measuring *how well the model predicts word occurrences*, and nothing in that objective rewards a topic being nameable. There is a well-known experimental result behind this: when human raters were asked to spot an intruder word inserted into a topic's top terms, the models that scored best on held-out likelihood were often the ones humans found *least* interpretable. Predictive fit and human-readable themes are genuinely different goals, and optimising the first can hurt the second. **Coherence** was built to close that gap. The common family scores each topic from the co-occurrence statistics of its own top `N` terms: if `credit`, `loan` and `rate` really do turn up in the same documents, the topic is coherent; if a topic's top words never appear together, it is an artefact. One standard form sums, over ordered pairs of top words, `log((co-document count + 1) / document count of the more frequent word)`, then averages across topics. Higher (less negative) is better. Other variants use pointwise mutual information over a sliding window in an external reference corpus. All of them are proxies for readability, and unlike perplexity they do peak: too few topics and each one blends unrelated vocabularies, too many and topics fragment into near-duplicates whose top words are dragged from thin evidence. ## What to actually do 1. **Take the coherence peak as the anchor, not the answer.** A peak at 18 says the neighbourhood of 18 is where topics hang together. Coherence curves are noisy — refit at each K with a couple of seeds before you trust a single point, because a bump of a few percent between K = 16 and K = 18 is often just initialisation. 2. **Read the topics.** This is the step candidates skip and interviewers are listening for. Print the top 10-15 terms per topic for two or three candidate K values and inspect: are there duplicate topics that clearly say the same thing? A junk topic made of boilerplate (signatures, greetings, template text)? A theme you know exists in the corpus that has no topic? Those symptoms tell you to move K, adjust the vocabulary, or change the prior — and no automatic score will tell you which. 3. **Let the downstream task arbitrate if there is one.** If topic proportions feed a classifier, a routing rule, or a retrieval system, that system has a real metric. Sweep K, measure that metric, pick the K that wins. It outranks both proxies because it is not a proxy. 4. **Prefer the smaller K when the curve is flat.** Between two K values with comparable coherence, fewer topics means fewer things for a human to name, maintain, and explain to stakeholders. Interpretability has an ongoing cost, not just a one-off one. 5. **Do not report perplexity as the selection criterion to a business audience.** It is still worth computing as a sanity check — a model whose held-out perplexity is wildly worse than its neighbours is broken — but presenting it as evidence that 60 topics are better than 18 will not survive contact with anyone reading the topics. ## Caveats worth voicing - **Perplexity comparisons must be like-for-like.** Different vocabularies, different held-out splits, or different estimators of the held-out likelihood make the numbers incomparable. If your sweep changed preprocessing along with K, the curve means nothing. - **Coherence has its own failure mode.** It rewards topics built from frequently co-occurring terms, so a topic full of common boilerplate can score beautifully and be useless. That is another reason the eyeball step is not optional. - **The right K may not exist.** If coherence is nearly flat from 10 to 40, the corpus may not have a clean topic structure at all, and forcing a number is a decision about presentation rather than a discovery about the data. Say so rather than manufacturing precision. - **Non-parametric alternatives exist.** A hierarchical Dirichlet process infers the number of topics from the data instead of fixing it, which converts the choice of K into the choice of a concentration parameter. It moves the problem rather than removing it, and adds fitting complexity. ## The one-sentence version Perplexity optimises prediction, coherence approximates readability, and the human reading the topic list is the actual acceptance test — so anchor on coherence, verify by reading, and defer to the downstream metric whenever one exists.
- How exactly is held-out perplexity computed for a topic model?You fit on training documents, then estimate the average log-probability the model assigns to each word in held-out text, and report the exponential of its negative: `exp(-total log-likelihood / number of held-out words)`. Lower is better. The estimate itself is approximate, since the held-out document's topic mixture must be inferred, and different approximations give different numbers — so only compare values produced the same way on the same vocabulary and split.
- What symptoms in the topic list tell you K is too large?Near-duplicate topics whose top terms overlap heavily, topics carried by a handful of documents, and topics made of boilerplate or fragments with no theme. Splitting one real theme across three topics also shows up as several topics sharing their strongest word. Too small a K shows the opposite: single topics that visibly mix two unrelated vocabularies.
- The coherence curve is flat from 10 to 40 topics — what do you conclude?That the corpus probably has no sharp topic structure at this granularity, and the choice of K is a presentation decision rather than a discovery. Report that honestly, pick the smallest K that is workable for the people who must name and maintain the topics, and check whether the documents are simply too short or too homogeneous for the model to separate themes.
saying these in an interview costs you the question
- Picks the K with the lowest held-out perplexity and stops
- Says perplexity measures how interpretable the topics are
- Never looks at the actual top words of any topic
- Treats a noisy single-seed coherence peak as exact
- Assumes higher perplexity is always better