How do you judge whether an unsupervised grouping of unlabelled data is any good?
answer
- no answer key exists here
- the objective is not the truth
- resample and see what persists
- does a downstream decision improve
- hand-label a few, for scoring only
basics
~20 sAn unsupervised grouping has no accuracy score, because there is no ground truth. Judge it on stability under resampling, on whether domain experts recognise the groups, and on whether acting on it improves a measurable downstream outcome.
solid answer
~50 sStart by accepting that no held-out score exists - with no labels there is nothing to be right or wrong about. So I use three kinds of evidence. **Stability**: re-run on bootstrap resamples and different random starts and check whether the same rows keep landing together; a grouping that reshuffles every run is describing noise. **Human validity**: put a sample of each group in front of domain experts and see whether they can name it; a group nobody can characterise is usually an artefact of scaling. **Downstream utility**: the grouping is an intermediate, so score the decision it feeds - if tickets are routed by discovered theme, does handling time drop against a control. I also hand-label a couple of hundred rows purely for evaluation, never for training. What I do not do is quote the method's own objective as evidence it worked.
go deeper
Recall the core fact: with no labels there is no accuracy to report. Be able to say that evidence has to come from stability, expert review, or the value of the decision the output feeds.
Explain the mechanics — why an internal geometric criterion echoes the method's own objective, and how a bootstrap resample plus pairwise co-assignment turns 'is this real' into a number you can quote.
Show you would design the downstream measurement before running anything: which decision changes, which slice of traffic acts as control, and which metric would tell you the structure earned its keep.
Own the reporting standard: what an unsupervised result is allowed to claim to the business, what evidence must accompany it, and when a descriptive answer should be refused in favour of funding labels.
## Why the usual scoreboard is unavailable Supervised evaluation works because the truth is written down: hide labelled rows, predict them, count errors. Unsupervised output has no such column. If you cluster 3 million free-text support tickets into themes nobody has named, there is no held-out label saying which theme each ticket 'really' belonged to. Any sentence of the form 'the clustering was 87% accurate' is meaningless unless someone quietly produced labels — in which case say so, and say how. That leaves four sources of evidence, in roughly increasing order of how much they should move you. ## 1. Internal criteria — weak, and easy to misuse Internal criteria score the geometry of the result using only the data: how tight the groups are and how separated they are from each other. They are worth computing, but they are the weakest evidence available, for a specific reason: **they are usually close relatives of the objective the method was minimising**. A method that packs points tightly around group centres will, unsurprisingly, produce tight groups around centres. Reporting that number as proof of correctness is scoring the exam you set yourself. Internal criteria are useful for comparing runs of the same method under the same assumptions; they cannot tell you the structure is real or that it means anything. ## 2. Stability — cheap, and genuinely informative A structure worth believing should not depend on which half of the data you looked at or which random start you used. The standard check is to draw bootstrap resamples or random subsamples, re-run the whole pipeline on each, and measure how often pairs of rows that were grouped together in one run are grouped together again in another. High co-assignment across resamples means the structure is a property of the data. Low co-assignment means you are describing sampling noise, and any story you tell about the groups will not survive next month's data. Stability is not sufficient — a badly scaled feature can dominate consistently, giving a stable but uninteresting split — but instability is close to disqualifying. ## 3. External validity — bring in truth you did not train on Even when the task has no labels, you can usually obtain *some* truth cheaply if you spend it wisely. Two moves: - **Hand-label a small evaluation sample.** Two hundred tickets read by a support lead is an afternoon's work. Used only for evaluation and never for fitting, it lets you ask concrete questions: do the discovered themes line up with what a human would have said, and where do they cut across each other. Keep the boundary strict — the moment those labels influence the pipeline, they stop being an independent check. - **Check against a variable you held out of the inputs.** If a recorded attribute exists that was deliberately excluded from the features — the queue a ticket was eventually escalated to, the product area it was billed against — you can see whether the discovered groups relate to it. Agreement is corroboration; complete independence is a warning sign. ## 4. Downstream utility — the evidence that actually settles it Unsupervised output is nearly always an intermediate step for a decision: which themes get their own help-centre article, which tickets are routed to a specialist queue, which cohort gets a follow-up. That decision has a measurable outcome, and that outcome is the honest evaluation. Route by discovered theme for one slice of traffic, hold another slice on the current routing, and compare resolution time or reopen rate. This is the only evidence that speaks the stakeholder's language, and it is also the only one that cannot be gamed by choosing a flattering criterion. A cheaper proxy when a controlled comparison is impossible: does the grouping change anybody's behaviour? A set of themes that nobody uses is a failed result no matter how tidy the geometry looked. ## Failure modes to name explicitly - **Confirming the objective.** Quoting the method's own optimisation criterion as evidence the answer is correct. - **Naming after the fact.** Studying a group until a plausible story appears, then presenting the story as a discovery. Any partition of real data admits a narrative; require the story to make a prediction that can fail. - **Ignoring preprocessing.** Feature scaling determines how much each column contributes to similarity, so a 'result' can be an artefact of leaving one feature on a much larger numeric range than the rest. - **Treating the group ids as truth downstream.** Once assignments are fed into another model as a feature, uncertainty about them disappears from the pipeline and reappears as unexplained error later. ## How to report it A credible write-up of an unsupervised result states the preprocessing choices that shaped similarity, shows a stability figure across resamples, shows expert reaction on a sampled set, and proposes the downstream measurement that would confirm the value. It does not lead with a single number, because there is no single number to lead with.
- Why is a tight within-group distance weak evidence that the grouping is correct?Because it is essentially the quantity the method was minimising. A procedure that packs points around group centres will report tight groups whatever the data looks like, so the number confirms that the optimiser ran, not that the structure is real or meaningful. It is fair for comparing runs under identical assumptions and unfair as proof of correctness.
- How exactly would you measure stability of a grouping?Draw bootstrap resamples or random subsamples, re-run the entire pipeline — including preprocessing — on each, then measure how often pairs of rows assigned together in one run are together again in another. High pairwise co-assignment across runs suggests the structure belongs to the data; low co-assignment suggests you are describing sampling noise that will not reappear.
- If you hand-label two hundred rows, why not just train a supervised model on them instead?You may eventually, but spending them on evaluation first is usually the better trade. Two hundred labels give a weak model and a strong measurement, and once they influence fitting they stop being an independent check. Measure first, decide whether the structure is real, and only then argue for funding the larger labelling effort.
saying these in an interview costs you the question
- Quotes an accuracy figure for unlabelled clustering
- Uses the method's own objective as proof it worked
- Says the groups look clearly separated, so it worked
- Never checks whether the groups survive a resample
- Invents a story for each group and calls it validation
- Ignores that feature scaling shaped the similarity