skip to content

A clustering of 8,000 merchants scores ARI 0.18 against merchant-category codes yet is stable under resampling — how do you judge it?

level: principalimportance: nice to knowfreq 26%

answer

  1. two numbers, two different questions
  2. agreement is not accuracy
  3. the taxonomy may encode another axis
  4. noisy labels cap the achievable score
  5. if you want the labels, model them

basics

~20 s

The two numbers answer different questions. Stability says the structure reproduces; a low adjusted Rand index says it does not reproduce that particular taxonomy. Neither is a verdict — the verdict comes from what the clustering was built to support.

solid answer

~40 s

Resist reading the adjusted Rand index as an accuracy score. A stable partition at 0.18 agreement means you found reproducible structure along an axis the category codes do not encode — spend, seasonality, refund behaviour — which is a finding, not a failure. Three checks first: read the contingency table, since a global 0.18 often hides a few categories captured almost perfectly while the rest smear; ask whether the features could express the taxonomy at all; and ask how good the codes are, since self-reported, coarse labels cap any achievable agreement. Then decide by purpose. If the goal was to recover the codes, that is a supervised problem and you should model them directly. If the goal was discovery, the codes are a sanity check and the real evidence is downstream usefulness.

go deeper

for a junior

Know that a low agreement score against existing labels does not by itself mean the clustering is wrong, and that the score is centred so unrelated partitions land near zero rather than at 50%.

for a middle

Be able to open the contingency table and describe the partial correspondence behind an aggregate score, and to explain why noisy or coarse labels put a ceiling on any agreement measure.

for a senior

Show the diagnostic sequence: partial correspondence, whether the features could encode the taxonomy at all, label quality, and stability across time rather than only across random row splits.

for a principal

Own the framing and the objective. Decide before the run what the partition is for and what evidence would retire it, and stop agreement scores from being circulated as accuracy figures or used to tune an unsupervised method into a bad classifier.

## Two numbers, two questions Resampling stability answers: would I find this partition again on different rows? External agreement answers: does this partition line up with a taxonomy someone else built? A clustering can score well on one and badly on the other in either direction, and neither ordering is automatically bad. The combination here — stable, but 0.18 agreement — is the most common and the most misread. The temptation is to treat the adjusted Rand index as an accuracy figure and conclude the clustering failed. That reading imports a supervised frame into an unsupervised problem. The merchant-category codes are one particular partition of the merchants, built for a particular purpose, usually regulatory or billing. The clustering has found a different reproducible partition. Those can both be valid descriptions of the same population. ## What to check before concluding anything **Read the contingency table, not just the scalar.** A global 0.18 is an average over very different local behaviours. It is common to find two or three categories that map almost one-to-one onto clusters, with the remainder spread evenly. That is a much more actionable finding than the aggregate, and it tells you which part of the taxonomy the features do encode. **Ask whether the features could express the taxonomy.** If the inputs are transaction volumes, timing and refund rates, and the codes distinguish businesses by legal industry, there may be no function of the inputs that recovers the codes. In that case no clustering, and no supervised model either, will agree with them, and 0.18 is a statement about the feature set rather than about the algorithm. A cheap test is to fit a supervised model to predict the code from the same features: if it also does poorly, the information is simply not there. **Ask how good the labels are.** Merchant-category codes are frequently self-selected at onboarding, coarse, stale, and inconsistently applied. Label noise puts a ceiling on any achievable agreement score. If two independent labelings of the same merchants exist, their mutual agreement estimates that ceiling; comparing your 0.18 against a ceiling of 0.4 is a very different conversation from comparing it against 1.0. **Ask what could still make the clustering junk.** Stability is not a clean bill of health. Check that the partition is not driven by one unscaled feature or a data artefact such as account age, that cluster sizes are not degenerate, and that the structure survives on a later time window and not only across random row splits. ## Deciding The decision rule is the purpose, not the metric: - **If the objective was to reproduce the codes** — for example to fill them in where missing — this is a supervised problem stated badly. Train a classifier on the codes and evaluate it as a classifier. A clustering tuned until its agreement rises is a worse classifier reached by a longer route, and tuning toward the label set overfits to a taxonomy you had already decided not to trust. - **If the objective was discovery** — finding behavioural structure the existing taxonomy misses — then a low agreement combined with high stability is close to the ideal result, because full agreement would mean the exercise merely rediscovered what you already had. Report it that way, and move the evidence to a downstream measure: does conditioning on the partition improve a decision, a forecast or a targeting rule, evaluated against an unpartitioned baseline? ## What to say to the room The organisational failure mode here is a stakeholder who read 0.18 as 18% correct. Head that off by stating up front what each number measures and what its baseline is: agreement is centred so that unrelated partitions score zero, so 0.18 is a weak but real relationship, not a failing grade; stability says the finding is reproducible; and neither is a measure of business value, which needs its own test. Fix that framing before the numbers are circulated, because it is much harder to reverse afterwards.

  • What evidence would move you from different axis to the clustering is junk?
    Instability across time rather than across random rows; a partition explained almost entirely by one dominant or unscaled feature; degenerate cluster sizes such as one cluster holding 97% of merchants; and no relationship to any external variable you try, not just the codes. Reproducibility on random splits alone does not rule these out, so I would test a later time window and inspect what actually separates the clusters.
  • Would you tune the clustering until its agreement with the codes improves?
    No, that is supervised learning by a longer route. Selecting the algorithm, feature set or k that maximises agreement with a label set is fitting to those labels while pretending the method is unsupervised, and it overfits to a taxonomy you had already judged imperfect. If recovering the codes is the objective, train a classifier on them and evaluate it honestly on held-out merchants.
  • How does noise in the merchant-category codes themselves affect the score you can expect?
    It caps it. If the codes are self-reported at onboarding and never refreshed, a sizeable share is simply wrong, and every wrong label counts against the clustering no matter how good the partition is. Estimate the ceiling where you can — two independent labelings of the same merchants, or a re-labelled audit sample — and judge the observed value against that ceiling rather than against 1.0.

saying these in an interview costs you the question

  • Reads an adjusted Rand index of 0.18 as 18% accuracy
  • Declares the clustering failed because it disagrees with existing labels
  • Treats the category codes as noise-free ground truth
  • Tunes k and features to maximise agreement with the labels
  • Assumes stability alone proves the partition is worth shipping

context