skip to content

Why is it wrong to compare k-means inertia between two runs built on different feature sets?

level: juniorimportance: should knowfreq 41%

answer

  1. an absolute quantity, not a score
  2. sum of squared distances to centroids
  3. doubling every feature quadruples it
  4. extra columns add extra squared terms
  5. comparable only inside one fixed setup

basics

~20 s

Inertia is the total within-cluster sum of squared distances from points to their cluster centroid. It is an unnormalised quantity whose magnitude grows with feature count, feature scale and row count, so two runs over different feature sets are simply on different scales.

solid answer

~50 s

Inertia is the objective k-means minimises: for every point, the squared distance to its own cluster's centroid, summed over all points. It is a raw sum in the units of the feature space, not a score. Multiply every feature by 10 and inertia grows about a hundredfold because the distances are squared. Add three more columns and every point picks up extra squared deviation along those dimensions. Double the number of rows and the sum roughly doubles. So a drop from 1200 to 340 between two differently built feature matrices says nothing about cluster quality — the two numbers live in different spaces. Inertia is comparable only within one fixed setup: same rows, same columns, same scaling, same cluster count. Its honest use there is comparing random restarts, since k-means only reaches a local optimum and the lowest-inertia restart is the best solution found.

go deeper

for a junior

Know that inertia is a sum of squared distances from points to their cluster centroids, that it has no fixed range, and that quoting it without saying what data and scaling produced it means nothing.

for a middle

Explain the mechanics of why it moves: squared distances make it grow with the square of any uniform rescaling, extra columns add non-negative terms, and extra rows add terms one per row.

for a senior

Demonstrate you would catch this in a review — someone comparing inertia across feature sets — and redirect to a dimensionless index or to a comparison that holds the geometry fixed, such as random restarts.

for a principal

Own the reporting convention: decide what quality numbers a segmentation is allowed to publish, so unnormalised objective values never circulate as evidence of improvement across differently built pipelines.

### The definition Inertia — also called the within-cluster sum of squares — is `sum over clusters, sum over points in that cluster, of squared_distance(point, centroid_of_that_cluster)`. It is exactly the quantity k-means exists to minimise: the assign-then-recentre loop lowers it at every step until assignments stop changing. Two properties follow directly from the formula and cause most of the misuse. **It is a sum, not an average or a ratio.** It has no upper bound and no normalisation. Zero is the floor, reached only when every point coincides with its centroid. There is no scale on which 340 is good and 1200 is bad; the number only means something relative to another number computed on identical geometry. **It is in squared units of the feature space.** Everything that changes those units changes the value: - *Scaling.* Multiply every feature by a constant `c` and every distance is multiplied by `c`, so every squared distance — and the total — is multiplied by `c*c`. Standardising a table therefore produces an inertia unrelated to the raw-table inertia, and usually a different partition too. - *Dimensionality.* Squared Euclidean distance is a sum over columns of squared per-column deviations. Adding columns adds non-negative terms, so a wider feature matrix generally carries a larger inertia even if it describes the same customers. - *Row count.* Each row contributes one term. Cluster the same structure on twice as many rows and the sum roughly doubles. Dividing by the number of rows to get a mean squared distance removes this one effect but leaves scale and dimensionality untouched. ### The failure in the wild The recurring version of this mistake: a first clustering run over eight behavioural columns reports an inertia of 1200; someone adds five more columns, rescales along the way, gets 340, and writes in the deck that the segmentation improved by a factor of three. Nothing of the kind was measured. The second number would be smaller or larger depending on the units of the new columns; it is not evidence about compactness relative to the first run at all. ### Where inertia is legitimately used - **Comparing restarts.** k-means converges to a local optimum that depends on initialisation. Running it many times from different seeds on identical data and keeping the lowest-inertia result is exactly the right use, because every restart shares the same geometry. - **Checking convergence.** Inertia decreases monotonically across the iterations of a single run; watching it flatten confirms the loop settled. - **As an ingredient.** It is the within-cluster dispersion term inside indices that normalise it, which is precisely what makes those indices comparable where the raw sum is not. ### What to report instead If the goal is to compare two candidate partitions built differently, use a dimensionless index rather than a raw sum: silhouette is bounded in `[-1, 1]`, and the Davies-Bouldin and Calinski-Harabasz indices are ratios, so multiplying every feature by one constant leaves them unchanged. Even then, be honest about the limit: an index computed in one feature space answers whether the partition is compact and separated *in that space*. It does not license the claim that one feature space is better than another — that judgement has to come from what the clustering is for, not from a geometry statistic.

  • When is comparing two inertia values actually legitimate?
    When everything except the thing you are comparing is fixed: same rows, same columns, same scaling, same cluster count and same distance. The standard case is random restarts of k-means, which reaches only a local optimum, so you run it repeatedly from different initialisations and keep the run with the lowest inertia. Same geometry, so the numbers mean the same thing.
  • Is inertia bounded, and what does a value of zero mean?
    It is bounded below by zero and unbounded above. Zero means every point coincides with its cluster centroid, which happens only when clusters contain identical points or a single point each. There is no upper limit and no normalisation to a 0-to-1 range, which is why the raw number cannot be read as a quality percentage.
  • What happens to inertia if you standardise the features first?
    It changes completely and unpredictably: standardising rescales each column separately, so distances, the partition k-means finds, and the resulting sum are all different. The pre-standardisation and post-standardisation numbers are not two measurements of the same thing. Choose the preprocessing first, then compare only within it.

saying these in an interview costs you the question

  • Treats lower inertia as better regardless of how the run was set up
  • Compares inertia across differently scaled or wider feature matrices
  • Believes inertia is normalised to a 0-to-1 quality score
  • Confuses inertia with a silhouette-style bounded measure
  • Forgets the sum grows simply because there are more rows

context