On a silhouette plot the average is 0.42 but one cluster's bars are negative and another's are all short — what does that tell you?
answer
- read the shape, not the mean
- one number is hiding five clusters
- negative bars name a preferred neighbour
- bars near zero mean frontier points
- block height is the point count
basics
~20 sThe negative-bar cluster holds points that are on average closer to another cluster, so it overlaps a neighbour. The short-bar cluster sits on a boundary with no space of its own. A respectable-looking average of 0.42 is being carried by the healthy clusters.
solid answer
~50 sA silhouette plot draws one horizontal bar per point, grouped and sorted by cluster, so it shows the whole distribution the average collapses. A cluster whose bars run negative is not a cluster of outliers — those points are on average nearer another cluster's members than their own, which usually means two groups overlap or one real group was split in two. A cluster whose bars are all short but positive sits in the frontier zone between neighbours: it exists, but it has no separated territory of its own, so it is largely boundary residue. Meanwhile the overall 0.42 is carried by whichever clusters are compact and separated, which is exactly why the headline number should never be reported alone. I would read the per-cluster means and the block heights, tabulate which neighbour the negative points prefer, and treat the overlapping pair as one candidate group rather than two.
go deeper
Know how to read the picture: each bar is one point's score, longer is better, and bars crossing to the left of zero mark points that fit a different cluster better than their own.
Explain why the overall mean can look fine while a cluster is broken, and distinguish the two defects on the plot: bars below zero versus bars bunched just above zero.
Turn the plot into an action. Identify the preferred neighbour of the negative points, weigh the defect against cluster size, check feature scaling and re-run stability before redesigning anything.
Set the standard for how partition quality is presented. Decide that segmentation reviews show the distribution and cluster sizes rather than one headline score, so a failing segment cannot pass unnoticed.
### What the plot actually draws A silhouette plot has one horizontal bar per point. Bars are grouped by cluster and sorted within each group, so each cluster appears as a wedge. Two things are readable at a glance: the **length** of the bars, which is each point's silhouette score, and the **height** of each block, which is how many points that cluster contains. A single average destroys both. Take a five-cluster partition of e-commerce session features — session length, pages viewed, cart events, time since last visit and so on. The overall silhouette is 0.42, which reads as acceptable. The plot tells a different story. ### The cluster with negative bars Negative silhouette means the cohesion term exceeded the separation term: on average these points are nearer another cluster's members than their own. A whole block running negative is a structural statement, not a handful of odd rows. The two usual causes: - **Two groups overlap.** The cluster occupies the same region as a neighbour and the boundary between them is arbitrary. Whatever produced the split is drawing a line through one continuous mass of sessions. - **One real group was cut in two.** Half the sessions ended up on each side of a division that does not exist in the data, so each half is closer to the other half than to itself. The useful next step is that silhouette already names the culprit: the separation term is defined by the *nearest* other cluster, so you can record, per negative point, which cluster that was. If four fifths of the negative points in cluster 4 point at cluster 2, you have located a specific overlapping pair rather than a vague quality problem. Size changes the reaction. If the negative block is 0.5 percent of sessions, it may be a thin residue of unusual traffic and the right response is to check whether it even survives a re-run from a different initialisation. If it is a fifth of the sessions, the partition itself is wrong. ### The cluster with uniformly short bars Scores clustered just above zero mean cohesion and separation are nearly equal for every member. The group is real enough to be assigned, but it lives in the frontier between two better-defined clusters — the residue left over once the distinct groups have claimed their points. It is not a small cluster, note; smallness would show as a short block, not as short bars. A boundary cluster like this is usually the least actionable segment of the set, because nothing about it is distinct. ### Why the average misled The overall silhouette is a mean over points, so a large, clean, well-separated cluster contributes many strongly positive values and can hold the headline number up while two other clusters are failing. Reporting per-cluster means is the minimum fix. Even that is not enough on its own: a cluster in which half the points score 0.8 and half score -0.1 averages roughly the same as one where every point scores 0.35, and those are very different situations. The distribution is the object of interest; the plot is how you see it. ### What to do with the diagnosis Read the per-cluster means and block heights together. Identify the preferred neighbour of the negative points, and treat the overlapping pair as a single candidate group unless there is a reason outside the geometry to keep them apart. Check whether the features are on comparable scales, since one dominating feature manufactures exactly this pattern — long thin regions carved arbitrarily. Re-run from several initialisations and see whether the negative block reappears in the same place; a defect that moves is an artefact of initialisation, a defect that stays is in the data. And be clear about what the plot cannot tell you: it evaluates geometry only, so a partition can look immaculate here and still be useless for the decision it was built to support.
- How would you find out which cluster the negative points would rather belong to?The separation term in the silhouette definition is the mean distance to the *nearest* other cluster, so that nearest cluster is already computed per point. Record it for every negative point and tabulate it. A concentration on one neighbour identifies a specific overlapping pair to merge or re-examine, instead of a general complaint that the clustering is poor.
- The negative cluster holds only 0.5 percent of the sessions. Does that change your response?Yes. A tiny negative block is more likely a residue of unusual traffic than a structural flaw, so the first check is stability: re-run from several initialisations and see whether it reappears in the same place with the same members. If it moves, it is an initialisation artefact. If a fifth of the sessions were negative, the partition itself would be the problem.
- Would reporting per-cluster mean silhouettes instead of one overall mean have been enough?Better, but still lossy. A cluster where half the points score 0.8 and half score -0.1 averages close to one where every point scores 0.35, and those are different problems. Per-cluster means catch a wholly broken cluster; only the distribution — the plot itself, or a summary of its spread — catches a cluster that is half solid and half misassigned.
saying these in an interview costs you the question
- Reports the overall average silhouette and stops there
- Reads negative bars as outliers instead of misassignment
- Assumes short bars mean the cluster has few points
- Treats 0.42 as a pass mark with no reference point
- Ignores the block heights, so cluster sizes go unnoticed