skip to content

How do you profile and name customer segments once a clustering run has produced them?

level: middleimportance: must knowfreq 64%

answer

  1. coordinates are not a description
  2. segment mean against the overall mean
  3. always carry the size column
  4. check variables the clustering never saw
  5. name from the two or three biggest gaps

basics

~20 s

Build a profile table: each segment's mean on every feature indexed against the overall mean, plus headcount and revenue share. Name each segment from the two or three dimensions where it differs most from the base.

solid answer

~50 s

Profiling turns coordinates into a description. I build a table with one row per segment and one column per variable, holding the segment's mean or median alongside an index against the population mean, so a reader sees instantly that a segment spends 2.4x the base average and visits half as often. Two columns are non-negotiable: headcount and share of revenue, because a striking profile over 0.4% of the base is a curiosity, not a segment. I always report the numbers in original units - dollars, sessions, days - not in whatever transformed space the clustering ran in. I also profile on variables that were *not* inputs, such as tenure, support tickets or acquisition channel, since agreement there is independent evidence the split tracks something real. Then each segment gets one sentence a marketer can repeat from memory, named for behaviour rather than `Cluster 3`.

go deeper

for a junior

Be ready to say that a centroid is a set of coordinates and needs translating into a description. Knowing that you compare each segment against the overall average, and that headcount belongs in the table, is enough at this level.

for a middle

Explain the mechanics: index-to-base, original units, median versus mean on skewed columns, categorical composition against base rates. Expect to be asked why profiling on non-input variables is worth doing.

for a senior

Demonstrate that you treat the profile as the deliverable rather than the model. Show how you catch a partition that is really a one-dimensional band, and how you decide a segment has no viable action and say so.

for a principal

Own the naming and the taxonomy across teams. Segment names become the organisation's shared vocabulary and end up in dashboards, targets and incentives, so argue for names that survive a re-fit and a governance owner who controls changes to them.

## What profiling is for A clustering run returns group assignments and a centroid per group - a vector of coordinates. Nobody can act on a vector of coordinates. Profiling is the step that converts each group into a described, sized, named thing that a marketer, a product manager or an executive can hold in their head and act on. It is where most of the business value of a segmentation is actually created, and it is the step most often rushed. ## The profile table The core artefact is one table: rows are segments, columns are variables, cells are the segment's central value on that variable. Three refinements make it readable: **Index to the base.** A raw mean carries no meaning on its own. Report alongside it the ratio of the segment mean to the overall population mean (an index where 100 means average, 240 means 2.4x the base), or the gap expressed in standard deviations. This is what lets a reader scan the row and immediately see which two or three numbers are the story. **Original units.** If the inputs were transformed before clustering, translate the centroid back before showing it. `-0.8` is unreadable; `$41 per order` is not. The transformed space is an implementation detail of the fitting step and should not leak into the deliverable. **Robust central values for skewed columns.** Spend is heavily right-skewed, so a segment mean can be dragged by a few large accounts and describe nobody. Show the median next to the mean for such columns, and for categorical variables show the composition - the share of the segment in each category against that category's share of the whole base. ## Size is part of the profile Every profile carries headcount, share of the base, and share of the business outcome that matters (revenue, margin, sessions). A segment that indexes at 400 on spend but holds 0.3% of customers is a very different object from one indexing at 140 over a third of the base, and the profile numbers alone do not distinguish them. Sizing is also what makes the deliverable honest: it is common for the most interesting-looking segment to be the least worth pursuing. ## Profile on variables the clustering never saw If the partition was built from purchase behaviour, profile it on tenure, acquisition channel, support-contact rate, region and product mix as well. This does two things. It provides an independent check that the split tracks something real rather than an artefact of the chosen inputs - if six segments look identical on every variable outside the input set, be suspicious. And it usually supplies both the memorable name and the practical hook: knowing a segment is 70% mobile-app-only tells the campaign team how to reach it, and no clustering input said that. ## Naming A good segment name is behavioural, mutually exclusive in meaning, and instantly repeatable. `Champions`, `At-Risk`, `Hibernating` work because each implies a different action. Rules that hold up: - Never ship `Cluster 0` through `Cluster 5`. Numbers carry no meaning and will be re-ordered by the next run. - Name from the two or three dimensions where the segment is most distinctive, not from one. A name built on a single variable describes a band, not a segment, and invites the reasonable question of why anyone clustered at all. - Avoid judgemental names. `Bad customers` cannot be said in a review, and it hides whether the segment is low-value, high-cost, or simply new. - Keep the set of names mutually intelligible: if two names could describe the same customer, the profile is not distinguishing them and the reader will not trust either. ## A useful sanity check If every segment sits near the population average on every variable except one, the partition is essentially a one-dimensional split of that variable. The honest response is often to replace the model with an explainable banding rule, which is cheaper to maintain and easier to defend. ## The deliverable What ships is not the model. It is: the profile table; one sentence per segment ("Champions - bought within 30 days, five-plus orders a year, top-decile spend; 8% of customers, 34% of revenue"); the sizes; and a recommended action per segment. If a segment cannot be given a recommended action, say so explicitly rather than letting the reader assume one exists.

  • Why profile segments on variables that were not clustering inputs?
    Because agreement there is independent evidence. If groups built from purchase behaviour also separate cleanly on tenure, support-contact rate or acquisition channel, the split is tracking something real rather than an artefact of the inputs chosen. Those outside variables usually supply the memorable name and the practical reach mechanism as well, since knowing a segment is mostly app-only tells the campaign team how to contact it.
  • Your six segments sit near the population average on everything except one variable - what does that mean?
    The partition is effectively a one-dimensional split on that variable, and the remaining dimensions are contributing noise rather than structure. The honest response is to say so: an explicit banding rule on that one variable gives the same groups, is reproducible, is far cheaper to maintain, and does not need re-fitting. Reserve clustering for when the groups genuinely need several dimensions to describe.
  • How do you present a segmentation to a marketing team rather than to analysts?
    One page per segment: the name, a one-sentence description in original units, headcount and revenue share, the two or three ways it differs from the base, how to reach it, and a recommended action. No feature-space language, no coordinates, no fitted-model detail. The test is whether someone can repeat a segment's description accurately a week later without the deck in front of them.

A profile table is a scouting report. The raw numbers matter far less than how each player stands against the league average, and the one-line summary is what the coach actually remembers.

saying these in an interview costs you the question

  • Ships segments labelled Cluster 0 through Cluster 5
  • Reports centroids with no headcount or revenue share
  • Reads segment means without comparing to the base average
  • Presents coordinates in transformed units nobody can interpret
  • Names a segment from one feature and ignores the rest
  • Uses the mean on a heavily skewed spend column and calls it typical

context