One of your six customer segments holds 37 of 400,000 customers - what do you do with it?
answer
- 37 people is not a campaign audience
- small enough to read row by row
- test accounts, duplicates or a broken field
- minimum size comes from channel economics
- document it rather than deleting it quietly
basics
~20 sInspect those 37 records individually first - a group that small is usually outliers, test accounts or a data fault. Either way, no campaign can economically address 37 people, so absorb them into the nearest segment or document them as exclusions.
solid answer
~50 sA group of 37 out of 400,000 is almost never a segment; it is the partition isolating extreme points. So I diagnose before I decide: 37 rows is small enough to look at one by one. Common causes are internal or test accounts, duplicated records, a pipeline fault producing impossible values, or a genuinely different business line such as wholesale buyers sitting inside a retail base. If it is bad data, fix it upstream and re-run. If it is real but extreme, those customers may deserve individual handling rather than a segment. What I will not do is ship it as one of six segments, because a segmentation exists to route action and the smallest useful segment is the smallest audience your cheapest channel can economically serve. Whatever I choose, the deliverable states what happened to those 37.
go deeper
Be ready to say that a group of 37 out of 400,000 is not something a campaign can use, and that the first move is to look at the records rather than to guess. Naming test accounts and data errors as likely causes is enough here.
Explain why distance-based partitions isolate extreme points in the first place, and walk through the diagnosis checklist. Be able to compare the options - fix upstream, exclude, absorb, keep a documented Other bucket - rather than naming only one.
Show that minimum segment size comes from channel economics and measurability, decided before you see the output. Demonstrate the judgment to recognise when a tiny group is instead a high-value named-account list worth escalating.
Own the rule the organisation applies. Decide who arbitrates when a team wants to target a group too small to measure, and make sure exclusions are documented and reconcilable so segment headcounts still add up to the customer base.
## Why tiny groups appear Partitioning methods that minimise within-group distance are structurally happy to spend a whole group on a few far-away points, because those points contribute enormous distance wherever else they are placed. A handful of records with values orders of magnitude beyond the rest - a corporate account buying at wholesale volume inside a consumer base, an internal test account with 40,000 orders, a row where a currency field was recorded in cents - is exactly the shape that gets isolated. So the first read on a 37-out-of-400,000 group is not "we found a rare and valuable niche", it is "something over there is far from everything else, and I should find out what". ## Diagnose before deciding Thirty-seven rows is a gift: you can read them. Work through a short checklist. - **Are they real customers?** Internal accounts, QA accounts, employee accounts and integration test records are the single most common answer. - **Are the values physically possible?** A purchase count that exceeds the number of days in the window, a negative recency, a spend figure a thousand times the next-largest - these say the pipeline is wrong, not the customer. - **Are they duplicates?** One entity split across many records, or many records collapsed into one, both produce extremes. - **Are they a different business?** Wholesale buyers, resellers, corporate accounts and partner integrations often live in the same table as retail customers and genuinely do not belong in the same segmentation. Only after that question is answered does the decision about what to do with them make sense. ## What "too small" actually means Minimum viable segment size is a **business** quantity, not a statistical one. It is set by the economics and reach of the channel that would act on the segment, plus the need to be able to read a result afterwards: - An email or in-app campaign can address a few hundred people economically, but you will struggle to measure whether it worked at that size. - A field-sales or account-management motion can address five people, profitably, so five is not too small there. - A product change gated on a segment needs enough volume to justify building and maintaining the branch. So the same 37 customers can be unusable for marketing and entirely usable for account management. Decide the threshold from the intended action *before* looking at the partition, so the number is not reverse-engineered to justify whatever came out. ## The options 1. **Fix and re-run.** If the group is a data fault, correcting it upstream is the only answer that also improves everything else built on that table. 2. **Exclude and handle by name.** Remove the records from the clustering population, document why, and give them explicit treatment - often a list handed to a human rather than a campaign. 3. **Absorb into the nearest segment.** Acceptable once you know what they are, and only then: merging a data-fault group hides the bug and drags the receiving segment's profile. 4. **Keep an explicit `Other` bucket.** Legitimate if it is named, sized and documented rather than quietly dropped. Every real segmentation has customers who fit nowhere; pretending otherwise is worse than admitting it. What rarely helps is re-running with a different number of groups purely to make the small one disappear. The extreme points do not go away; they just get attached to a neighbour whose profile they then distort. ## The exception worth knowing Sometimes the 37 are the most valuable output of the whole run. If they carry a large share of revenue or risk - 37 accounts worth a fifth of the business - then surfacing them was the run doing its job. They still are not a marketing segment; they are a named-account list handled one-to-one, and the right deliverable is that list plus the reason they are separate. Recognising this distinction, rather than reflexively discarding anything small, is what separates a senior answer from a rule-following one. ## Report it either way Whatever you decide, the segmentation deliverable carries the headcount of every group and an explicit note about the 37: what they were, what you did, and what the downstream consumer should expect. A quietly deleted group is a surprise waiting for whoever reconciles segment headcounts against the customer table.
- When is a 37-member group the most valuable thing the run produced?When those 37 carry a disproportionate share of revenue or risk. Thirty-seven accounts worth a fifth of the business are worth knowing about individually, and the clustering has done its job by surfacing them. The treatment changes rather than the value: they become a named-account list handled one-to-one by a human, not a marketing segment, and the deliverable explains why they sit outside the main partition.
- Would you just merge them into the nearest segment and move on?Only after diagnosing them. Merging first is how a data fault gets laundered into a segment profile: the receiving group's spend average shifts, its one-line description stops matching its members, and the underlying pipeline bug survives into the next run. Diagnose, then merge if they are genuine, and note the merge in the deliverable so the profile shift is explainable.
- How do you set the minimum usable segment size before you start?Work backwards from the intended action. Ask which channel would act on the segment, what it costs per contact, and how large an audience is needed to read a result at all. That gives a threshold in people, decided before the partition exists so it cannot be reverse-engineered to justify the output. Different actions justify wildly different thresholds within the same company.
saying these in an interview costs you the question
- Ships a 37-person group as one of the marketing segments
- Assumes any small group is noise without inspecting it
- Merges the small group in before diagnosing what it is
- Deletes the rows without documenting the exclusion
- Treats minimum segment size as a statistical rather than business question