skip to content

How do you design a user-state taxonomy so its segments are mutually exclusive and exhaustive?

level: principalimportance: nice to knowfreq 26%

answer

  1. exactly one label per user
  2. at most one, and at least one
  3. one question, one point in time
  4. close the list with an explicit other
  5. overlapping tags are not a pie

basics

~20 s

Define states as answers to one question about a user at one point in time, so exactly one applies, and close the list with an explicit unknown bucket. Labels like paid plan and churned fail both tests.

solid answer

~50 s

A partition needs both properties at once: mutually exclusive, so no two states share a user, and collectively exhaustive, so every user lands somewhere. The practical recipe is to make the taxonomy the answer to a single question evaluated at a single timestamp, for example "what is this user's state as of the last day of the month?", and to close it with an explicit unknown or other bucket so coverage is guaranteed rather than hoped for. "Paid plan" and "churned" fail both tests: a churned user may well have been paying, and a free-trial user who is still active is neither. Only a genuine partition makes shares sum to 100%, makes segment counts safely additive, and makes conditional slices meaningful. When attributes genuinely overlap, such as feature-usage tags, keep them as independent flags and never present them as a pie.

go deeper

for a junior

Know the two definitions and be able to test a proposed pair of labels against both: can a user hold both, and can a user hold neither. Say partition when both conditions hold.

for a middle

Explain why only a true partition lets segment probabilities sum to one and lets counts be added safely, and show how a first-match-wins rule with a catch-all guarantees exactly one label per user.

for a senior

Diagnose a live taxonomy: find the overlapping or missing cases, anchor every state to an as-of date, add the unknown bucket, and add pipeline assertions so drift is caught rather than discovered in a review.

for a principal

Own the taxonomy as shared vocabulary: one definition with one owner, a deliberate choice between restating history and versioning when a state changes, and the discipline to keep the partition as small as the decisions require.

## The two properties, kept separate A collection of events `A1, ..., Ak` over a sample space `S` can have either property independently: - **Mutually exclusive** (pairwise disjoint): no two of them share an outcome. At most one occurs. - **Collectively exhaustive**: their union is all of `S`. At least one occurs. When both hold, the collection is a **partition** of the sample space, and exactly one event occurs on every trial. Only then does `P(A1) + P(A2) + ... + P(Ak) = 1` Many real taxonomies satisfy neither, which is precisely why so many dashboards show shares that do not total 100%. ## The canonical failing pair Consider labelling every user as "paid plan" or "churned". - **Not mutually exclusive**: someone who was on a paid plan and then cancelled satisfies both descriptions, depending on the tense you read them in. If "paid plan" means "has ever paid", the overlap is large. - **Not exhaustive**: a user in a free trial who is still active, or a free-tier user who has never paid and never left, is neither. The symptom in a report is unmistakable: two supposedly complementary segments whose shares total something other than 100%, and a stakeholder asking which number is real. And because they overlap, `P(paid or churned)` is *not* `P(paid) + P(churned)`; the sum double counts the churned payers. ## Fixing it: one question, one timestamp The reliable construction is to define the taxonomy as the possible answers to a **single question about one user at one moment**: > As of the last day of the reporting month, what is this account's state? Candidate answers: never activated, in trial, active free, active paid, dormant, cancelled. Because a user has exactly one state on that date, exclusivity is structural rather than a naming convention you hope people respect. Three design rules follow: 1. **Anchor to a timestamp.** "Churned" is meaningless without an as-of date, since churn is a transition, not an intrinsic property. Once the state is evaluated on a date, a user who paid in March and cancelled in April is "active paid" in the March cut and "cancelled" in the April cut, with no contradiction. 2. **Order the rules and make them total.** Implement the assignment as a first-match-wins sequence of conditions ending in a catch-all. That guarantees exactly one label per user by construction. 3. **Give unknown a home.** An explicit "unknown" or "other" bucket makes exhaustiveness real. Silently dropping unclassifiable users is the most common way a taxonomy becomes non-exhaustive without anyone noticing, and it biases every share upward. ## Mixing dimensions is the usual root cause Broken taxonomies almost always mix two questions into one list. "Paid plan" answers *how does this user pay*; "churned" answers *is this user still with us*. Two questions produce a grid, not a list. Keep them as separate dimensions, each individually a partition, and cross them when needed. The cross-product of two partitions is itself a partition, which is exactly why the discipline pays off. ## When a partition is the wrong model Not everything should be forced into exclusive buckets. Feature-usage tags, acquisition touchpoints and interest labels are genuinely overlapping: a user can hold several at once, and collapsing them into "primary tag" throws away information and invents an arbitrary priority order. The rule is about presentation as much as modelling. Overlapping attributes belong in a bar chart with a stated base, where each bar is a share of the total and the bars are not expected to add up. A pie chart, a stacked bar to 100%, or any "share of users" table is a claim that the categories partition the base, and using one for overlapping tags is a factual error, not a stylistic one. ## What a lead actually owns here - **One definition, one owner.** Segment definitions should live in a single modelled layer that every dashboard reads, not be re-implemented per query. Two teams with two definitions of active is a reporting failure that no amount of chart polish fixes. - **Automated invariants.** Assert in the pipeline that segment counts sum to the base and that no user appears twice. These are cheap tests and they catch taxonomy drift the moment a new state is introduced. - **Change management.** Redefining a state breaks every historical trend built on it. Decide deliberately whether to restate history or to version the taxonomy and show the break, and communicate which was chosen. - **Knowing when to stop.** Every extra state costs shared vocabulary. Prefer the smallest partition that supports the decisions being made, with the option to cross it against a second dimension for depth. ## Traps to avoid - Treating a set of tags as a partition because they were listed together. - Defining a state without an as-of date. - Letting unclassifiable users disappear instead of assigning them an explicit bucket. - Presenting overlapping labels as shares of a whole.

  • Why is "churned" ambiguous unless you attach an as-of date?
    Churn is a transition between states, not a permanent attribute, so a user can be active in one month's cut and cancelled in the next. Without an as-of date, two reports built from the same table can legitimately disagree, and the same user counts in two segments. Evaluating state on a fixed date makes exclusivity structural.
  • How do you handle users who fit none of your defined states?
    Give them an explicit unknown or other bucket rather than dropping them. Exhaustiveness is what makes shares total 100%, and a silent drop biases every reported share upward while hiding a data-quality problem. A visible unknown bucket that grows is a useful alarm; an invisible one is a slow corruption of every metric.
  • When should attributes stay overlapping instead of being forced into one partition?
    Whenever an entity genuinely holds several at once, such as feature-usage or interest tags. Collapsing them into a single primary label invents an arbitrary priority and destroys information. Keep them as independent flags, report each as a share of a stated base, and never present them in a pie or a stacked-to-100% chart.
  • What automated check catches a taxonomy that has quietly stopped being a partition?
    Assert two invariants in the pipeline: the segment counts sum exactly to the base population, and no entity identifier appears in more than one segment. The first catches lost exhaustiveness when a new state or filter is added; the second catches lost exclusivity when a definition broadens. Both are cheap and fail loudly at the source.

saying these in an interview costs you the question

  • Calling a list of tags a segmentation without checking overlap
  • Defining churned or active without an as-of date
  • Dropping unclassifiable users instead of bucketing them
  • Showing overlapping labels as a pie chart of users
  • Mixing two dimensions into a single flat list of states

context