skip to content

How do you count users reachable by email, push or SMS without double counting?

level: seniorimportance: should knowfreq 41%

answer

  1. singles minus pairs plus triple
  2. the all-three group nets to zero
  3. add it back exactly once
  4. or count who is reachable nowhere
  5. check every region is non-negative

basics

~20 s

Use three-set inclusion-exclusion: add the three channel reaches, subtract the three pairwise overlaps, then add the triple overlap back once. With 60k, 45k and 30k reachable, pairwise 20k, 12k and 9k, and 5k on all three, the union is 99k.

solid answer

~50 s

Inclusion-exclusion for three sets: `|E ∪ P ∪ S| = |E| + |P| + |S| - |E∩P| - |E∩S| - |P∩S| + |E∩P∩S|`. The triple overlap is added back because subtracting all three pairwise terms removes it one time too many. With email 60k, push 45k, SMS 30k, pairwise overlaps 20k, 12k and 9k, and 5k reachable on all three: `135 - 41 + 5 = 99k` distinct users, not the 135k a naive sum reports. Two practical points matter more than the algebra. First, every overlap must be computed on a single resolved identity key; if email and device records are not joined to the same user, the overlaps are unmeasurable and the union is a guess. Second, when the triple overlap is unknown, the union bound caps reach at 135k and the truncated sum gives a floor of 94k, so quote the interval rather than a point.

go deeper

for a junior

Know the three-set formula and apply it to given numbers without dropping or mis-signing a term. Remember the pattern: add singles, subtract pairs, add the triple.

for a middle

Explain why the triple term is added back by tracking one user through the count, and show the seven-region decomposition as a check that the supplied overlaps are consistent.

for a senior

Lead with the measurement problem: whether identities resolve across channels, how reachable is defined, and whether the windows align. When the triple overlap is missing, quote bounds rather than a fabricated point value.

for a principal

Own the reach definition across the organisation, decide whether identity resolution is worth building before reach numbers drive budget, and set the convention for publishing deduplicated reach with its uncertainty.

## The formula For three events or sets `E`, `P` and `S`: `|E ∪ P ∪ S| = |E| + |P| + |S| - |E∩P| - |E∩S| - |P∩S| + |E∩P∩S|` The same identity holds with probabilities in place of counts. The alternating signs continue for more sets: add all single terms, subtract all pairs, add all triples, subtract all quadruples, and so on. ## Why the last term is added back Track a single user who is reachable on all three channels. In the three single-channel terms they are counted three times. Each of the three pairwise overlaps also contains them, so subtracting all three pairs removes them three times, leaving a net count of zero. They exist and must be counted once, so the triple-overlap term adds them back. A user reachable on exactly two channels is counted twice by the singles, removed once by the one pairwise term containing them, and finishes at one, which is already right, and the triple term does not touch them. ## A worked example A growth team wants the number of distinct users reachable through at least one messaging channel. - Email: 60,000. Push: 45,000. SMS: 30,000. - Email and push: 20,000. Email and SMS: 12,000. Push and SMS: 9,000. - All three: 5,000. Apply the formula: `60 + 45 + 30 = 135` (thousands) `135 - (20 + 12 + 9) = 135 - 41 = 94` `94 + 5 = 99` So 99,000 distinct users are reachable. The naive sum of 135,000 overstates reach by 36%. A good sanity check is to decompose into the seven disjoint regions and confirm none is negative and the total matches: - All three: 5,000. - Email and push only: `20 - 5 = 15,000`. Email and SMS only: `12 - 5 = 7,000`. Push and SMS only: `9 - 5 = 4,000`. - Email only: `60 - 15 - 7 - 5 = 33,000`. Push only: `45 - 15 - 4 - 5 = 21,000`. SMS only: `30 - 7 - 4 - 5 = 14,000`. Those seven regions total `33 + 21 + 14 + 15 + 7 + 4 + 5 = 99,000`. If any region had come out negative, the supplied overlap numbers would be mutually inconsistent, which is a genuinely common finding when the three counts arrive from three different systems. ## The complement shortcut Often the easier route is the other side. If the total user base is 150,000 and 51,000 users are reachable on no channel, then reach is `150 - 51 = 99,000` immediately. Counting the complement avoids the alternating signs entirely and is usually the better query when a single table already carries all three flags. Reach for inclusion-exclusion when the channel counts arrive separately and no joined table exists. ## When you do not know all the overlaps Partial information still constrains the answer, and quoting bounds is a stronger answer than inventing a joint number. - **Upper bound (union bound)**: the union never exceeds the sum of the parts, so reach is at most 135,000. This needs no overlap information at all. - **Lower bound**: the union is at least the largest single channel, 60,000. If the pairwise overlaps are known but the triple is not, truncating inclusion-exclusion after the subtraction gives a floor of 94,000. In general, stopping after a subtracted term gives a lower bound and stopping after an added term gives an upper bound; these are the Bonferroni inequalities. So with everything except the triple overlap, honest reach is "between 94,000 and 99,000", since the triple overlap cannot exceed the smallest pairwise overlap of 9,000, and the arithmetic bounds it further. ## The part that actually goes wrong in practice The algebra is the easy half. The failure mode is identity. - **Overlaps require a joined key.** "Email reach" is a count of addresses, "push reach" a count of device tokens, "SMS reach" a count of phone numbers. Unless all three resolve to one user identifier, the intersections are not measurable and any union you report is an assumption dressed as a number. - **Reachable must be defined once.** Does a valid address count, or a deliverable one, or one that is opted in and not bounced? Different definitions can move each term more than the overlap correction does. - **Duplicates within a channel** inflate a single term before the overlaps are even considered, so deduplicate inside each set first. - **Time alignment.** If email reach is measured over the last 30 days and push over the last 7, the overlap is not comparable to either. A senior answer states the formula in a sentence and spends most of its time on identity resolution, the definition of reachable, and the sanity check for negative regions. ## Traps to avoid - Subtracting the triple overlap instead of adding it, which undercounts by exactly that amount. - Subtracting only the largest pairwise overlap and treating the rest as noise. - Presenting the naive sum as "total reach" in a plan or forecast. - Accepting overlap counts that imply a negative region without questioning the source.

  • Why is the triple-overlap term added rather than subtracted?
    A user on all three channels is counted three times by the single-channel terms, then removed three times by the three pairwise terms, netting zero. Adding the triple overlap back restores them to a count of one. Users on exactly two channels are already correct after the pairwise subtraction, so the final term leaves them untouched.
  • You know each channel's reach but none of the overlaps. What can you still report?
    A range. The union is at most the sum of the three reaches, by the union bound, and at least the largest single channel, since the union contains each set. With 60k, 45k and 30k that is between 60k and 135k. Report the interval and the missing joint data rather than a point estimate that quietly assumes disjointness.
  • When would you count the complement instead of applying inclusion-exclusion?
    Whenever a single table already flags each user per channel. Counting users reachable on no channel and subtracting from the base gives the union in one pass, with no alternating signs and no chance of a sign error. Inclusion-exclusion earns its place when the counts arrive from separate systems and no joined per-user view exists.
  • The supplied overlaps imply a negative region. What does that tell you?
    The inputs are mutually inconsistent, not merely imprecise. Decomposing into the seven disjoint regions is a hard constraint: every one must be non-negative. A negative region means at least one count is wrong, most often because the three numbers were measured over different windows, with different definitions of reachable, or without a shared identity key.

Three overlapping spotlights on a stage: adding their areas counts the doubly-lit patches twice and the centre patch three times, so you subtract the overlaps and then restore the centre you erased.

saying these in an interview costs you the question

  • Subtracting the triple overlap instead of adding it back
  • Reporting the naive sum of channel reaches as total reach
  • Computing overlaps without a resolved shared user identifier
  • Assuming channels are disjoint because overlap data is unavailable
  • Accepting overlap counts that imply a negative region

context