skip to content

Why does a multi-tenant spam filter cut its null-rate and score series per tenant instead of watching the aggregate?

level: seniorimportance: should knowfreq 44%

answer

  1. averages hide small segments
  2. a mixture moves when weights move
  3. 0.5% of traffic, half a point global
  4. each slice against its own history
  5. volume floor plus a remainder bucket

basics

~10 s

Aggregates fail in both directions: a small tenant's total breakage is diluted into noise, and the global curve can move purely because the mix of traffic changed while no tenant's mail changed at all.

solid answer

~40 s

Volume-weighting hides small segments. A tenant at 0.5% of 40 million messages a day is 200,000 messages; if its enrichment breaks completely, the global null rate moves from about 1% to about 1.5% - half a point, inside ordinary daily variation - while every one of that customer's messages is being scored blind. The opposite error is just as common: the aggregate score histogram can shift with **no** tenant's distribution moving, simply because a large tenant's nightly batch changed the mix. Slicing also buys attribution, because the cuts that matter - tenant, inbound path, language, model version - tend to map one-to-one onto the upstream that broke. The cost is series count, so slice by dimensions that name a cause and put tenants below a volume floor into a watched remainder bucket.

go deeper

for a junior

Remember that a global average is weighted by volume, so a small customer can be completely broken while the overall number barely twitches.

for a middle

Work the dilution arithmetic out loud, and explain the opposite case too: an aggregate curve is a mixture, so changing the mix moves it even when no component moved.

for a senior

Name the cuts that map onto causes - tenant, inbound path, language, model version - and say how you bound the series count with a volume floor and a watched remainder bucket.

for a principal

Own the trade-off between coverage and readability: past some slicing depth every extra dimension buys noise, and decide which segments are worth their own reference and which are covered in aggregate.

## Two opposite failures of the aggregate A single global series is volume-weighted by construction, and that produces two symmetric errors. **Dilution - a real break that the aggregate cannot show.** Take a mail service scoring 40 million messages a day and a tenant that is 0.5% of that, so 200,000 messages daily. Suppose the enrichment feeding the sender-reputation feature fails for that tenant only, taking its null rate from about 1% to 100%. The global series moves to `0.995 x 1% + 0.005 x 100%`, which is about **1.5%** against a baseline of 1%. Half a percentage point is the kind of wobble a busy global series produces on an ordinary Tuesday, and no sensible threshold on it would fire - yet that customer's mail is being filtered with its strongest feature missing. **Mix shift - an aggregate move with nothing wrong underneath.** The global score histogram is a mixture of per-tenant histograms weighted by volume. Change the weights and the mixture changes even if every component is identical to yesterday: a large tenant starts a nightly bulk send, a region wakes up, one customer onboards a new domain. The curve moves, the investigation starts, and there is nothing to find inside any tenant. This is the case that burns credibility fastest, because the signal was real and the conclusion was not. ## Segments are where causes live The cuts worth carrying are the ones that map onto something that can break: - **Tenant** - a per-customer configuration, a policy flag, an onboarding change. - **Inbound path** - usually a distinct upstream and a distinct enrichment chain, so a failure lands on exactly one path. - **Language or region** - text features behave differently, and a change often arrives in one locale first. - **Model version** - so a step change starting exactly at a rollout boundary is attributable instead of being confounded with the world changing that day. When a global curve moves, the first useful action is to re-cut the *same* signal along these dimensions. If the move is confined to one, the cause is usually within one dependency of it. If the move is present everywhere in proportion, it is more likely genuinely global - or a mix change, which the per-slice curves will tell you apart because they stay still. ## Each slice needs its own reference Tenants legitimately differ. A newsletter house's score distribution sits permanently higher than a law firm's; a transactional-mail tenant has almost no mass in the high-score tail at all. Reading every tenant against the global reference therefore produces a standing false signal on the tenants that are merely different from average. Each slice is read against **its own** prior behaviour, which is also what makes a small tenant's break visible: against its own history, going from 1% nulls to 100% is unmistakable regardless of how little volume it carries. ## The cost, and how to bound it Slicing multiplies series: tenants times languages times inbound paths times model versions grows faster than anyone expects, and most of the resulting series are too thin to read. Two rules keep it bounded: 1. **A volume floor.** Below some messages-per-window, a slice's null rate and histogram are dominated by sampling noise and will manufacture movement that is not there. Tenants below the floor go into a **remainder bucket** which is itself watched - so they are covered in aggregate even when they are not covered individually. 2. **Cut by cause, not by curiosity.** Add a dimension because a failure can be localised to it, not because the field exists. A cut that has never once named a cause is series you pay for forever. | approach | catches a small tenant's break | survives a traffic-mix change | series cost | |---|---|---|---| | global series only | no | no - it moves with the mix | lowest | | per-slice, no volume floor | yes | yes | highest, and noisy | | per-slice above a floor, plus remainder bucket | yes, for tenants that matter individually | yes | bounded | ## Two questions the cut answers at once The per-slice view is what lets you say both "how bad is it" and "who is affected" in the same chart. A two-point fall in a global blocked rate is unreadable as harm until you know it is a near-total collapse confined to a fifth of traffic. Slicing converts an averaged-away number into a statement about a named set of customers - which is also the form in which the consequence has to be reported to anyone outside the team.

  • Why does each slice need its own reference rather than being compared to the global curve?
    Tenants legitimately differ - a newsletter sender's score distribution sits permanently higher than a transactional sender's. Comparing each to the global average produces a standing false signal on every tenant that is simply not average, and buries the one that actually changed.
  • How do you stop the number of series exploding as tenants multiply?
    Cut by dimensions that can name a cause - inbound path, language, model version - and by tenant only above a volume floor. Everyone below the floor goes into a remainder bucket that is watched as a group. A slice too thin to read is noise, not coverage.
  • The global score histogram moved but every per-tenant histogram held steady. What does that mean?
    The mixture weights changed, not the components: the traffic mix across tenants shifted, for instance a large tenant starting a bulk send. Nothing inside any tenant moved, and chasing the aggregate further will find nothing.

saying these in an interview costs you the question

  • Treats a flat global series as evidence that every tenant is healthy.
  • Investigates an aggregate move without first re-cutting it by segment.
  • Compares every tenant against the global reference distribution.
  • Adds a slice for every available field and drowns in thin series.
  • Reads a tiny slice's noisy series as a real movement.