An inbound spam filter's blocked rate falls from 12% to 10% overnight with no deploy and no errors, while one feature's null rate jumps to 22% - what happened?
answer
- green health checks, wrong answers
- a default is a legal value
- one fifth of traffic lost its evidence
- confident tail thins, modes collapse
- split the histogram by present versus defaulted
basics
~20 sAn upstream feature source degraded and its missing values are being filled with a default, so about a fifth of messages are scored without their strongest evidence. Nothing errors, because a default is a legal value, and detection quietly falls.
solid answer
~40 sThe sender-reputation lookup started timing out, and the feature pipeline substitutes a neutral default on a miss. Scoring therefore succeeds for every message - there is no error to count and no request to retry - but for the affected 22% of traffic the model has lost the feature that pushed confident spam to the top of the range. Their scores compress toward the middle, fewer cross the quarantine cutoff, and the blocked rate falls about two points overall. Inside the affected slice the drop is far larger: roughly 12% blocked down to about 3%. The confirmation is one chart - the score histogram cut by whether the feature was present or defaulted - and the matured verdicts that would have proved a quality regression are still days away.
code
pseudocode · 18 lineson_message(msg):
rep = lookup(reputation_source, msg.sender_domain, timeout_ms = 50)
if rep is null: # timeout or miss - not an error
features.sender_reputation = NEUTRAL_DEFAULT
count("feature_defaulted", feature = "sender_reputation",
tenant = msg.tenant, path = msg.inbound_path)
present = false
else:
features.sender_reputation = rep
present = true
score = score_with(model_version, features) # always returns a number
observe("score_histogram", score, tenant = msg.tenant, present = present)
if score > QUARANTINE_CUTOFF:
count("blocked", tenant = msg.tenant)
return scorego deeper
Remember the pattern: a missing feature is usually filled with a default, the request still succeeds, and the model quietly answers with less evidence. Green health checks do not mean correct answers.
Explain the chain from a timed-out lookup through the substituted default to compressed scores and fewer messages crossing the quarantine cutoff, and say why no error counter moves anywhere along it.
Work the numbers - a two-point global fall implies a near-collapse inside the affected fifth - and name the one chart that proves the link: the score histogram split by feature-present versus defaulted.
Take the position on defaulting itself: keeping the default protects delivery, so the obligation is to make every substitution a first-class monitored series rather than an invisible runtime convenience.
## The shape of a silent failure Nothing in this incident looks like a failure from the outside. Every scoring request returns a score. The enrichment call has a short timeout and a null fallback, so end-to-end latency stays inside its budget. No exception is thrown, no retry fires, no error counter moves. The only things that moved are two label-free series: a feature's **null rate**, from well under 1% to 22%, and the **blocked rate**, from 12% of scored messages to 10%. The mechanism is the substitution. A missing feature value is filled with a default - a neutral reputation, a training mean, a zero - and that default is a perfectly legal input. The model consumes it and returns a confident-looking number. **A feature that fails to a legal default degrades the model without degrading the service**, which is precisely the failure class that request-level health checks cannot see. ## Checking the arithmetic The two numbers have to be consistent with each other, and they are: - 22% of traffic lost the feature; 78% was untouched and keeps blocking at its usual 12%. - Overall blocked rate: `0.78 x 12% + 0.22 x b = 10%`, so `b` is about **2.9%**. - So inside the affected fifth, blocking fell from roughly 12% to roughly 3% - a near-collapse - which averages out to the modest-looking two-point dip in the global series. - On 40 million messages a day that is 4.8 million quarantined falling to 4.0 million: about **800,000 spam messages delivered per day** that would previously have been held. That gap between how small the aggregate looks and how large the local effect is, is the entire lesson of this incident, and it is the reason the same series are also emitted per segment. ## What the score histogram shows A healthy spam filter's score curve is strongly bimodal: a large mass of clean mail near zero and a distinct tail of confident spam near one. When a high-signal feature is replaced by a constant for a fifth of traffic, the affected messages lose the evidence that separated the modes, so their scores **compress toward the middle** and the confident tail thins. The mean score may barely move while the shape has changed completely - which is why the histogram is kept as buckets, not as a mean. The decisive chart is not the global curve but the same histogram **cut by whether the feature was present or defaulted**. Two populations, one keeping its normal bimodal shape and one collapsed into the middle, is a picture no other explanation produces. ## The order of investigation 1. **Which series moved, and when.** The null rate names the feature and timestamps the break to the window. 2. **Where it lives.** Cut null rate and score histogram by tenant, inbound path and language. A cause that is one upstream dependency usually maps onto one of those cuts. 3. **Confirm the link.** Split the score histogram by feature-present versus defaulted; if the collapsed population is exactly the defaulted one, the chain is closed. 4. **Quantify the consequence.** Convert the blocked-rate change into messages per day, per tenant, using the numbers above. Notice what is *not* in that list: waiting for outcomes. The verdicts that would demonstrate a real precision or recall regression are hours to days away, and by the time they matured, a full day of mail would already have been delivered. ## Why the obvious readings are wrong | reading | why it fails here | |---|---| | "spam volume simply fell" | the total scored volume did not move; only the share crossing the cutoff did | | "the model drifted, retrain it" | the model is a fixed artifact and has not changed; retraining on defaulted inputs would bake the failure in | | "health checks are green, so scoring is fine" | a defaulted feature produces a valid score, so nothing about the request looks unhealthy | | "someone changed the cutoff" | a cutoff change moves the blocked rate but cannot raise a feature's null rate | ## The design conclusion Defaulting a missing feature is usually the right runtime behaviour - failing every message because an optional enrichment timed out converts a quality loss into an outage. The design obligation is to make the substitution **visible**: emit the default rate per feature as a first-class series, and carry a feature-present flag into the slice keys so a degraded input is legible in the charts rather than hidden inside a healthy-looking score. An input that can silently become a constant is not a monitoring detail; it is the most common way an ML system goes quietly wrong.
- Why does the per-feature null rate close this investigation faster than the score histogram does?The histogram tells you that something moved; the null rate tells you which input moved and in which window. Detection and attribution are separate jobs, and the null rate is the cheapest attribution signal there is because it points at one named upstream dependency.
- The default is what hid the failure - should the scorer reject the message instead?Usually not. Rejecting every message when an optional enrichment times out trades a partial quality loss for a full outage on a delivery path. The fix is visibility, not failure: emit the default rate as a first-class series and carry a feature-present flag into the slice keys so the degraded population can be charted separately.
- Would the aggregate blocked-rate series have been enough to catch this?Only barely, and only because a fifth of traffic was affected. A two-point move can sit inside ordinary daily variation, and if the same failure had hit one small tenant the global series would not have moved at all. The per-slice cut is what makes the size of the break visible.
saying these in an interview costs you the question
- Blames the model and asks for a retrain before checking input signals.
- Says flat error rates and latency rule out a scoring problem.
- Treats a defaulted feature value as equivalent to a row that was skipped.
- Waits for matured verdicts to confirm a regression that is visible now.
- Reads the blocked-rate fall as inbound spam volume falling.
- Assumes a steady mean score means the distribution's shape held.