Which signals show a priority inbox's retraining loop is narrowing its own training data before recipients notice missed mail?
answer
- watch composition, not service health
- diff successive training snapshots per class
- exposures and opens tracked separately
- confidence up, live quality flat
- randomised slice is the only control
basics
~20 sCompare successive training snapshots: a sender class whose share of positives collapses refresh after refresh, promotions concentrating into fewer classes, and rising model confidence while the live open rate on the promoted view stays flat. The randomised promotion slice is the control that says which it is.
solid answer
~50 sWatch the retraining pipeline's own artifacts, not just the serving dashboards. Per refresh, report the share of positives and the count of exposures for each sender class, then diff that composition against the previous snapshot: a class sliding toward zero positives across two or three refreshes is the signature. Alongside it, track concentration — what share of promotions the top few classes take — and the gap between an improving offline number and a flat live open rate on the promoted view, which says the model is agreeing with itself on data it selected. None of that separates a loop from a real change in recipient behaviour; the randomised promotion slice does. If a class's open rate on randomly promoted mail is steady while its share of training positives collapses, the narrowing is the pipeline's doing.
code
pseudocode · 15 linesfor each sender_class c:
share_prev = positives(snapshot[n-1], c) / positives(snapshot[n-1])
share_now = positives(snapshot[n], c) / positives(snapshot[n])
if share_now >= 0.5 * share_prev:
continue // composition held; nothing to read
// the class is collapsing in the training data - ask the control why
seen = promotions(random_slice[n], c)
if seen < MIN_OBSERVATIONS:
flag(c, "exploration share too small to judge")
else if opens(random_slice[n], c) / seen >= 0.8 * rate_prev(c):
flag(c, "suppressed by placement - interest is steady")
else:
flag(c, "genuine drop in recipient interest")go deeper
Know that a model can get quietly worse with no errors and no latency change. The symptom lives in which mail the system still learns about, not in service health.
Explain why exposures and opens must be tracked separately per sender class: falling opens with falling exposures is censoring, while falling opens with steady exposures is a real behaviour change.
Design the report the refresh emits, diff it against the previous snapshot, and state plainly that a metric computed on the same censored traffic cannot detect this. Bring the randomised slice in as the control.
Decide what the flag is allowed to do — warn, hold the refit, or block promotion — and who is accountable for the exploration share that keeps the control readable in the first place.
## What you are actually looking for A self-reinforcing cycle is not an outage. Error rates stay flat, latency is unchanged, the refresh completes, and the offline number often improves. What changes is the **composition of the evidence**: the range of mail the system still learns anything about shrinks, refresh after refresh. So the detection has to sit where the composition lives — in the training snapshots and in the placement log — rather than in service health. ## Signals inside the retraining pipeline Emit these as part of every refresh, and compare each to the previous refresh rather than to an absolute threshold: - **Share of positives per sender class.** A class whose share halves, then halves again, is the signature. A one-off dip is noise; a monotone slide across three refreshes is not. - **Count of classes carrying any meaningful positives.** If the number of sender classes contributing more than a handful of opens falls each refresh, the model is learning about a narrower world than last month. - **Exposures per class, separately from opens.** This is what distinguishes the cause: a class whose opens fell *because its exposures fell* is being censored, while a class whose exposures held and opens fell is a genuine behaviour change. - **Concentration of promotions.** The share of promotion slots taken by the top few sender classes, tracked refresh over refresh. - **Score distribution of the refit.** Predictions piling up at the extremes usually means the refit found a cleanly separable training set — which is what a censored one looks like. ## Signals in the live system - **An improving offline number beside a flat promoted-view open rate.** The refresh looks better every month and the recipient's experience does not. If the evaluation split was carved out of the same censored traffic, the offline number is largely measuring agreement with the previous model. - **Rising rescue actions.** Recipients digging messages out of the secondary view, or marking demoted senders important by hand, is a paid-for signal that the demotion is wrong. It is a lagging one: by the time it is visible in aggregate the loop has been running for refreshes. - **Complaint and support volume**, which lags further still and arrives already attributed to the wrong cause. ## The randomised slice is the control Every signal above is ambiguous on its own, because a class's positives can also fall for the honest reason that recipients stopped caring. The small randomised share of promotions settles it: those messages were exposed without reference to the score, so their open rate is a measurement of interest that the ranker did not shape. | Training-set share of the class | Open rate on the randomised slice | Reading | |---|---|---| | collapsing | steady | a self-reinforcing cycle — the pipeline suppressed the class | | collapsing | falling too | a genuine drop in recipient interest | | steady | falling | a change upstream of placement — content, senders or the label definition | | collapsing | too few observations to read | the exploration share is too small to protect this class | That last row matters as much as the others: a control with no data is not a control, and it is the honest trigger for revisiting the exploration share. ## Wiring it into the refresh 1. At the end of each refresh, write a composition report for the snapshot it was fitted on: per class, exposures, positives, share of positives, and the same three from the randomised slice alone. 2. Diff it against the previous refresh's report and flag any class crossing a relative drop threshold — the flag is on the **relative move across refreshes**, not on a fixed level, because absolute class sizes differ by orders of magnitude. 3. Route the flag somewhere a human reads before the refit becomes the live ranker, and hold the flagged classes' exploration share fixed while the question is open. ## The standing trap The most common way this is missed is measuring the refresh against a held-out split of the same logged traffic. The split inherits the censoring exactly, so the metric improves as the loop tightens and the report reads as a success. Any number intended to catch this has to be computed on exposure the ranker did not choose, or on outcomes the recipient produced without the ranker's help.
- Why is a fixed threshold on a class's share of positives a poor alarm?Sender classes differ in size by orders of magnitude, so a level that is alarming for a large class is normal for a small one, and any single threshold either pages constantly or never fires. Alarm on the relative move between consecutive refreshes instead, and require the move to persist across two or three of them before it counts.
- The refit's offline number improves every month. What is that number likely measuring?Agreement with the previous ranker's choices. When the evaluation split is carved from the same logged traffic, it carries the same placement censoring, so both the fit and the yardstick have been narrowed together. A reading that survives this has to be computed on the randomised slice or on outcomes the recipient reached without the ranker's placement.
- Recipients are rescuing more mail by hand from the secondary view. How useful is that signal?Directionally very useful and operationally late. A rescue is a recipient paying effort to contradict the ranker, so it is strong evidence the demotion is wrong for that sender class. But it only appears once the demotion is already common and annoying, and the volume is far too small to re-seed a training set — treat it as confirmation, not as the detector.
saying these in an interview costs you the question
- Looks only at serving dashboards, where nothing is erroring
- Trusts an offline number computed on the same censored traffic
- Alarms on an absolute class share instead of the refresh-to-refresh move
- Reads falling opens as lost interest without checking exposures
- Waits for complaints, which arrive refreshes after the narrowing