During a scorer outage the suggestion strip still fills from the frequency list and no errors fire — how would you see the quality loss?
answer
- the fallback returns a success
- error rate never moves
- tag the rung on every response
- split acceptance rate by rung
- alert on the degraded share
basics
~20 sTag every rendered strip with the rung that produced it and the reason, then split acceptance rate by rung. A blended number absorbs the loss; the share of keystrokes served below the top rung is the signal that actually moves.
solid answer
~40 sThe fallback is a success path — it returns a filled strip — so error-rate and availability monitors stay flat through the whole incident. The only thing that moves is the mix of rungs, which means the mix has to be a first-class signal, not a debugging detail. I would tag each response with its rung and the reason it was chosen, then carry two measurements: the **share of keystrokes served below the top rung**, alerted on directly, and **acceptance rate per rung**, which says what that share costs. A single blended acceptance rate is the trap: it averages two populations, so a big loss on a small slice shows up as a fraction of a point and reads as noise. The same tag also lets downstream work exclude or down-weight degraded rows.
code
json · 9 lines{
"prefix_length": 4,
"rung": "frequency_list",
"rung_reason": "scorer_deadline_exceeded",
"feature_groups_present": ["session", "locale"],
"feature_groups_absent": ["user_history"],
"candidates_rendered": 3,
"accepted_index": null
}go deeper
Know that a fallback returns a normal successful response, so 'no errors' does not mean nothing went wrong on the prediction path.
Explain why one blended quality number averages two populations, and what a per-rung split plus a mix signal show that the blend cannot.
Operate it: alert on the degraded share with a reason breakdown, and check the mix before accepting that a quality drop is a model regression.
Decide what standing budget of degraded traffic is acceptable and how a permanently elevated share gets reviewed rather than paged forever.
## The outage that produces no errors A fallback ladder converts failures into successes by design. When the personalised scorer stops answering, the request does not fail — it is served from the cached or frequency-list rung and returns a perfectly well-formed strip. Every monitor that watches for failure therefore sees a healthy system: request success is unchanged, latency may even improve because the cheap rung is faster, and the client reports nothing. The system is doing exactly what it was built to do, and that is why the quality loss has to be measured on purpose rather than discovered. The only quantity that moves is the **mix** — which rung produced each rendered strip. Treat it as a primary signal. ## How a blended metric absorbs the loss Suppose the top rung's acceptance rate — the share of rendered strips whose suggestion the user taps — is 22%, and the global frequency-list rung achieves 9%. In normal operation 1% of keystrokes are served below the top rung: - blended = 0.99 x 22 + 0.01 x 9 = **21.87%** Now a feature job fails and the fallback share rises to 6%: - blended = 0.94 x 22 + 0.06 x 9 = **21.22%** The headline number fell by 0.65 points, roughly 3% relative — well inside the day-to-day band on most products, and indistinguishable from a weekday effect. Meanwhile 6% of keystrokes lost more than half their acceptance rate. The blend is not lying; it is averaging two populations that should never have been averaged. ## Two signals, not one | signal | what it answers | how it is used | |---|---|---| | share of responses below the top rung | how much traffic is degraded right now | alert on it directly; it is the only thing an outage moves | | acceptance rate per rung | what being degraded costs | sets the floor each rung must hold and prices the outage | | reason tag on each degraded response | why it degraded — timeout, absent feature, shed for load | separates a capacity problem from a dependency problem | The first two are complementary: the mix without per-rung quality tells you something changed but not what it cost; per-rung quality without the mix tells you the price but not the volume. ## The regression that was really a fallback share The expensive version of this mistake is not the incident but the slow one. A feature writer breaks for one shard, 5% of traffic quietly falls to the frequency-list rung, and it stays there. Blended acceptance drifts down by half a point and never recovers. Read as a model metric, that looks exactly like a model gone stale, so the team retrains — and the retrain does nothing, because the model was never the problem. Worse, the new model is evaluated against a baseline already polluted by the degraded slice. The first question when a quality metric moves should therefore always be: **did the rung mix move at the same time?** If it did, the serving path changed, not the model. ## What to record on every rendered strip Four fields are enough: - the **rung** that produced the candidates; - the **reason** that rung was chosen; - which **feature groups** were present when the decision was made; - whether the user accepted a candidate, and which one. That record is what makes every statement above computable. Without it, an incident review can establish that the fallback fired but not how much of it there was or what it cost. ## The rows that flow downstream Accepted suggestions are also the raw material for later training data. An acceptance on a generic frequency-list suggestion is evidence about the frequency list, not about what the personalised scorer would have proposed. Left unmarked, those rows enter the next training set as if a personalised model had produced them. The rung tag is what lets the pipeline drop or down-weight them, and it costs nothing extra because the tag already exists for the metric. ## Alerting on the mix Because degraded responses are successes, the alert cannot be an error-rate threshold. It is a threshold on the degraded share over a short window, with the reason breakdown attached so the page says which rung and why. Pair it with a slow-moving check on the same share over a long window, which is what catches the 5%-forever case that no incident alert will ever fire on.
- Acceptance rate is down two points and a team starts a retrain. What should they check first?The rung mix over the same window. If the share of keystrokes served below the top rung rose at the same time, the model did not regress — a feature job or the scorer fleet did, and the blended metric is averaging two populations. Retraining on that signal answers a serving fault with a modelling fix and buries the real cause.
- Why should degraded responses be marked in the logs that later training sets are built from?Because acceptance of a generic frequency-list suggestion is evidence about that list, not about what the personalised scorer would have proposed. Left unmarked, those rows enter the next training set as if the scorer had produced them. The rung tag already exists for the metric, so the pipeline can drop or down-weight them at no extra cost.
- What alert fires for a degraded share that has been elevated for two months?None, if the only alert is a short-window threshold — a level that never changes never crosses it. Carry a second, slow check comparing the degraded share against a baseline over weeks, or a standing budget for degraded traffic that is reviewed rather than paged. Permanent degradation is a review item, not an incident.
saying these in an interview costs you the question
- Assumes an outage always shows up in the error rate
- Reads one blended acceptance rate across all rungs
- Treats a rising fallback share as a model regression
- Logs the suggestion without the rung that produced it
- Thinks a filled strip means the system is healthy