skip to content

How would you design a drift alert on hourly sampled quality scores for a live assistant?

level: seniorimportance: must knowfreq 52%

answer

  1. small sample, hourly cadence
  2. compare against same hour last week
  3. rolling window beats single bucket
  4. pin the scorer's version
  5. mix shift is not a regression

basics

~20 s

Score a fixed small sample of live sessions each hour, compare the rolling mean against a same-hour-of-week baseline rather than a flat threshold, and fire only when the deviation persists across several windows. Pin the scorer version, or scorer drift will masquerade as quality drift.

solid answer

~50 s

Sample a fixed fraction of live conversations — around 2% is a common starting point — score them hourly, and store each score with the segment dimensions you will need later. Alert on a rolling mean over several hours rather than a single bucket, because hourly samples are small and noisy: a rule like "the six-hour rolling mean sits 1.5 standard deviations below the same-hour-of-week baseline" survives normal variance and seasonality far better than a fixed threshold. Three design details decide whether the alert is trustworthy. Pin the scoring model, its prompt and its version, so a provider-side change cannot look like a product regression. Track the sample size per bucket and suppress alerts when it is too small. And record traffic composition alongside the score, because a shift in intent or channel mix moves aggregate quality without anything getting worse. This reflects mid-2026 practice.

code

python · 13 lines
python
import statistics

# mean judged score per hour bucket, keyed by (weekday, hour)
baseline = {(1, 9): [0.82, 0.80, 0.84, 0.81, 0.83]}
rolling_now = 0.71
samples_in_window = 240

hist = baseline[(1, 9)]
mu = statistics.fmean(hist)
sigma = statistics.stdev(hist)
z = (rolling_now - mu) / sigma
alert = z <= -1.5 and samples_in_window >= 150
print(f"mu={mu:.3f} sigma={sigma:.3f} z={z:.2f} alert={alert}")

go deeper

for a junior

Know the idea: score a small random sample of real conversations continuously so you have a quality number that does not depend on users volunteering feedback. Recall that the sample must be random, not the sessions that already look bad.

for a middle

Explain why a fixed threshold fails and a seasonal baseline works, why hourly buckets are noisy at small sample rates, and what a rolling mean plus a sigma threshold buys you. Be ready to name scorer-version pinning as a requirement.

for a senior

Show operational judgement: derive the sample rate from the regression size you need to catch, design per-segment alerts, add minimum-sample guards, and describe the triage runbook that separates a mix shift from a real regression before anyone rolls back.

for a principal

Own the economics and governance: what continuous judging costs relative to serving, when the sample rate should flex, who is paged and with what error budget, and how scorer migrations are handled so a year of quality history stays comparable.

## What sampled online judging is for Offline evaluation tells you how a change performs on cases you already know about. Sampled online judging tells you how the system is doing right now, on real traffic, including the cases nobody anticipated. The mechanism is simple: take a small random slice of live sessions, score each one automatically, and treat the resulting time series as a continuous quality read. It fills the gap between explicit ratings (too sparse and too biased to trend) and human review (too slow and too expensive to run continuously). How the scorer itself is designed and validated — rubric shape, bias controls, agreement with human labels — is its own discipline. What matters here is treating its output as a monitored signal like any other. ## Choosing the sample Three parameters: rate, unit, and stratification. - **Rate.** Around 2% of conversations is a workable default: enough volume for hourly buckets on a high-traffic assistant, cheap enough to run indefinitely. Derive it rather than guess it — decide the smallest regression you want to catch and how fast, and check that the resulting bucket size gives that detection power. Low-traffic products often cannot support hourly buckets at all and should judge daily. - **Unit.** Sample whole sessions, not individual turns, if the failure you fear is conversational (looping, forgetting, failing to resolve). Sample turns if you care about per-response correctness. Mixing them makes the series uninterpretable. - **Stratification.** Pure random sampling under-covers small but important segments. Stratifying by intent or tier, then weighting back, keeps the headline number honest while giving each segment enough volume to alert on independently. Always sample *before* you know the outcome. Judging only sessions that already look suspicious produces a series that measures your filter, not your system. ## Building a baseline that survives seasonality A flat threshold ("alert below 0.80") fails immediately, because quality has structure. Overnight traffic differs from business hours. Monday differs from Saturday. A marketing push changes the intent mix for a day. A flat line either fires every night or is set so low it never fires. The usual fix is a **same-hour-of-week baseline**: compare this Tuesday 09:00 bucket against the distribution of previous Tuesday 09:00 buckets, and express the deviation in standard deviations of that history. That absorbs daily and weekly seasonality automatically. Keep several weeks of history, and exclude known-bad periods from the baseline or the incident you just had will normalise the regression it caused. ## Alert rules that do not cry wolf Hourly samples are small, so single-bucket alerts are dominated by noise. Practical shape: - Compute a **rolling mean over several hours** (six is a reasonable default) and compare that, not the raw bucket. - Fire when the rolling deviation crosses a threshold such as **1.5 sigma below baseline** and stays there. Requiring persistence across consecutive windows trades a little detection latency for a large reduction in false pages. - Add a **minimum-sample guard**: suppress the alert when the window's judged count falls below what the threshold assumes, or a quiet night will page you. - Alert **per segment as well as in aggregate**. A 20% regression confined to one intent that is 5% of traffic barely moves the headline number, and that is exactly the regression you want to catch early. ## The confounds that will bite you Three things move this series without quality changing: 1. **Scorer drift.** If the scoring model, its prompt, or its decoding settings change, the scale moves. Pin the model version and prompt hash, record both with every score, and re-baseline deliberately whenever you change them — never silently. Run the old and new scorer in parallel over the same window to translate between scales. 2. **Traffic-mix shift.** A campaign, an outage elsewhere, or a new channel changes which questions arrive. Aggregate quality drops because hard questions grew as a share, not because answers got worse. Record composition alongside the score and check it first when an alert fires; segment-level series usually settle the question immediately. 3. **Sampling changes.** Someone edits the sample rate or the stratification and the series steps. Version the sampling config the same way you version the scorer. ## What happens after the page An alert is only useful if it routes to a diagnosis. The runbook that pays for itself: check sample size and composition first; then compare segment series to localise; then pull the lowest-scoring judged traces from the window and read them. If they show a common failure mode, that is both your root cause and your regression-case candidate, which is how a drift alert feeds the eval loop rather than just interrupting someone's evening. ## Cost Judging is inference, so it has a bill. At a couple of percent of traffic it is typically a small fraction of serving cost, but the number should be explicit and monitored, and the sample rate should be a tunable knob — raised temporarily during a risky rollout, lowered when the series is quiet.

  • The alert fired. What do you check before declaring a quality regression?
    Sample size and traffic composition first. If the window is thin, or the intent, channel or language mix shifted, the aggregate can drop with nothing broken. Then compare per-segment series to localise it, and confirm the scorer version and prompt did not change during the window. Only after those come back clean do you read the lowest-scoring traces and start treating it as a real regression.
  • Why not alert on the raw hourly bucket instead of a multi-hour rolling mean?
    Because an hourly bucket at a 2% sample rate holds few enough sessions that ordinary variance regularly crosses any useful threshold. You would page constantly and the team would mute the alert, which is worse than having none. A rolling mean over several hours plus a persistence requirement costs a few hours of detection latency and buys an alert people still trust after a month.
  • How do you change the scoring model without corrupting the historical series?
    Treat it as a metric migration. Run old and new scorers in parallel over the same sampled traffic for a period, measure the offset and spread, then either rescale history or start a new series with a clearly marked boundary. Re-derive the baseline on the new scorer before enabling alerts on it. Never swap in place and keep alerting against the old baseline.

saying these in an interview costs you the question

  • Alerting on a fixed absolute score threshold with no seasonal baseline
  • Paging on a single hourly bucket at a small sample rate
  • Letting the scoring model or prompt change without re-baselining
  • Judging only sessions that already look suspicious
  • Reading an aggregate drop as a regression before checking traffic mix

context