skip to content

Thumbs ratings hold flat while sampled judge scores fall 8%. Which signal do you trust?

level: principalimportance: nice to knowfreq 30%

answer

  1. audit the instrument before the reading
  2. low coverage means low sensitivity
  3. find a signal sharing no failure mode
  4. thirty traces beat a week of dashboards
  5. name the system of record in advance

basics

~20 s

Neither, until each signal's validity chain is checked: rating rate and rater mix on one side, scorer version and traffic composition on the other. Then break the tie with a third independent signal and a human read of flagged sessions, and decide in advance which signal is the system of record.

solid answer

~50 s

Disagreement between quality signals is information, not an inconvenience. Start by asking what each signal could be doing other than measuring quality. Flat thumbs are weakly informative by construction: at a sub-1% rating rate, an 8% shift among a self-selected minority is easily invisible, and the rated population itself moves. A falling judged score can come from a scorer change, a traffic-mix shift toward harder requests, or a real regression. Check those mechanically. Then triangulate with a signal that shares no failure mode with either — escalation rate, session-repair rate, ticket reopen rate — and read a sample of the lowest-scoring sessions yourself, because thirty traces usually settle in an hour what a week of dashboard argument will not. The organisational half matters as much: agree beforehand which signal is the system of record and what evidence justifies a rollback, so this argument is not had during an incident.

go deeper

for a junior

Understand that a system has several quality signals and that they measure different things, so they can disagree. Say that you would check whether each signal moved for measurement reasons before concluding quality changed.

for a middle

Explain the concrete checks: rating coverage and rater mix on one side, scorer version and traffic composition on the other. Be ready to name a behavioural signal you would use to break the tie.

for a senior

Demonstrate incident judgement: audit the instruments, corroborate with an independent signal, read the lowest-scoring traces yourself, and let the stakes of the surface decide whether you roll back now or buy more evidence.

for a principal

Own the governance. Name which signal is the system of record and why, write the rollback criterion before the incident, assign ownership of each signal's validity, keep one signal outside the optimisation loop, and be candid that no consensus weighting formula exists.

## Disagreement is the normal case Any mature quality programme runs several signals with different coverage, different bias and different cost: explicit ratings, implicit behavioural signals, sampled automated scoring, and periodic human review. Because they measure different things through different mechanisms, they will disagree regularly. Treating disagreement as a data-quality bug leads teams to delete signals until only one remains — usually the most convenient one — and then to be blindsided when it turns out to have been the wrong one. The skill is arbitration: knowing what each signal can and cannot see, and having decided in advance how conflicts resolve. ## Step 1: audit each signal's validity chain Before comparing the numbers, ask what could move each one without quality changing. **Flat ratings.** At a fraction of a percent coverage, ratings have low statistical power: a genuine 8% quality drop can leave the rating percentage inside its normal noise band, especially if it is concentrated in a segment that rarely rates. Ratings are also composition-sensitive — if the drop drives some users to abandon before reaching the rating control, survivorship pushes the rated score *up* while experience worsens. Flat ratings are therefore weak evidence of "nothing changed"; they are mostly evidence that ratings are insensitive. **Falling judged score.** Three mundane explanations first. Did the scorer, its prompt, or its configuration change in the window? Did the traffic mix shift toward harder requests, so per-segment quality is unchanged and only the weighting moved? Did the sampling configuration change? Any of these produces a real-looking drop with nothing broken. ## Step 2: bring in an independent third signal The strongest tie-breaker is a signal whose failure modes overlap neither candidate. Behavioural signals qualify: human escalation rate, the rate at which users rephrase the same request, session abandonment, downstream ticket reopen rate, or task completion where the end state is verifiable. These are emitted by all traffic and do not depend on volunteering or on a scoring model. If escalation is up over the same window and in the same segments as the score drop, the judged signal is corroborated and the flat ratings are simply insensitive. If every behavioural signal is flat, suspicion shifts back onto the scorer or the traffic mix. ## Step 3: read the traces At some point stop comparing aggregates. Pull the lowest-scoring judged sessions from the window and read thirty of them. Either they show a recognisable, real failure mode — in which case the regression is real and you have your root cause and your regression-case candidates in the same hour — or they look fine and the scorer is drifting. This is cheap, fast, and settles disputes that dashboards cannot. It is also why the flagged-sample review layer exists in the first place. ## Step 4: decide with asymmetric costs in mind The decision is rarely "which number is true" and usually "what do we do now, given uncertainty". The asymmetry decides it. If the suspected regression touches a high-stakes surface — money, health, safety, legal commitments — the cost of a false alarm is a rollback and a wasted afternoon, while the cost of a missed regression is customer harm you will read about later. Act on the pessimistic signal. If the surface is low-stakes and rollback is disruptive, buy more evidence: raise the sample rate, add a human review batch, wait for a longer window. ## Step 5: settle the governance before the incident The part that separates a principal answer from a senior one is that most of this should have been decided in advance: - **Name the system of record.** One signal is the reported quality number, with its known biases documented. Others are diagnostics. Without this, every incident restarts the argument about which dashboard counts. - **Write the rollback criterion.** "Rolling judged score 1.5 sigma below baseline for six hours, corroborated by escalation rate, on a segment above X% of traffic" is a decision rule. "Quality seems worse" is not. - **Assign ownership.** Someone owns each signal's validity — its version pinning, its baseline, its documented biases — the same way a service has an owner. - **Expect Goodhart.** Whatever becomes the reported number will be optimised. Keep at least one signal off the scorecard and outside the optimisation loop so you retain an uncorrupted read. - **Re-audit periodically.** Signals decay: rating UIs move, scorers get upgraded, traffic changes. A quality signal that has not been validated against human labels in a year is a number, not a measurement. ## The honest position There is no consensus formula for weighting these signals, and anyone who offers one is overselling. What a strong answer conveys is the shape: signals are instruments with known distortions, arbitration means auditing the instruments before the readings, human reading of traces is the ground truth of last resort, and the decision rule should be written down while nobody is panicking.

  • Why is flat explicit feedback weak evidence that nothing regressed?
    Because its statistical power is tiny and its bias runs the wrong way. At sub-1% coverage, a real drop concentrated in one segment stays inside normal noise. Worse, a regression that makes users give up early removes them before they reach the rating control, so survivorship can push the rated score up while the experience gets worse. Flat ratings mostly tell you ratings are insensitive.
  • What makes a good tie-breaking third signal?
    One that shares no failure mode with either candidate. Escalation rate, repair behaviour and downstream reopen rate are emitted by all traffic, depend on no volunteering and on no scoring model, so they cannot be moved by rater self-selection or scorer drift. Verified task completion is even better where the end state is checkable, since it is an outcome rather than a proxy.
  • How do the stakes of the surface change what you do under this uncertainty?
    They set which error you would rather make. On money, health, safety or binding commitments, a false alarm costs a rollback while a missed regression costs real harm, so you act on the pessimistic signal and investigate afterwards. On a low-stakes surface with disruptive rollback, you buy evidence instead: raise the sample rate, commission a human review batch, and wait for a longer window before acting.

saying these in an interview costs you the question

  • Declaring the signal that agrees with your expectations the correct one
  • Deleting a disagreeing signal instead of auditing why it disagrees
  • Reading flat explicit ratings as proof that nothing regressed
  • Arguing about which dashboard counts during the incident itself
  • Putting every quality signal on the same scorecard the team optimises

context