Why can't thumbs-up/down ratings alone measure quality of a production LLM assistant?
answer
- coverage versus honesty
- who bothers to click
- under 1% rated, 100% observed
- the angriest users left before rating
- behaviour beats declaration for trends
basics
~20 sExplicit ratings cover a tiny, self-selected slice of traffic, often under 1% of sessions, and the people who bother to click skew angry or delighted. Implicit signals such as retries, abandonment and human edits are emitted by every session, so they carry the trend.
solid answer
~50 sExplicit feedback has a coverage problem and a selection problem. On a telecom help-center assistant, expect well under 1% of sessions to leave a thumbs rating, and the raters are not a random sample: people click when they are annoyed enough to complain or delighted enough to praise. The rated population then drifts with UI placement, campaigns and incident news, independent of actual quality, which makes rating percentage a poor trend line and a terrible absolute number. Implicit signals — the user rephrasing the same question, abandoning mid-session, escalating to a human, or an agent heavily editing a drafted reply — are emitted by 100% of sessions and are behavioural rather than declarative. The working split is: implicit signals for coverage and trend, explicit ratings as sparse but high-value labels for triage and for calibrating implicit ones, and a sampled judged score as a continuous read.
go deeper
Know the two families by name — explicit ratings versus implicit behavioural signals — and be able to give two examples of each. Say plainly that most users never rate, so ratings alone cannot tell you how the assistant is doing.
Explain the mechanics: what fraction of sessions actually rate, why the raters are a biased sample, why survivorship pushes the score upward, and how implicit signals restore coverage at the cost of ambiguity. Be ready to sketch which signals you would instrument first.
Show production judgement: how you would design the signal stack for a real assistant, which signal you alert on, how you avoid reading single sessions from noisy proxies, and how you check that an implicit proxy actually tracks quality rather than session length.
Own the strategic call: quality instrumentation is a budget line decided before launch, not after the first bad week. Argue for what becomes the reported number, who owns it, and why a cheap, biased metric on an executive dashboard is worse than no metric at all.
## The two families of quality signal A production LLM system emits two very different kinds of quality evidence. **Explicit feedback** is anything the user deliberately submits: a thumbs up or down, a star rating, a "was this helpful?" choice, a free-text complaint. **Implicit feedback** is behaviour you observe without asking: the user retypes the same question, abandons the session, escalates to a human, copies the answer, or — in an assist flow where a human sends the final message — edits the draft before sending. The distinction matters because they fail in opposite directions. Explicit feedback is precise about intent but almost absent. Implicit feedback is abundant but ambiguous. ## Coverage: the 1% problem On a consumer-facing assistant, a realistic explicit-rating rate is a fraction of one percent of sessions. Take a telecom help-center assistant where 0.9% of sessions leave a thumbs rating while 100% of sessions emit timing, turn structure, escalation and abandonment events. At that ratio, a day with 20,000 sessions yields about 180 ratings. Slice those by intent, language and channel and each cell holds a handful of votes — far too few to detect the kind of regression that matters (a few percent drop on one intent) before it has run for days. Coverage alone would be survivable if the missing 99% were missing at random. They are not. ## Selection and survivorship bias Who clicks? People at the emotional extremes, people who found the control, people still present at the end of the session. Three biases follow: - **Selection bias.** The rated set over-represents strong opinions. The modal outcome — a mildly useful answer — is almost never rated, so the rating distribution is bimodal and unrepresentative of the median experience. - **Survivorship bias.** Users who gave up in turn two and closed the tab never reach the rating control. The worst experiences are systematically excluded, which biases the score *upward* exactly when things are worst. - **Instrumentation bias.** Moving the control, changing its wording, or prompting for feedback shifts who rates. A rating percentage that jumps after a UI change tells you about the UI, not the model. The practical consequence: an absent rating is not a satisfied user. Treating no-rating as implicit approval is the single most common misreading of this data. ## What implicit signals buy you Implicit signals invert the tradeoff. Every session produces them, so they support segment slicing and hourly trends. And because they are behavioural, they are harder to fake and less sensitive to who feels like clicking. Useful families: - **Repair behaviour** — the user rephrases the same question within a short window, or asks "no, I meant…". Rephrasing seconds after a reply is often a stronger negative than an unclicked thumbs-down, because it is evidence the answer failed *for a user who kept trying*. - **Abandonment** — session ends immediately after a response, no follow-up, no resolution event. - **Escalation / containment** — the user asks for a human, or the deflection rate drops. - **Human correction** — in assist flows, the distance between what the model drafted and what the human actually sent. - **Downstream outcome** — the ticket reopens, the recommended action is reversed, the order is refunded anyway. ## Ambiguity is the cost Every implicit signal has innocent explanations. A short session may mean the answer was perfect. A rephrase may mean the user changed their mind. Abandonment may mean the phone rang. That is why implicit signals are used as *trends and comparisons*, not as verdicts on single sessions, and why each one should be validated against labelled examples before it is trusted (does high edit distance actually co-occur with human-labelled bad answers?). ## How the signals compose A workable production stack layers them: 1. **Implicit signals on all traffic** — the high-coverage trend line, sliced by segment, and the thing you alert on for volume-sensitive regressions. 2. **Sampled online judging** — a small percentage of live sessions scored continuously to give a quality number that is not behaviour-mediated. 3. **Explicit ratings** — sparse, but each one is a labelled example. Use them to triage (a thumbs-down with free text is a bug report), and as ground truth to check whether the implicit proxies actually correlate with perceived quality. 4. **Human review of a flagged sample** — the small, expensive top layer, aimed at whatever the other three disagree about. Read that way, the thumbs button is not a metric; it is a cheap labelling channel and a complaint funnel. The metric lives in the signals that every session emits. ## Common mistakes Reporting "94% thumbs-up" as satisfaction; comparing rating percentage across periods where the rating rate itself moved; adding more feedback prompts and believing the bias went away; discarding implicit signals as "too noisy" without ever testing them against labels.
- If a 0.9% rating rate is too sparse to trend, what is it good for at all?Two things. First, triage: a thumbs-down with free text is effectively a bug report, and routing those into a review queue finds real defects cheaply. Second, labels: those few hundred rated sessions are human judgements you can use to check whether your implicit proxies and your sampled judge scores actually move with perceived quality. Sparse data is weak as a trend and valuable as ground truth.
- What breaks if you attach the thumbs control to every message instead of once per session?Rating volume rises, but the semantics blur. Users often rate the turn where frustration peaked rather than the turn that caused it, so attribution drifts to the wrong span. Per-message ratings also over-weight long conversations, since a ten-turn session can cast ten votes. If you do it, record the message and trace id with each rating and analyse per-session, not per-vote.
- How would you tell whether a rise in thumbs-up percentage is real improvement or a shift in who rates?Check the rating rate and rater mix alongside the score. If the denominator moved — more or fewer sessions rating, different intents or channels represented — the composition changed and the percentage is not comparable. Confirm against a high-coverage signal that does not depend on volunteering, such as escalation rate or a sampled judged score, over the same window and the same segments.
saying these in an interview costs you the question
- Assuming a session with no rating was a satisfied user
- Reporting thumbs-up percentage as overall customer satisfaction
- Comparing rating percentages across periods when rating rate changed
- Dismissing implicit signals as too noisy without testing them against labels
- Believing a more prominent feedback prompt removes selection bias