skip to content

Observability & Tracing

Seeing inside a running LLM system: traces over prompt and response chains, token dashboards, latency percentiles, error rates, and drift in answer quality. Interviewers ask what you would open first when users report that the assistant 'got worse' overnight.

on this pageshow

questions

15

Why can't thumbs-up/down ratings alone measure quality of a production LLM assistant?

level: middleimportance: must knowfreq 62%

answer

  1. coverage versus honesty
  2. who bothers to click
  3. under 1% rated, 100% observed
  4. the angriest users left before rating
  5. behaviour beats declaration for trends

basics

~20 s

Explicit ratings cover a tiny, self-selected slice of traffic, often under 1% of sessions, and the people who bother to click skew angry or delighted. Implicit signals such as retries, abandonment and human edits are emitted by every session, so they carry the trend.

solid answer

~50 s

Explicit feedback has a coverage problem and a selection problem. On a telecom help-center assistant, expect well under 1% of sessions to leave a thumbs rating, and the raters are not a random sample: people click when they are annoyed enough to complain or delighted enough to praise. The rated population then drifts with UI placement, campaigns and incident news, independent of actual quality, which makes rating percentage a poor trend line and a terrible absolute number. Implicit signals — the user rephrasing the same question, abandoning mid-session, escalating to a human, or an agent heavily editing a drafted reply — are emitted by 100% of sessions and are behavioural rather than declarative. The working split is: implicit signals for coverage and trend, explicit ratings as sparse but high-value labels for triage and for calibrating implicit ones, and a sampled judged score as a continuous read.

go deeper

for a junior

Know the two families by name — explicit ratings versus implicit behavioural signals — and be able to give two examples of each. Say plainly that most users never rate, so ratings alone cannot tell you how the assistant is doing.

for a middle

Explain the mechanics: what fraction of sessions actually rate, why the raters are a biased sample, why survivorship pushes the score upward, and how implicit signals restore coverage at the cost of ambiguity. Be ready to sketch which signals you would instrument first.

for a senior

Show production judgement: how you would design the signal stack for a real assistant, which signal you alert on, how you avoid reading single sessions from noisy proxies, and how you check that an implicit proxy actually tracks quality rather than session length.

for a principal

Own the strategic call: quality instrumentation is a budget line decided before launch, not after the first bad week. Argue for what becomes the reported number, who owns it, and why a cheap, biased metric on an executive dashboard is worse than no metric at all.

## The two families of quality signal A production LLM system emits two very different kinds of quality evidence. **Explicit feedback** is anything the user deliberately submits: a thumbs up or down, a star rating, a "was this helpful?" choice, a free-text complaint. **Implicit feedback** is behaviour you observe without asking: the user retypes the same question, abandons the session, escalates to a human, copies the answer, or — in an assist flow where a human sends the final message — edits the draft before sending. The distinction matters because they fail in opposite directions. Explicit feedback is precise about intent but almost absent. Implicit feedback is abundant but ambiguous. ## Coverage: the 1% problem On a consumer-facing assistant, a realistic explicit-rating rate is a fraction of one percent of sessions. Take a telecom help-center assistant where 0.9% of sessions leave a thumbs rating while 100% of sessions emit timing, turn structure, escalation and abandonment events. At that ratio, a day with 20,000 sessions yields about 180 ratings. Slice those by intent, language and channel and each cell holds a handful of votes — far too few to detect the kind of regression that matters (a few percent drop on one intent) before it has run for days. Coverage alone would be survivable if the missing 99% were missing at random. They are not. ## Selection and survivorship bias Who clicks? People at the emotional extremes, people who found the control, people still present at the end of the session. Three biases follow: - **Selection bias.** The rated set over-represents strong opinions. The modal outcome — a mildly useful answer — is almost never rated, so the rating distribution is bimodal and unrepresentative of the median experience. - **Survivorship bias.** Users who gave up in turn two and closed the tab never reach the rating control. The worst experiences are systematically excluded, which biases the score *upward* exactly when things are worst. - **Instrumentation bias.** Moving the control, changing its wording, or prompting for feedback shifts who rates. A rating percentage that jumps after a UI change tells you about the UI, not the model. The practical consequence: an absent rating is not a satisfied user. Treating no-rating as implicit approval is the single most common misreading of this data. ## What implicit signals buy you Implicit signals invert the tradeoff. Every session produces them, so they support segment slicing and hourly trends. And because they are behavioural, they are harder to fake and less sensitive to who feels like clicking. Useful families: - **Repair behaviour** — the user rephrases the same question within a short window, or asks "no, I meant…". Rephrasing seconds after a reply is often a stronger negative than an unclicked thumbs-down, because it is evidence the answer failed *for a user who kept trying*. - **Abandonment** — session ends immediately after a response, no follow-up, no resolution event. - **Escalation / containment** — the user asks for a human, or the deflection rate drops. - **Human correction** — in assist flows, the distance between what the model drafted and what the human actually sent. - **Downstream outcome** — the ticket reopens, the recommended action is reversed, the order is refunded anyway. ## Ambiguity is the cost Every implicit signal has innocent explanations. A short session may mean the answer was perfect. A rephrase may mean the user changed their mind. Abandonment may mean the phone rang. That is why implicit signals are used as *trends and comparisons*, not as verdicts on single sessions, and why each one should be validated against labelled examples before it is trusted (does high edit distance actually co-occur with human-labelled bad answers?). ## How the signals compose A workable production stack layers them: 1. **Implicit signals on all traffic** — the high-coverage trend line, sliced by segment, and the thing you alert on for volume-sensitive regressions. 2. **Sampled online judging** — a small percentage of live sessions scored continuously to give a quality number that is not behaviour-mediated. 3. **Explicit ratings** — sparse, but each one is a labelled example. Use them to triage (a thumbs-down with free text is a bug report), and as ground truth to check whether the implicit proxies actually correlate with perceived quality. 4. **Human review of a flagged sample** — the small, expensive top layer, aimed at whatever the other three disagree about. Read that way, the thumbs button is not a metric; it is a cheap labelling channel and a complaint funnel. The metric lives in the signals that every session emits. ## Common mistakes Reporting "94% thumbs-up" as satisfaction; comparing rating percentage across periods where the rating rate itself moved; adding more feedback prompts and believing the bias went away; discarding implicit signals as "too noisy" without ever testing them against labels.

  • If a 0.9% rating rate is too sparse to trend, what is it good for at all?
    Two things. First, triage: a thumbs-down with free text is effectively a bug report, and routing those into a review queue finds real defects cheaply. Second, labels: those few hundred rated sessions are human judgements you can use to check whether your implicit proxies and your sampled judge scores actually move with perceived quality. Sparse data is weak as a trend and valuable as ground truth.
  • What breaks if you attach the thumbs control to every message instead of once per session?
    Rating volume rises, but the semantics blur. Users often rate the turn where frustration peaked rather than the turn that caused it, so attribution drifts to the wrong span. Per-message ratings also over-weight long conversations, since a ten-turn session can cast ten votes. If you do it, record the message and trace id with each rating and analyse per-session, not per-vote.
  • How would you tell whether a rise in thumbs-up percentage is real improvement or a shift in who rates?
    Check the rating rate and rater mix alongside the score. If the denominator moved — more or fewer sessions rating, different intents or channels represented — the composition changed and the percentage is not comparable. Confirm against a high-coverage signal that does not depend on volunteering, such as escalation rate or a sampled judged score, over the same window and the same segments.

saying these in an interview costs you the question

  • Assuming a session with no rating was a satisfied user
  • Reporting thumbs-up percentage as overall customer satisfaction
  • Comparing rating percentages across periods when rating rate changed
  • Dismissing implicit signals as too noisy without testing them against labels
  • Believing a more prominent feedback prompt removes selection bias

context

open as a page

Which token counters belong on an LLM cost dashboard beyond input and output?

level: middleimportance: must knowfreq 66%

basics

~20 s

Reasoning (thinking) tokens, which are billed at generation rates but never shown, and the cache-write versus cache-read split of input tokens, which carry very different unit prices. Without those three extra counters, a dashboard cannot explain the invoice.

open as a page

How would you model one turn of a tool-using LLM agent as a trace of spans?

level: middleimportance: must knowfreq 58%

basics

~20 s

Make the turn one trace: a root run span, agent spans for each reasoning cycle, and child spans for every model call, tool call and retrieval. Each span's parent is whatever caused it and outlives it.

open as a page

How would you design a drift alert on hourly sampled quality scores for a live assistant?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Score a fixed small sample of live sessions each hour, compare the rolling mean against a same-hour-of-week baseline rather than a flat threshold, and fire only when the deviation persists across several windows. Pin the scorer version, or scorer drift will masquerade as quality drift.

open as a page

A trip-planning agent recommended a fully booked hotel — how do you localize the fault in its trace?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Open the turn's trace and split the question in two: did the availability fact ever reach the model? The retrieval and tool spans answer that; the model-call span's recorded input answers whether the fact was present and ignored.

open as a page

Why meter LLM spend from the provider's reported token usage rather than a local tokenizer estimate?

level: juniorimportance: should knowfreq 42%

basics

~20 s

The response's usage record carries the counts you are actually billed for, including tokens you never see, such as reasoning tokens and cached-prefix tokens. A local tokenizer only guesses the prompt, misses server-side additions, and drifts as models change.

open as a page

How do you turn a production failure trace into a permanent regression test case?

level: middleimportance: should knowfreq 48%

basics

~20 s

Freeze the trace: capture the inputs and the tool responses it saw, redact personal data, then write an assertion about the property that actually failed rather than the exact wording. Confirm the case fails before the fix and passes after, and tag it by failure mode so the suite stays prunable.

open as a page

How do session and run ids relate to trace ids across a multi-turn LLM conversation?

level: middleimportance: should knowfreq 44%

basics

~20 s

Each turn gets its own trace and trace id. A session id, recorded on the spans, groups the turns of one conversation; a run id names one agent invocation, which can cover more than one trace when the invocation is retried.

open as a page

How would you validate edit distance as a quality proxy for AI-drafted agent replies?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Treat it as a hypothesis, not a metric. Normalise the distance for length, then check on a few hundred human-labelled drafts whether high distance actually co-occurs with bad answers, per segment. If the correlation is weak or driven by boilerplate, the proxy is measuring style, not quality.

open as a page

What makes a burn-rate LLM spend alert better than a monthly-budget threshold?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A cumulative threshold fires long after the damage starts and names no cause. Burn-rate alerting compares current spend per hour against a trailing baseline, sliced by tenant and feature, so an anomaly pages within an hour and arrives with the culprit attached.

open as a page

Token-cost metrics labelled by model, prompt_version and tenant exploded to 180k Prometheus series — how do you fix it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Split the pipeline. Keep time-series metrics to bounded dimensions you alert on — model, effort, cache-hit bucket, feature, status — and move unbounded ones like tenant id and prompt version into per-call usage events in a warehouse or log store where you group at query time.

open as a page

Why follow the OpenTelemetry GenAI semantic conventions when tracing LLM calls?

level: seniorimportance: should knowfreq 37%

basics

~20 s

They fix the attribute names on a model-call span — provider, requested model, token usage, finish reason — so any backend can render and query your LLM spans without per-vendor mapping. As of mid-2026 they are still pre-stable, so pin a version.

open as a page

How do you establish cost per tenant per invocation before pricing a new LLM feature?

level: principalimportance: should knowfreq 34%

basics

~20 s

Meter first, price second. Ship the feature behind a flag to real tenants with attribution on every call, roll costs up to the user-visible invocation, and price off the p95 tenant's distribution plus tail share — never the mean.

open as a page

How do you set a prompt and completion payload capture policy for LLM traces?

level: principalimportance: should knowfreq 29%

basics

~20 s

Capture full prompt and completion payloads on a small sampled share of traces, plus always on errors and flagged turns. Keep attribute-only spans everywhere else, cap payload size, redact on the way out, and retain payloads for less time.

open as a page

Thumbs ratings hold flat while sampled judge scores fall 8%. Which signal do you trust?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Neither, until each signal's validity chain is checked: rating rate and rater mix on one side, scorer version and traffic composition on the other. Then break the tie with a third independent signal and a human read of flagged sessions, and decide in advance which signal is the system of record.

open as a page