skip to content

Why can't a LangSmith online evaluator score live traffic against a reference answer?

level: seniorimportance: must knowfreq 48%

answer

  1. production has no second half
  2. the reference lives on the example
  3. judge only what is on the run
  4. grounding yes, exact match no
  5. truth arrives late, by review or outcome

basics

~20 s

Production runs carry an input and an output but no ground truth, so a metric that needs an expected answer has nothing to compare against. Online rules need reference-free signals; the reference only exists on dataset examples, which is where reference-based metrics belong.

solid answer

~60 s

An automation rule points an evaluator at real traffic. A real run has inputs, an output, whatever context was retrieved, and metadata — but nobody wrote down what the correct answer was, because if you knew it you would not have called the model. So an evaluator built around comparing the output to an expected output either errors or, worse, scores against nothing and produces a confident, meaningless number that lands on your chart. What works online is anything judged from material present on the run itself: is the answer grounded in the context that was retrieved, does it address the question asked, does it follow the required format or policy, did it refuse when it should have, were the tool calls well-formed. Ground truth arrives later and by other means — a human in an annotation queue writing a correction, or a delayed real-world outcome posted back with `create_feedback`. Once you have those labels, promote the run into a dataset and let reference-based metrics run there, offline, where the reference actually exists.

go deeper

for a junior

Understand the basic asymmetry: a dataset example has an expected answer attached, a live production run does not, so metrics needing one cannot run on live traffic.

for a middle

Be able to sort metrics into what is computable from the run alone — grounding in retrieved context, relevance, format compliance, tool-call validity — versus what needs a curated expected answer.

for a senior

Warn that a reference-based judge on production usually fails silently rather than erroring, and describe the loop that manufactures truth afterwards: review or delayed outcome, then promotion into a dataset for offline scoring.

for a principal

Own the instrumentation decision that makes delayed ground truth possible at all — which downstream outcomes get captured and written back — because that channel is worth more than any judge configuration and only exists if someone designed for it.

## The asymmetry between online and offline An offline experiment evaluates a dataset. A dataset example has two halves: the inputs, and the expected outputs someone curated. That second half is what makes exact match, semantic similarity to a gold answer, or reference-based recall computable at all. A production run has no second half. It has what the user sent, what the system did, and what came out. The reference is the one thing production never provides, and no configuration flag conjures it. This is why the same evaluator can be perfectly valid in an experiment and worthless attached to an automation rule — the metric did not change, the availability of its input did. ## The failure is quiet, which is the danger If a reference-requiring evaluator simply crashed on every production run, this would be a five-minute lesson. Often it does not. The judge prompt has a slot for the reference, the slot is empty or filled with something meaningless, and a language model asked to compare an answer against nothing will still return a number in the requested format. You get a full chart. It trends. People make decisions from it. The score is noise wearing the costume of a metric, and the only tell is that it does not move when quality obviously does. So the practical advice is not only "choose the right metric" but "before trusting any online series, look at a handful of the individual judge outputs and confirm the judge was actually given what it needed". ## What is legitimately computable online Everything whose evidence is on the run: - **Grounding against retrieved context.** The retrieved documents are on the trace, so a judge can ask whether each claim in the answer is supported by them. This needs no gold answer at all. - **Relevance to the question asked.** The input is on the run, so "does this response actually address it" is answerable. - **Format and schema compliance.** Deterministic, cheap, no model needed — did it return valid JSON, did it include the required disclaimer, is the citation list non-empty. - **Policy and safety checks.** Did it disclose what it must, did it refuse an out-of-scope request, did it leak an internal identifier. - **Tool-call validity.** Were the arguments well-formed, did the tool error, did the agent loop. - **Self-consistency and structural signals.** Length anomalies, repetition, an unusual number of retries. Notice that most of these are checks on *behaviour* rather than on *truth*. That is the honest limit of online evaluation: it catches the system misbehaving, not the system being confidently wrong in a way that looks perfectly well-formed. Accepting that limit is what separates a realistic online setup from a fantasy one. ## Manufacturing ground truth after the fact There are two ways truth arrives late. **A human writes it.** Route interesting runs into an annotation queue. The reviewer scores it and, crucially, writes a correction — the answer that should have been given. Now that run has a reference. Promote it into a dataset, where its inputs become the example's inputs and the correction becomes the expected output. Reference-based metrics now work on it, offline, forever. **The world reveals it.** Many applications produce a delayed outcome that is better than any label: the user accepted the drafted email unchanged, the support ticket did not reopen, the extracted invoice total matched what accounting later booked, the suggested code was merged. These are ground truth, they cost nothing to obtain, and they arrive minutes to days after the run. Because feedback is attached by run id with no time limit, a downstream job can write them back onto the original run under their own feedback key whenever they become known. Designing for that second channel is one of the highest-leverage things you can do in a production LLM system, and it is invisible to anyone who thinks of evaluation as something that happens at request time. ## The loop this implies Online reference-free evaluation detects that something is off and selects candidates. Human review or delayed outcomes supply labels for a subset. Those labelled runs become dataset examples. Offline experiments over that dataset — where references exist — are what you actually gate a change on. Then the change ships and online evaluation watches it in the wild. Each half does the job the other cannot: online has all the traffic and none of the truth; offline has all the truth and none of the traffic. ## What a weak answer sounds like "You just configure the evaluator with the expected output." Ask where that expected output comes from for a request that arrived four seconds ago, and the answer dissolves. The second weak answer is "use a stronger judge model" — a better judge with no reference is still a judge with no reference.

  • If not on the run, where does a reference actually live?
    On a dataset example. An example is inputs plus expected outputs, and that expected half is what reference-based metrics consume. It is written by a curator, by a human reviewer's correction promoted from a production run, or by a known-good historical outcome. Which is exactly why reference-based scoring is an offline experiment activity rather than an online rule activity.
  • Can you ever obtain ground truth for live traffic?
    Yes, late. Many systems produce a downstream outcome that reveals the truth: the user sent the drafted email unchanged, the ticket did not reopen, the extracted total matched the booked figure. Feedback is attached by run id with no deadline, so a downstream job can write that outcome back onto the original run under its own key hours later. Designing for that channel beats any judge.
  • What is the risk of pointing a reference-based judge at production anyway?
    It usually does not fail loudly. The judge is handed an empty reference, still returns a well-formed score, and you get a plausible chart made of noise. The tell is that the series does not move when quality visibly does. Before trusting any online metric, read a few individual judge outputs and confirm the judge received the evidence it needed.

saying these in an interview costs you the question

  • Assuming LangSmith can supply the expected answer for a live run
  • Believing a stronger judge model compensates for a missing reference
  • Trusting an online score series without inspecting individual judge outputs
  • Thinking online evaluation can replace an offline suite before a deploy
  • Overlooking delayed real-world outcomes as a ground-truth source

context