skip to content

Online Eval & Annotation

Scoring production traffic instead of a fixed dataset: rules that sample live runs into an automatic evaluator or a human review queue, and the feedback those produce. The tradeoff conversation is sampling rate versus judge cost.

on this pageshow

questions

6

How do you attach an end-user thumbs-down to a LangSmith run from your app?

level: juniorimportance: must knowfreq 52%

answer

  1. point at a run, name a metric
  2. the key is the chart series
  3. root run, not a nested call
  4. score, value, comment, correction
  5. presigned token keeps the key server-side

basics

~20 s

Capture the traced call's run id, then call Client.create_feedback with that run id, a feedback key such as "user_score", and a score. The key names the metric, the score is the number, and comment or correction can carry the user's words or the fixed answer.

solid answer

~60 s

Two steps: get the run id, then write the feedback. The run id comes from the tracing context of the call you want to rate — in Python, `get_current_run_tree()` inside a traced function gives you the run whose id you should return to your front end alongside the answer. Then `Client.create_feedback(run_id, key="user_score", score=0, comment=...)` attaches it. The `key` is the name of the metric and becomes its own series in the project's charts and its own filter in the runs table, so keep your keys stable and few. `score` is numeric, `value` holds a categorical label, `comment` holds free text, and `correction` holds what the answer should have been. Never do this straight from a browser, because it needs your LangSmith API key: mint a presigned feedback token server-side with `create_presigned_feedback_token` and let the client post its score to that URL instead. Feedback written this way is the same object an online judge or a human reviewer produces, so it shows up in the same charts and can drive the same rules.

code

python · 11 lines
python
from langsmith import Client

client = Client()

client.create_feedback(
    run_id="9f7c4a2e-1b3d-4c5e-8a90-1f2b3c4d5e6f",
    key="user_score",
    score=0,
    comment="Quoted the wrong refund policy",
    correction={"output": "Refunds are available within 30 days."},
)

go deeper

for a junior

Know the shape: get the run id from the traced call, then call create_feedback with a run id, a key, and a score. Be able to say what key, score, comment and correction each hold.

for a middle

Explain that the key defines the chart series and must stay stable, that feedback belongs on the root run, and that a presigned feedback token is how a browser posts a score without your API key.

for a senior

Argue for implicit and delayed outcome signals over sparse thumbs, handle idempotency with your own feedback id, and connect low-scoring feedback to rules that route those runs into review or a dataset.

for a principal

Own the feedback vocabulary across teams — a small agreed set of keys and their meanings — because inconsistent keys make cross-service quality reporting impossible and no chart will reveal the drift.

## What feedback is In LangSmith, feedback is a record attached to a run: a named metric with a value, plus optional prose. It is deliberately generic, because the same object is written by three different sources — your application (a user's rating, an implicit signal), an online evaluator (a judge score), and a human in an annotation queue (a review verdict). Everything downstream treats them alike: the monitor charts plot feedback keys over time, the runs table filters and sorts on them, and an automation rule can use existing feedback as its filter. ## Getting the run id Feedback needs to point at a specific run. Inside a traced function, `get_current_run_tree()` returns the run currently being traced, and its id is what you hand back to your client alongside the model's answer — typically in the API response body next to the text the user is about to read. When the user clicks thumbs-down, the client sends that id back and you write the feedback. The practical wrinkle is which run to point at. A single interaction produces a whole tree — retriever, LLM calls, tools. User feedback belongs on the *root* run, because that is the unit the user experienced. Putting it on a nested LLM call makes it awkward to chart and impossible to compare against an end-to-end judge score. ## Writing it `Client.create_feedback` takes the run id, a `key`, and then whichever payload fields apply: - `key` — the metric's name. This is the single most consequential argument, because it is the identity of the series. `user_score` written by your app and `correctness` written by a judge are separate series that never mix. Choose a small, stable vocabulary; a typo creates a new metric silently. - `score` — a number. Thumbs are conventionally 1 and 0, which makes the average of the series a satisfaction rate. - `value` — a categorical label when a number is the wrong shape (`"wrong_tone"`, `"hallucinated"`). - `comment` — free text, usually the user's own words from a feedback box. This is what you actually read when you go hunting for why the score dropped. - `correction` — the output the user or reviewer says it should have been. This is the field that makes a bad run useful later, because it is the expected output when the run is promoted into a dataset. You may also supply your own `feedback_id`, which matters for idempotency: a user who clicks twice, or a retried request, otherwise produces two rows and a skewed average. Deriving a stable id from your own message id makes a resubmission overwrite rather than accumulate. ## Doing it from a browser safely The obvious implementation — call `create_feedback` from front-end JavaScript — requires shipping your LangSmith API key to every user, which hands them full access to your traces. The supported alternative is a presigned feedback token: server-side, you call `create_presigned_feedback_token` for the run and feedback key, and give the client the resulting URL. The browser posts its score to that URL. The token is scoped to one run and one key and expires, so the worst case is a bogus rating on a single run rather than a leaked project. ## Implicit signals are often better than explicit ones Explicit thumbs are sparse — a small single-digit percentage of interactions, and biased toward people who are annoyed. The same API takes signals you already have: whether the user regenerated, whether they copied the answer, whether they edited a drafted reply before sending, whether the support ticket reopened a day later. Each becomes its own feedback key with its own score, written from wherever in your system that outcome becomes known — including hours after the run, which is fine, because feedback is attached by run id and has no deadline. Delayed outcome signals are the closest thing to ground truth that production traffic ever produces. ## How this connects to the rest of the loop Once a run carries a low `user_score`, it is a candidate for everything else: an automation rule can filter on that feedback and route those runs into an annotation queue for a human to look at, or into a dataset so the failure becomes a permanent test case. That is the practical reason to write feedback at all — not to admire a satisfaction chart, but to give your rules something to select on. ## Common mistakes Inventing a new key per deploy, so no series is longer than a week. Attaching feedback to whichever nested run was convenient. Writing scores from the browser with a production key. Not handling double submission. And treating the thumbs rate alone as your quality metric, when it measures who bothered to click as much as it measures the answers.

  • How do you record what the answer should have been, not just that it was wrong?
    Use the `correction` field on the feedback. It stays attached to the run alongside the score, and it is what you promote as the expected output when that run is later added to a dataset. A score tells you the rate went down; a correction is what turns one bad run into a permanent regression case.
  • A user clicks thumbs-down twice, or the request is retried. What happens?
    By default you get two feedback records and a skewed average on that key. Supply your own `feedback_id`, derived from something stable like your message id, so the second write replaces the first instead of appending. The same reasoning applies to a presigned token, which is scoped to one run and one key.
  • Explicit thumbs cover maybe two percent of interactions. What else can you write through the same API?
    Implicit and delayed signals, each under its own key: the user regenerated, they copied the answer, they edited a drafted reply before sending, the ticket reopened the next day. Feedback is attached by run id with no time limit, so an outcome that only becomes known hours later can still be written back to the run that caused it.

saying these in an interview costs you the question

  • Calling create_feedback from the browser with the project API key
  • Inventing a new feedback key per release so no series survives
  • Attaching user feedback to a nested run instead of the root run
  • Ignoring duplicate submissions and skewing the average
  • Treating the thumbs rate as the whole quality picture

context

open as a page

In LangSmith, what does an automation rule do to the runs in a tracing project?

level: middleimportance: must knowfreq 62%

basics

~20 s

A LangSmith automation rule watches a tracing project, keeps the runs matching its filter, samples a percentage of those, and sends each sampled run to an action: an online evaluator, an annotation queue, or a dataset.

open as a page

Why can't a LangSmith online evaluator score live traffic against a reference answer?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Production runs carry an input and an output but no ground truth, so a metric that needs an expected answer has nothing to compare against. Online rules need reference-free signals; the reference only exists on dataset examples, which is where reference-based metrics belong.

open as a page

In LangSmith, how do you pick an online evaluator's sampling rate at high volume?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Work backwards from the judge bill: sampled runs equal eligible runs times the rate, and each sampled run costs at least one judge model call. Narrow the rule's filter first, then set a rate that yields a few hundred to a few thousand scored runs a day.

open as a page

What is a LangSmith annotation queue, and how do runs end up in one?

level: middleimportance: should knowfreq 45%

basics

~20 s

An annotation queue is a review worklist of production runs. A human opens it, sees one run's input and output at a time, applies the queue's rubric, and their verdict is written back as human-sourced feedback on that run.

open as a page

In LangSmith, how do you split coverage between an online judge and human review?

level: principalimportance: should knowfreq 36%

basics

~20 s

Two different budgets: dollars for the judge, reviewer-hours for the queue. Send the judge wide and cheap across all traffic for a trend line, and send humans narrow and deep into the slices where being wrong is expensive and where their labels calibrate the judge.

open as a page