skip to content

Evaluators

The scoring functions attached to an experiment: cheap deterministic checks, LLM-as-judge prompts, and summary evaluators that score the run as a whole. Also where pairwise comparison of two experiments fits.

on this pageshow

questions

6

What signature and return value does a custom LangSmith evaluator function need?

level: juniorimportance: must knowfreq 78%

answer

  1. a plain function, no base class
  2. parameters are injected by name
  3. dict with a name and a number
  4. booleans average into a pass rate
  5. many scores from one call

basics

~20 s

A LangSmith evaluator is an ordinary Python function that declares any subset of the arguments inputs, outputs and reference_outputs, and returns a dict such as {"key": "exact_match", "score": 1}. You pass it in the evaluators list of evaluate().

solid answer

~50 s

An evaluator is just a callable. The SDK inspects your parameter **names** and injects only what you ask for: `inputs` (the example's input dict), `outputs` (whatever your target function returned for that example), and `reference_outputs` (the example's stored expected output). The legacy `run` and `example` objects are still accepted for evaluators written against the older interface. The simplest return value is a dict with a `key` (the feedback name shown in the experiment column) and a `score` — a number or a boolean. You can add `comment` for a free-text justification, or `value` for a categorical label. Returning a bare bool/int/float also works: the key is then inferred from the function name. To emit several scores from one call, return `{"results": [...]}` with one dict or `EvaluationResult` per score. Each evaluator runs once per dataset row and its score is written as feedback on that row's run.

code

python · 23 lines
python
from langsmith import evaluate


def my_app(inputs: dict) -> dict:
    return {"answer": inputs["question"].upper()}


def exact_match(outputs: dict, reference_outputs: dict) -> dict:
    return {
        "key": "exact_match",
        "score": outputs["answer"].strip() == reference_outputs["answer"].strip(),
    }


def under_200_chars(outputs: dict) -> bool:
    return len(outputs["answer"]) < 200


results = evaluate(
    my_app,
    data="my-dataset",
    evaluators=[exact_match, under_200_chars],
)

go deeper

for a junior

Be able to write one from memory: a function taking outputs and reference_outputs that returns {"key": ..., "score": ...}, passed in the evaluators list. Say plainly that the SDK matches your parameter names.

for a middle

Explain the four accepted return shapes and when you would use each, especially returning several results from one judge call to avoid paying for the model twice. Mention that booleans aggregate into a pass rate.

for a senior

Show how you keep evaluators pure and defensively parsed so a flaky judge produces a readable failure rather than an experiment full of errors, and how comments make low scores triageable.

for a principal

Own the conventions: one metric per key, a shared score scale across the org, and a small library of deterministic assertions everyone reuses so experiment tables from different teams can actually be compared.

## Where evaluators sit A LangSmith experiment has two halves. The **target** is your application: it takes an example's inputs and produces outputs, and the SDK records one run per dataset row. The **evaluators** are scoring functions applied to those recorded runs. You supply them through the `evaluators` argument, and each one contributes one column of scores to the experiment table. That separation matters: an evaluator never calls your app. It only ever sees what the app produced, plus what the dataset says the answer should have been. ## The signature is name-matched, not positional The modern interface is deliberately undemanding — write a plain function and name its parameters after the things you want: - `inputs` — the example's `inputs` dict from the dataset. - `outputs` — the dict your target returned for that example. - `reference_outputs` — the example's stored `outputs`, i.e. the expected answer. Present only if the dataset row has one. Declare any subset, in any order. A length check needs only `outputs`; an exact-match check needs `outputs` and `reference_outputs`; a groundedness judge might want all three. Because injection is by name, a typo (`reference_output`, singular) silently fails rather than quietly working — that is the single most common beginner error. The older interface, still supported, takes `run` and `example` objects instead and reaches into `run.outputs` and `example.outputs` itself. Prefer the plain-dict form for new code; you only need the run object when you want run metadata such as latency, token usage or child runs. ## What you may return Four shapes are accepted: 1. **A dict** — `{"key": "correctness", "score": 1}`. `key` names the feedback (this is the column header and the field you later filter and chart on); `score` is numeric or boolean. Optional extras: `comment` (a string explaining the score, invaluable when the score is produced by a judge model) and `value` (a categorical label, for evaluators whose output is a class such as "refusal" rather than a number). 2. **A bare primitive** — `True`, `0.8`, `3`. The key is taken from the function's name, so name the function after the metric. 3. **An `EvaluationResult`** — the typed equivalent of the dict, importable from `langsmith.evaluation`. 4. **Several results at once** — `{"results": [{"key": "a", "score": 1}, {"key": "b", "score": 0}]}`, or an `EvaluationResults` object. Use this when one expensive judge call can produce multiple sub-scores, so you pay for the call once instead of once per metric. ## Boolean, numeric, categorical Booleans are stored as 0/1 and average into a pass rate, which is usually what you want for assertions ("did it cite a source?"). Continuous scores in 0–1 are the convention for judged quality. Mixing scales across evaluators is legal but makes the experiment table hard to read; pick a convention per project and stick to it. ## Cheap checks and judged checks are the same interface Nothing in the interface knows whether your function is a string comparison or a model call. A deterministic check — JSON parses, required field present, output under a length limit, regex matches — is a two-line function that costs nothing. A judge is the same function with a model call inside, returning a score parsed from the model's reply. This is why experiments usually carry a mix: several free assertions plus one or two judged metrics. ## Failure and partial results An evaluator that raises produces an error on that row rather than a score; the rest of the run continues, and the missing scores show up as gaps in the table. Because judge calls are the flakiest part, defensive parsing (and a clear `comment` on failure) keeps an experiment readable. ## Practical conventions Name the function after the metric so the inferred key is sensible. Keep one metric per key, so aggregates mean something. Put the reasoning in `comment`, not in the key. And keep evaluators pure — they should read the run and the example and nothing else, so re-scoring later gives the same answer.

  • What happens if your evaluator declares a parameter the SDK does not recognise?
    Injection is by name, so only inputs, outputs, reference_outputs and the legacy run/example are filled in. An unrecognised name is not populated and the call fails when the function is invoked, which surfaces as an error on every row of the experiment rather than a score. Spelling reference_outputs correctly is the usual fix.
  • How would you return both a score and the judge's reasoning?
    Include a comment field alongside key and score: {"key": "groundedness", "score": 0.4, "comment": "claim about pricing is not in the retrieved context"}. The comment is stored with the feedback and shown next to the score in the experiment view, which is what makes a judged failure triageable instead of just a low number.
  • Do you have to write judge evaluators yourself?
    Not necessarily. LangChain publishes prebuilt judge prompts in the separate openevals package, which you install alongside langsmith and pass into the evaluators list like any other callable. They are a starting point rather than a finished metric — the prompt still has to be adapted to your task's definition of correct.

saying these in an interview costs you the question

  • Thinks the evaluator calls the application itself
  • Uses positional arguments and assumes order matters
  • Spells the parameter reference_output and never gets the expected answer
  • Returns a bare string and expects it to be scored
  • Believes an evaluator must subclass something

context

open as a page

In LangSmith's evaluate_comparative, what does the evaluator receive and return?

level: middleimportance: should knowfreq 40%

basics

~20 s

evaluate_comparative takes two or more already-finished experiments over the same dataset and, for each example, hands the evaluator that example's inputs plus a list of outputs — one per experiment, in order. The evaluator returns a ranking, a score per experiment position.

open as a page

How do you score a finished LangSmith experiment with a new evaluator?

level: middleimportance: should knowfreq 42%

basics

~20 s

Call evaluate_existing with the finished experiment's name or ID and the new evaluators, or aevaluate_existing for async ones. It fetches the recorded runs, applies the evaluators, and writes the new feedback onto the same experiment without executing your application again.

open as a page

What does a LangSmith summary evaluator score, and what is it passed?

level: middleimportance: should knowfreq 55%

basics

~20 s

A summary evaluator scores an experiment as a whole rather than one row. Passed through the summary_evaluators argument, it runs once after all rows finish and receives the full lists of inputs, outputs and reference_outputs, returning a single key/score for the run.

open as a page

How do you keep judge cost sane when adding LLM evaluators to a LangSmith experiment?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Count first: judge calls equal rows times repetitions times judged evaluators. Then cut the multipliers — put free deterministic checks first, keep one judged metric rather than four, emit several sub-scores from a single judge call, run the expensive judge on a sample or on a nightly schedule.

open as a page

In the langsmith SDK, when do you need RunEvaluator instead of a plain function?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Rarely. The SDK wraps any plain function into a RunEvaluator for you. Subclass RunEvaluator, or use the @run_evaluator decorator, when the evaluator needs constructor-held state such as a configured judge client, or must implement evaluate_run against the Run and Example objects directly.

open as a page