What signature and return value does a custom LangSmith evaluator function need?
answer
- a plain function, no base class
- parameters are injected by name
- dict with a name and a number
- booleans average into a pass rate
- many scores from one call
basics
~20 sA LangSmith evaluator is an ordinary Python function that declares any subset of the arguments inputs, outputs and reference_outputs, and returns a dict such as {"key": "exact_match", "score": 1}. You pass it in the evaluators list of evaluate().
solid answer
~50 sAn evaluator is just a callable. The SDK inspects your parameter **names** and injects only what you ask for: `inputs` (the example's input dict), `outputs` (whatever your target function returned for that example), and `reference_outputs` (the example's stored expected output). The legacy `run` and `example` objects are still accepted for evaluators written against the older interface. The simplest return value is a dict with a `key` (the feedback name shown in the experiment column) and a `score` — a number or a boolean. You can add `comment` for a free-text justification, or `value` for a categorical label. Returning a bare bool/int/float also works: the key is then inferred from the function name. To emit several scores from one call, return `{"results": [...]}` with one dict or `EvaluationResult` per score. Each evaluator runs once per dataset row and its score is written as feedback on that row's run.
code
python · 23 linesfrom langsmith import evaluate
def my_app(inputs: dict) -> dict:
return {"answer": inputs["question"].upper()}
def exact_match(outputs: dict, reference_outputs: dict) -> dict:
return {
"key": "exact_match",
"score": outputs["answer"].strip() == reference_outputs["answer"].strip(),
}
def under_200_chars(outputs: dict) -> bool:
return len(outputs["answer"]) < 200
results = evaluate(
my_app,
data="my-dataset",
evaluators=[exact_match, under_200_chars],
)go deeper
Be able to write one from memory: a function taking outputs and reference_outputs that returns {"key": ..., "score": ...}, passed in the evaluators list. Say plainly that the SDK matches your parameter names.
Explain the four accepted return shapes and when you would use each, especially returning several results from one judge call to avoid paying for the model twice. Mention that booleans aggregate into a pass rate.
Show how you keep evaluators pure and defensively parsed so a flaky judge produces a readable failure rather than an experiment full of errors, and how comments make low scores triageable.
Own the conventions: one metric per key, a shared score scale across the org, and a small library of deterministic assertions everyone reuses so experiment tables from different teams can actually be compared.
## Where evaluators sit A LangSmith experiment has two halves. The **target** is your application: it takes an example's inputs and produces outputs, and the SDK records one run per dataset row. The **evaluators** are scoring functions applied to those recorded runs. You supply them through the `evaluators` argument, and each one contributes one column of scores to the experiment table. That separation matters: an evaluator never calls your app. It only ever sees what the app produced, plus what the dataset says the answer should have been. ## The signature is name-matched, not positional The modern interface is deliberately undemanding — write a plain function and name its parameters after the things you want: - `inputs` — the example's `inputs` dict from the dataset. - `outputs` — the dict your target returned for that example. - `reference_outputs` — the example's stored `outputs`, i.e. the expected answer. Present only if the dataset row has one. Declare any subset, in any order. A length check needs only `outputs`; an exact-match check needs `outputs` and `reference_outputs`; a groundedness judge might want all three. Because injection is by name, a typo (`reference_output`, singular) silently fails rather than quietly working — that is the single most common beginner error. The older interface, still supported, takes `run` and `example` objects instead and reaches into `run.outputs` and `example.outputs` itself. Prefer the plain-dict form for new code; you only need the run object when you want run metadata such as latency, token usage or child runs. ## What you may return Four shapes are accepted: 1. **A dict** — `{"key": "correctness", "score": 1}`. `key` names the feedback (this is the column header and the field you later filter and chart on); `score` is numeric or boolean. Optional extras: `comment` (a string explaining the score, invaluable when the score is produced by a judge model) and `value` (a categorical label, for evaluators whose output is a class such as "refusal" rather than a number). 2. **A bare primitive** — `True`, `0.8`, `3`. The key is taken from the function's name, so name the function after the metric. 3. **An `EvaluationResult`** — the typed equivalent of the dict, importable from `langsmith.evaluation`. 4. **Several results at once** — `{"results": [{"key": "a", "score": 1}, {"key": "b", "score": 0}]}`, or an `EvaluationResults` object. Use this when one expensive judge call can produce multiple sub-scores, so you pay for the call once instead of once per metric. ## Boolean, numeric, categorical Booleans are stored as 0/1 and average into a pass rate, which is usually what you want for assertions ("did it cite a source?"). Continuous scores in 0–1 are the convention for judged quality. Mixing scales across evaluators is legal but makes the experiment table hard to read; pick a convention per project and stick to it. ## Cheap checks and judged checks are the same interface Nothing in the interface knows whether your function is a string comparison or a model call. A deterministic check — JSON parses, required field present, output under a length limit, regex matches — is a two-line function that costs nothing. A judge is the same function with a model call inside, returning a score parsed from the model's reply. This is why experiments usually carry a mix: several free assertions plus one or two judged metrics. ## Failure and partial results An evaluator that raises produces an error on that row rather than a score; the rest of the run continues, and the missing scores show up as gaps in the table. Because judge calls are the flakiest part, defensive parsing (and a clear `comment` on failure) keeps an experiment readable. ## Practical conventions Name the function after the metric so the inferred key is sensible. Keep one metric per key, so aggregates mean something. Put the reasoning in `comment`, not in the key. And keep evaluators pure — they should read the run and the example and nothing else, so re-scoring later gives the same answer.
- What happens if your evaluator declares a parameter the SDK does not recognise?Injection is by name, so only inputs, outputs, reference_outputs and the legacy run/example are filled in. An unrecognised name is not populated and the call fails when the function is invoked, which surfaces as an error on every row of the experiment rather than a score. Spelling reference_outputs correctly is the usual fix.
- How would you return both a score and the judge's reasoning?Include a comment field alongside key and score: {"key": "groundedness", "score": 0.4, "comment": "claim about pricing is not in the retrieved context"}. The comment is stored with the feedback and shown next to the score in the experiment view, which is what makes a judged failure triageable instead of just a low number.
- Do you have to write judge evaluators yourself?Not necessarily. LangChain publishes prebuilt judge prompts in the separate openevals package, which you install alongside langsmith and pass into the evaluators list like any other callable. They are a starting point rather than a finished metric — the prompt still has to be adapted to your task's definition of correct.
saying these in an interview costs you the question
- Thinks the evaluator calls the application itself
- Uses positional arguments and assumes order matters
- Spells the parameter reference_output and never gets the expected answer
- Returns a bare string and expects it to be scored
- Believes an evaluator must subclass something