What does a LangSmith summary evaluator score, and what is it passed?
answer
- one call, not one per row
- lists in, single score out
- attached to the run, not a row
- think precision, recall, F1
- accuracy does not need it
basics
~20 sA summary evaluator scores an experiment as a whole rather than one row. Passed through the summary_evaluators argument, it runs once after all rows finish and receives the full lists of inputs, outputs and reference_outputs, returning a single key/score for the run.
solid answer
~50 sOrdinary evaluators are per-row: each produces one score for one dataset example, and the experiment column you see is the average. A **summary evaluator** is called once, after every row has been executed and scored, and is given the whole run: parameter names `inputs`, `outputs` and `reference_outputs` receive **lists** aligned by example (the legacy form takes `runs` and `examples` lists instead). It returns the same shape as any evaluator — `{"key": ..., "score": ...}` — but the feedback is attached to the experiment, not to any individual run. You need one whenever the metric is not decomposable per row. Precision, recall and F1 over a classification dataset need the full confusion matrix. So do pass-rate-at-a-threshold, coverage of a label distribution, and any comparison of the output distribution against the reference distribution. Accuracy, by contrast, is per-row correctness averaged, so it needs no summary evaluator at all.
code
python · 28 linesfrom langsmith import evaluate
def my_app(inputs: dict) -> dict:
return {"label": "spam" if "free" in inputs["text"] else "ham"}
def f1(outputs: list, reference_outputs: list) -> dict:
tp = sum(
1
for o, r in zip(outputs, reference_outputs)
if o.get("label") == "spam" and r["label"] == "spam"
)
fp = sum(
1
for o, r in zip(outputs, reference_outputs)
if o.get("label") == "spam" and r["label"] != "spam"
)
fn = sum(
1
for o, r in zip(outputs, reference_outputs)
if o.get("label") != "spam" and r["label"] == "spam"
)
denom = 2 * tp + fp + fn
return {"key": "f1", "score": (2 * tp) / denom if denom else 0.0}
results = evaluate(my_app, data="spam-dataset", summary_evaluators=[f1])go deeper
Know that the summary_evaluators argument exists and that a function there is called once for the whole experiment rather than once per example. Naming F1 as the classic example is enough at this level.
Explain the injected lists, the aligned indices, and the single key/score return attached to the experiment. Be ready to say which metrics decompose per row and therefore should not be summary evaluators.
Show the failure handling: errored rows, missing reference outputs and empty lists all reach the summary evaluator last, after all the spend, so the arithmetic must be defensive and the treatment of errors stated explicitly.
Decide which summary metric is the org's headline number and how it is defined once, so experiments from different teams compare. Own the tradeoff that summary metrics are the verdict while per-row scores are the diagnosis.
## Two levels of scoring An experiment produces one run per dataset example. Evaluators in the `evaluators` list are invoked once per run and their scores are stored as feedback on that run; the number displayed at the top of the experiment column is simply the mean of those per-row scores. That covers most metrics, because most metrics really are averages of a per-example judgement. Some are not. Pass a `summary_evaluators` list to the same `evaluate()` call and each function there is invoked exactly once, after all rows have completed, with the entire experiment in hand. ## What a summary evaluator receives The same name-based injection applies, but every injected value is a list aligned by example: - `inputs` — list of every example's input dict. - `outputs` — list of every target output dict. - `reference_outputs` — list of every example's expected output. The older interface takes `runs` and `examples` — lists of the Run and Example objects — and is what you want when the metric depends on run-level metadata such as latency or token counts rather than on the payloads. Indices correspond: `outputs[i]` is the output produced for `inputs[i]`. Do not assume the dataset order is the source order; treat the lists as a set, and if you need to join on identity, use the run/example form and match on example IDs. ## What it returns and where the score lands The return shape is identical to a per-row evaluator: a dict with `key` and `score`, a bare number or boolean (key inferred from the function name), an `EvaluationResult`, or `{"results": [...]}` for several summary metrics from one pass. The difference is the attachment point — the feedback belongs to the experiment as a whole, so it shows as a single experiment-level number rather than a column of per-row values. That means you cannot click into a row to see why the summary score is what it is; if you need drill-down, emit a per-row evaluator alongside it. ## Metrics that genuinely need it - **Precision, recall, F1** — each requires counting true positives, false positives and false negatives across the whole set. No individual row knows whether it is a false positive relative to the corpus. - **Pass rate at a threshold** — "at least 95% of rows scored above 0.8" is a property of the distribution. - **Distribution comparisons** — does the model's label mix match the reference mix? Does the length distribution shift? - **Coverage** — did the run touch every category or intent in the dataset at least once? - **Costly aggregate judging** — one judge call shown a sample of the whole run, rather than one call per row. ## Metrics that do not Accuracy, mean latency, mean token count and pass/fail rate for a boolean assertion are all averages of per-row values. Writing them as summary evaluators is a mistake, not a shortcut: you lose the per-row column, so failures become invisible and you cannot filter to the rows that broke. The rule of thumb: if the metric can be computed for one row in isolation, make it a per-row evaluator and let the aggregate emerge. ## Cost and failure behaviour A summary evaluator runs once, so a judged summary metric is one model call per experiment rather than one per row — the cheapest place to put an expensive judge, at the price of no per-row explanation. It is also the last thing to run, which means an exception there wastes the whole experiment's target and per-row evaluator spend. Guard the arithmetic: empty lists, missing reference outputs on some rows, and rows whose target errored (and so have no usable output) all reach a summary evaluator and all break naive code. Filter for missing values explicitly rather than letting a division by zero take the run down at the end. ## Reading the results Because summary feedback is attached to the experiment, it is what you compare between experiments in the experiment list — two prompt versions side by side, each with its F1. Per-row scores are for diagnosis; summary scores are for the verdict.
- Why is accuracy a poor candidate for a summary evaluator?Accuracy decomposes: each row is either right or wrong, so a per-row boolean evaluator gives you the same top-line number plus a per-row column you can filter and drill into. Computing it as a summary metric throws away the diagnosis surface and gives nothing back. Reserve summary evaluators for metrics no single row can produce, such as F1.
- How do you handle rows whose target function errored when writing a summary evaluator?They still arrive in the lists, with missing or empty outputs, so the arithmetic has to account for them explicitly. Decide and document the policy: count errors as failures (usually right for a gate) or exclude them and report the excluded count. Silently skipping them inflates the metric, and dividing by an unfiltered length deflates it.
- Where would you put a judge that must see many outputs at once, such as one checking answer diversity?In a summary evaluator. It is the only place with the whole run in scope, and it costs one model call per experiment instead of one per row. Sample the outputs if the run is large enough to blow the context window, and emit the sample size in the comment so the number is interpretable later.
saying these in an interview costs you the question
- Thinks summary evaluators run once per dataset row
- Computes accuracy as a summary metric and loses per-row drill-down
- Assumes the injected lists are single dicts
- Ignores errored rows and reports an inflated aggregate
- Expects to click a summary score and see which row failed