In LangSmith's evaluate_comparative, what does the evaluator receive and return?
answer
- no target function is executed
- two finished experiments, one dataset
- outputs arrives as a list
- aligned by position, not by name
- the return is a ranking
basics
~20 sevaluate_comparative takes two or more already-finished experiments over the same dataset and, for each example, hands the evaluator that example's inputs plus a list of outputs — one per experiment, in order. The evaluator returns a ranking, a score per experiment position.
solid answer
~50 s`evaluate_comparative(experiments, evaluators=[...])` does not run your application. It joins **existing** experiments — identified by name or ID, and required to share a dataset — row by row on the dataset example, then calls each comparative evaluator once per example. The evaluator's injected `outputs` is a **list**, one entry per experiment in the order you passed them, alongside the shared `inputs` (and `reference_outputs` if the example has one). Instead of a single score it returns a ranking across those positions — typically a list of scores aligned to the experiment order, such as `[1, 0]` for "the first experiment won". Evaluators written against the older interface receive the paired `runs` and the `example` and return a mapping of run ID to score. Useful arguments: `randomize_order`, which shuffles the order outputs are presented to a judge; `num_repetitions`; `max_concurrency`; and `experiment_prefix` for naming the comparison.
code
python · 15 linesfrom langsmith import evaluate_comparative
def prefer_shorter(inputs: dict, outputs: list) -> list:
lengths = [len(o.get("answer", "")) for o in outputs]
best = min(range(len(lengths)), key=lambda i: lengths[i])
return [1 if i == best else 0 for i in range(len(outputs))]
evaluate_comparative(
("prompt-v1-abc123", "prompt-v2-def456"),
evaluators=[prefer_shorter],
randomize_order=True,
max_concurrency=4,
)go deeper
Know that evaluate_comparative compares two experiments that have already been run, over the same dataset, and that the evaluator sees both outputs for the same input at once.
Explain the injected list of outputs aligned to the experiment order, the ranking-shaped return value, and the fact that no target function is executed because both sides are already recorded.
Cover the operational details: randomize_order and num_repetitions when the evaluator is a model, max_concurrency as the throttle, and checking how many examples actually paired before quoting a win rate.
Own the reporting consequence — relative win rates decide a single A-vs-B call but cannot serve as a tracked quality metric, so the org needs a pointwise number alongside them and a retention policy for experiments worth comparing against.
## What the call actually does `evaluate_comparative` is the pairwise entry point in the LangSmith Python SDK. Unlike `evaluate()`, it takes no target function, because there is nothing to execute — both sides have already run. You give it a sequence of experiments (names or IDs, most often exactly two), and it produces a comparison whose scores are relative rather than absolute. The join is on the dataset example. Experiment A's run for example 42 is paired with experiment B's run for example 42, and the evaluator is invoked once for that pair. This is why the experiments must have been run over the same dataset: with no shared examples there is nothing to pair, and rows present in only one experiment cannot be compared. ## The evaluator contract Name-based injection works as it does elsewhere, but the shapes change: - `inputs` — the single shared example input dict; both sides saw the same input. - `outputs` — a **list** of output dicts, positionally aligned with the `experiments` argument. `outputs[0]` came from the first experiment. - `reference_outputs` — the example's expected output, if the dataset carries one. Pairwise comparison is most often used precisely because there is no usable reference, but nothing stops you using one. The return value expresses a ranking over those positions rather than a quality score: a list of numbers aligned to the experiment order, where a higher number means preferred — `[1, 0]` for a first-experiment win, `[0, 1]` for the second, and equal values for a tie. The older run-based interface receives `runs` and `example` and returns a key plus a mapping from run ID to score, which is the same information addressed by identity rather than position. The result is stored as comparative feedback, and the LangSmith comparison view renders it as a per-example win/loss along with the aggregate win rate. ## The arguments that matter - `experiments` — the sequence to compare. Two is the common case; more is allowed and turns the ranking into an ordering over several positions. - `randomize_order` — shuffles the order in which the paired outputs are handed to the evaluator. When the evaluator is a model, position in the prompt is not neutral, so this is the mechanical control the SDK gives you over that; the evaluator must therefore never assume position 0 is a particular experiment when reading the shuffled outputs, only that the returned scores are aligned to what it was given. - `num_repetitions` — repeats the comparison; with a stochastic judge, the aggregate over repetitions is more stable than a single pass. - `max_concurrency` — how many comparisons run in parallel, the throttle when the evaluator is a model call. - `experiment_prefix` — names the resulting comparison so it is findable later. ## Practical consequences **Both experiments must already exist.** The workflow is: run `evaluate()` for version A, run it again for version B, then compare. If you have not kept the earlier experiment, there is nothing to compare against, which is a real argument for naming and retaining experiments rather than treating them as throwaway. **One judge call per example, not two.** A pointwise judge scores each side separately — two calls per example across two experiments. A comparative judge sees both outputs in one prompt and answers once, which is often cheaper as well as easier for a model to answer than an absolute score. **Scores are relative and non-transportable.** A 62% win rate says B beat A on this dataset with this evaluator. It is not a quality level, it does not compose across experiment pairs, and it cannot be tracked as a trend line the way an absolute metric can. Teams that want a dashboard number need a pointwise metric too. **Ties need a policy.** Decide up front whether your evaluator may return equal scores and what a tie means for the decision; an evaluator forced to pick a winner on indistinguishable outputs produces noise that looks like signal. ## When it does not apply If the two experiments ran over different datasets, or one is a subset of the other, the comparison covers only the intersection and the aggregate is over fewer rows than you think. Check the compared count, not just the win rate.
- Why must the compared experiments share a dataset?The pairing is done per dataset example: run A for example 42 is matched with run B for example 42. Without shared examples there is no pair to hand the evaluator. If the datasets only partly overlap, the comparison silently covers the intersection, so check how many examples were actually compared before trusting the win rate.
- What does randomize_order change for the evaluator you write?It shuffles which experiment's output appears at which position before the evaluator sees them, so your code must not assume outputs[0] is a specific experiment — it may only return scores aligned to the list it was given. The SDK maps those back to the right runs. It exists because a model reading two candidates is sensitive to their order.
- You have a 58% win rate for B over A. What can you not conclude from it?That B is good. A win rate is purely relative: it says B was preferred more often on this dataset by this evaluator, with no absolute quality level attached, and it will not compose with a comparison against some third version. Keep a pointwise metric alongside it if you need a number to track over time.
saying these in an interview costs you the question
- Thinks evaluate_comparative runs the application for both versions
- Expects a single outputs dict instead of a list
- Compares experiments built from different datasets
- Reads a win rate as an absolute quality score
- Hardcodes outputs[0] as a particular experiment despite randomized order