skip to content

How do you score a finished LangSmith experiment with a new evaluator?

level: middleimportance: should knowfreq 42%

answer

  1. the generations are already stored
  2. score again, generate once
  3. the async twin exists too
  4. new columns, same experiment
  5. frozen outputs, so no prompt change

basics

~20 s

Call evaluate_existing with the finished experiment's name or ID and the new evaluators, or aevaluate_existing for async ones. It fetches the recorded runs, applies the evaluators, and writes the new feedback onto the same experiment without executing your application again.

solid answer

~40 s

`evaluate_existing("my-experiment-name", evaluators=[...])` takes an experiment that already ran — by name or ID — pulls its runs from LangSmith, and applies the evaluators you pass, exactly as `evaluate()` would have. `summary_evaluators`, `max_concurrency` and `metadata` work the same way; `aevaluate_existing` is the async form. Because the experiment is bound to its dataset, `reference_outputs` is still available to the evaluators — the examples are joined back automatically. The new scores appear as additional feedback columns on the same experiment, so the previous evaluators' results remain. The limitation is that the outputs are frozen. Re-scoring tells you what a different metric says about the same generations; it cannot tell you anything about a changed prompt, model or retriever. For that you need a new `evaluate()` run.

code

python · 13 lines
python
from langsmith.evaluation import evaluate_existing


def mentions_source(outputs: dict) -> dict:
    text = outputs.get("answer", "")
    return {"key": "cites_source_v2", "score": "[source:" in text}


evaluate_existing(
    "rag-prompt-v3-8f2c1d",
    evaluators=[mentions_source],
    max_concurrency=8,
)

go deeper

for a junior

Know that a finished LangSmith experiment can be scored again with new evaluators via evaluate_existing, without running the application a second time.

for a middle

Explain what the evaluators receive — recorded outputs, plus reference outputs joined back from the dataset — and that new feedback is added as extra columns on the same experiment.

for a senior

Own the discipline around it: keys are contracts, so a changed metric definition needs a new key; and the target should record its intermediate artefacts so future evaluators have material.

for a principal

Frame it as a workflow change — generate once, iterate the metric cheaply against real outputs — and set the retention and naming policy that keeps months-old experiments re-scorable when the quality bar moves.

## The problem it solves Evaluators are the part of an experiment you get wrong first. You run 2,000 examples, look at the results, and realise the metric you needed was not the one you wrote. Re-running `evaluate()` would execute the application over all 2,000 examples again — the same generations, at the same cost, for no new information about the application at all. `evaluate_existing` separates those two costs. The runs are already stored in LangSmith; scoring them again requires only the evaluator calls. ## The call Pass the experiment's name (the auto-generated one with the prefix and suffix, or the one your `experiment_prefix` produced) or its ID, plus the evaluators: - `evaluators` — the per-row scoring functions, identical in form to the ones `evaluate()` takes. - `summary_evaluators` — run-level metrics, also supported. - `max_concurrency` — throttle for judged evaluators, which are model calls like any other. - `metadata` — annotate the scoring pass. `aevaluate_existing` is the async counterpart, for async evaluators. ## What the evaluators see Exactly what they would have seen live. The `outputs` are the recorded target outputs; `inputs` come from the example; and because the experiment is tied to the dataset it ran against, `reference_outputs` is joined back from the examples. Evaluators written against `run` and `example` also work, and here the `Run` object carries the original timing and token usage — this is what makes it possible to add a latency or cost evaluator after the fact. ## Where the feedback lands Onto the same experiment, as additional feedback keys. Existing scores are not replaced, so the experiment accumulates columns. Two consequences: - **Keys collide by name.** Re-running an evaluator with the same key adds more feedback under that key rather than overwriting the old value. If you fixed a buggy evaluator, give the corrected version a distinct key, or the column mixes two definitions of the metric and the aggregate becomes meaningless. - **Metric definitions must be stable over time.** Because historical experiments are compared on keys, a key whose meaning changed silently corrupts every trend built on it. Treat an evaluator's key as a contract. ## What it cannot do The outputs are frozen. If you want to know whether a new prompt is better, `evaluate_existing` is useless — nothing about the application is re-executed. It answers only "what does this other metric say about these same generations?" It also cannot score anything the original run did not record. If the target returned only the final answer and your new evaluator needs the retrieved documents, they are not there; you have to change the target to record them and run again. This is a real argument for having the target return more than the bare answer in the first place — retrieved context, the model used, intermediate decisions — so later evaluators have material to work with. ## The working pattern It changes the rhythm of evaluation work. Generate once, score many times: 1. Run the application over the dataset with a cheap baseline evaluator, or none. 2. Look at the outputs, and work out what actually distinguishes a good one from a bad one. 3. Write the evaluator, apply it with `evaluate_existing`, and read the scores. 4. Disagree with the scores, refine the evaluator, apply it again under a new key. Steps 3 and 4 cost judge calls only, so iterating on the metric is cheap enough to do properly. That matters, because an evaluator you iterated on three times against real outputs is worth far more than one you wrote from imagination and never checked. ## Operational notes Judged evaluators applied to a large historical experiment are a burst of provider traffic; use `max_concurrency` to keep it inside your rate limits. And keep experiments you may want to re-score — a clear `experiment_prefix` makes them findable months later, when the question is "how would the metric we use today have rated the version we shipped in March?"

  • You re-run a fixed version of an evaluator under the same key. What goes wrong?
    The new feedback accumulates under that key alongside the old, so the column mixes scores produced by two different definitions of the metric and its aggregate means nothing. Give the corrected evaluator a distinct key, or score into a fresh experiment. Treat a key as a contract: its meaning must stay stable, because historical comparisons join on it.
  • Your new evaluator needs the retrieved documents, but the finished experiment only recorded the final answer. What now?
    Re-scoring cannot help — evaluators see only what was recorded. You have to change the target to return the retrieved context alongside the answer and run evaluate() again. It is a good argument for making the target return its intermediate artefacts from the start, so later metrics have something to work with.
  • When must you use evaluate() rather than evaluate_existing?
    Whenever the thing you changed is the application: a new prompt, model, temperature, retriever or tool. evaluate_existing never executes the target, so the outputs it scores are the old ones. It answers what a different metric says about the same generations, never what a different version would have generated.

saying these in an interview costs you the question

  • Thinks evaluate_existing re-runs the application with the new evaluator
  • Reuses a key after changing what the evaluator measures
  • Expects to score data the original run never recorded
  • Believes re-scoring can validate a changed prompt
  • Assumes reference outputs are unavailable after the run finished

context