In the langsmith SDK, when do you need RunEvaluator instead of a plain function?
answer
- the function form is sugar over the class
- state held across rows
- run metadata the payloads do not carry
- a small, memorable export list
- one widely-cited class is not in this SDK
basics
~20 sRarely. The SDK wraps any plain function into a RunEvaluator for you. Subclass RunEvaluator, or use the @run_evaluator decorator, when the evaluator needs constructor-held state such as a configured judge client, or must implement evaluate_run against the Run and Example objects directly.
solid answer
~40 s`langsmith.evaluation` exports `run_evaluator`, `RunEvaluator`, `StringEvaluator`, `EvaluationResult` and `EvaluationResults` alongside the `evaluate` family. Under the hood every evaluator you pass becomes a `RunEvaluator` with an `evaluate_run(run, example)` method; a plain function is simply adapted into one. So the class-based form buys you three things: **state** (a judge client, a compiled prompt, a threshold configured once and reused across rows, rather than rebuilt per call), **direct access** to the `Run` object — latency, token usage, child runs, errors — which the plain `inputs`/`outputs`/`reference_outputs` injection does not surface, and a natural place to return `EvaluationResults` with several scores. `StringEvaluator` is a ready-made `RunEvaluator` subclass for string-in/score-out grading. One trap: `LangChainStringEvaluator` is **not** importable from `langsmith.evaluation`. It belongs to the LangChain side of the ecosystem, not this SDK.
code
python · 16 linesfrom langsmith.evaluation import EvaluationResult, RunEvaluator
class LatencyBudget(RunEvaluator):
def __init__(self, budget_seconds: float) -> None:
self.budget_seconds = budget_seconds
def evaluate_run(self, run, example=None) -> EvaluationResult:
if run.end_time is None or run.start_time is None:
return EvaluationResult(key="within_latency_budget", score=0)
elapsed = (run.end_time - run.start_time).total_seconds()
return EvaluationResult(
key="within_latency_budget",
score=int(elapsed <= self.budget_seconds),
comment=f"{elapsed:.2f}s against a {self.budget_seconds}s budget",
)go deeper
Know that a plain function is the normal way to write a LangSmith evaluator and that no base class is required. Recognising the name RunEvaluator as the underlying interface is enough.
Explain that the SDK wraps functions into RunEvaluator objects, and give the concrete reasons to subclass: held state, access to the Run object, and returning several results. Name the real exports rather than guessing.
Show how you package shared evaluators for several teams — constructor-configured classes with explicit keys — and be precise that LangChainStringEvaluator is not part of the Python SDK.
Own the evaluator library as a dependency: versioning, who may change a metric definition, and the rule that a metric's key never changes meaning, since historical experiments are compared on those keys.
## Everything becomes a RunEvaluator When you pass a callable in the `evaluators` list, the SDK inspects its parameter names, wraps it, and ends up with an object exposing `evaluate_run(run, example)`. That is the real internal interface; the plain-function signature is sugar over it. Knowing this explains the whole design: nothing you can do with a class is impossible with a function, and nothing about a function is second-class. ## What `langsmith.evaluation` actually exports The public surface is small and worth memorising, because guessing here is how people invent imports that do not exist: `run_evaluator`, `EvaluationResult`, `EvaluationResults`, `RunEvaluator`, `StringEvaluator`, `evaluate`, `aevaluate`, `evaluate_existing`, `aevaluate_existing`, `evaluate_comparative`. That is the list. - **`RunEvaluator`** — the base class. Implement `evaluate_run(self, run, example)` and return an `EvaluationResult` or `EvaluationResults`. - **`run_evaluator`** — a decorator that turns a function taking `(run, example)` into a `RunEvaluator` without writing a class. - **`StringEvaluator`** — a concrete `RunEvaluator` subclass for the common case of pulling a string prediction (and optionally a string reference) out of the run and handing it to a grading function. - **`EvaluationResult` / `EvaluationResults`** — the typed return objects with `key`, `score`, `value`, `comment` and related fields; `EvaluationResults` carries a list of them. ## The three reasons to reach for the class **Held state.** A judged evaluator usually owns a model client, a prompt template, a threshold and maybe a parser. Constructing those inside a plain function means rebuilding them on every row. A class constructs once in `__init__` and reuses them, and lets you instantiate the same evaluator twice with different configuration — `Judge(model="small")` and `Judge(model="large")` — in the same experiment. **Run-level data.** The plain form injects payloads: inputs, outputs, references. It does not give you how long the run took, how many tokens it used, whether it errored, or what its child runs were. An evaluator that scores latency budget compliance, or that inspects the retrieval step nested inside the chain, needs the `Run` object — so either take a `run` parameter in the plain function or subclass. **Reusable library code.** When several projects share evaluators, a class with a clear constructor is a better distribution unit than a function that reads module-level globals. ## What you do not need it for Everything else. A length check, an exact match, a JSON-parses assertion, a regex, even a one-off judge call — all read better as plain functions, and the injected-by-name signature is the documented modern style. Reaching for a class first is over-engineering, and interviewers notice. ## The import trap `LangChainStringEvaluator` is a name that circulates widely in LangSmith material, and it is not in the Python `langsmith` package. Trying to import it from `langsmith.evaluation` fails. It belongs to the LangChain/JS side of the ecosystem; the Python class that does exist here is `StringEvaluator`. This is the kind of detail worth being precise about, because "I would use LangChainStringEvaluator" in an interview is a claim that does not survive a `python -c "from langsmith.evaluation import ..."`. For prebuilt LLM-as-judge evaluators in Python, LangChain publishes them in the separate `openevals` package, installed alongside `langsmith` — they are ordinary callables you drop into the `evaluators` list, not part of the `langsmith` SDK. ## Async `aevaluate` accepts async evaluators. A class-based evaluator that needs to await a model call can expose an async evaluation entry point; the plain async function is again the simpler path. Mixing sync and async evaluators in one run is possible but the sync ones block, so a judged evaluator in an async run should be async too.
- What can a RunEvaluator see that the plain inputs/outputs injection does not give you?The Run object itself: start and end time, latency, token usage, error state, tags and metadata, and the tree of child runs. That is what you need to score a latency budget, or to reach into the retrieval step nested inside a chain rather than only judging the final answer. A plain function can also take a run parameter to get the same access.
- Where do prebuilt LLM-as-judge evaluators come from for LangSmith in Python?From the separate openevals package that LangChain publishes, installed alongside langsmith. Its judges are ordinary callables you place in the evaluators list. They are not part of the langsmith SDK itself, and the class often named in older material, LangChainStringEvaluator, cannot be imported from langsmith.evaluation.
- Why would you instantiate the same evaluator class twice in one experiment?To compare configurations in a single run: a strict and a lenient threshold, or a small and a large judge model, each with its own key so both columns appear side by side. Constructor arguments make that natural, whereas the function form would need two near-duplicate definitions or a closure factory.
saying these in an interview costs you the question
- Claims LangChainStringEvaluator is importable from langsmith.evaluation
- Thinks a plain function is a lesser or unsupported form
- Rebuilds the judge client inside the function on every row
- Assumes the class form is required for LLM-as-judge evaluators
- Expects run latency or token counts in the injected outputs dict