In LangSmith's evaluate(), what must the target function accept and return?
answer
- one callable, one dict in, one dict out
- called once per example
- produces an experiment, not a number
- name it with a prefix that says what changed
- key names are a three-way contract
basics
~20 sThe target takes one argument — a single example's inputs dict — and returns a dict of its outputs. LangSmith calls it once per example in the dataset and records each call as a run inside one named experiment.
solid answer
~60 s`evaluate(target, data=..., evaluators=[...])` is the whole loop: `data` names the dataset, and `target` is any callable that takes **one example's `inputs` dict** and returns a **dict** of results. LangSmith invokes it once per example, traces each invocation as a run, joins the returned dict with that example's reference `outputs`, and hands both to the evaluators. The output of the run is not a score — it is an **experiment**: a named, browsable set of runs over that dataset, with aggregate scores you can line up against other experiments in the comparison view. Two consequences matter in practice. First, the target is where you put the *thing under test* — a prompt, a chain, an HTTP call to your own service — and it should be the thinnest wrapper possible, so what you compare is the change and not your test harness. Second, key names are a contract: what the target reads out of `inputs` and what it puts in its returned dict have to match the dataset and the evaluators, or you get errored rows instead of low scores.
code
python · 23 linesfrom langsmith import evaluate
def my_app(question: str) -> str:
return "Use the Forgot password link on the sign-in page."
def target(inputs: dict) -> dict:
return {"answer": my_app(inputs["question"])}
def exact_match(outputs: dict, reference_outputs: dict) -> bool:
return outputs["answer"] == reference_outputs["answer"]
results = evaluate(
target,
data="support-qa",
evaluators=[exact_match],
experiment_prefix="prompt-v7",
metadata={"model": "gpt-4o-mini", "git_sha": "abc1234"},
max_concurrency=4,
)go deeper
Memorise the shape: a function that takes one inputs dict and returns a dict, passed to evaluate() with a dataset name. Say that LangSmith calls it once per example.
Explain what the run produces — a named experiment of individually traced runs — and the three-way key contract between dataset inputs, the target's returned dict and the reference outputs the evaluators read.
Show judgment about what goes inside the target: it defines the system under test, so retries and repair belong there only if production has them, and you check the errored-row count before trusting any aggregate.
Own the naming and cost discipline across a team — prefixes and metadata that keep months of experiments comparable, and the per-example call arithmetic that decides which suites can run on every commit.
## The one-line shape ``` def target(inputs: dict) -> dict: return {"answer": my_app(inputs["question"])} evaluate(target, data="support-qa", evaluators=[...], experiment_prefix="v3") ``` That is the entire integration surface. `evaluate` fetches the dataset named by `data`, calls `target` once per example with that example's `inputs` dict, and collects each returned dict as the run's outputs. ## What `data` can be `data` accepts a dataset name string, a dataset UUID, or an iterable of examples — which is how you narrow a run to part of a dataset, by passing the result of a filtered example listing instead of a bare name. Whatever you pass, the examples must all belong to the same dataset for the results to be comparable. ## What comes back `evaluate` returns a results object you can iterate over, but the durable artifact lives on the server: an **experiment**. An experiment is a named collection of runs, one per example (times any repetitions), each fully traced — you can open any row and see the exact prompt, the model call, the latency, the token counts and the scores attached to it. This is the reason people use the tool instead of a script that prints an average: when the number moves, you can click into the individual example that moved it. `experiment_prefix` controls the name. LangSmith appends a generated suffix so names stay unique across reruns, and the prefix is what you actually read in the experiment list, so it should say what changed: `"prompt-v7"`, `"gpt-4o-mini"`, `"with-reranker"`. A default, unset prefix produces an auto-generated name that tells a colleague nothing three weeks later. `metadata={...}` is the machine-readable version of the same idea — record the model, the temperature, the git sha — and it is what lets you filter experiments later. ## Sync and async `evaluate` is the synchronous entry point. If your application is `async def`, use `aevaluate` with an async target and await it inside an event loop; wrapping an async target in `evaluate` gives you a coroutine object as the run output, not a result, and every evaluator then scores gibberish. ## Keep the target thin The most common design mistake is to put orchestration inside the target: retries, fallback models, output post-processing that fixes what the prompt got wrong. Everything inside the target is part of the system under test. If the target silently retries three times, your experiment measures the retry loop, not the prompt, and a real regression in first-attempt quality is invisible. Put the smallest possible surface in the target and let the failures show. The reverse is also true: if production *does* retry and post-process, and your target does not, the experiment measures something no user ever experiences. The honest rule is that the target should be the same code path as production, imported rather than reimplemented — which usually means your application exposes a function you can call, and the target is one line that calls it. ## The three-way key contract There are three dictionaries in flight per example: the example's `inputs`, the target's returned dict, and the example's reference `outputs`. Evaluators read the target's dict as the prediction and the reference as ground truth. Nothing validates the key names for you, so a typo does not lower the score — it raises, and the row shows in the experiment as an error. Always eyeball the first experiment for an error count before you read the average; an average computed over the eight rows that happened not to raise is worse than no number. ## Costs are per example One experiment over N examples is at least N calls to whatever your target invokes, plus a call per LLM-based evaluator per example. A 500-example dataset with two judge metrics is 1,500 model calls per run. That is the arithmetic to do before you wire an experiment into a per-commit job, and it is why people keep a small fast subset for the frequent loop and reserve the full set for less frequent runs. ## Determinism Even at temperature 0, a target that calls a hosted model is not reproducible run to run — providers change model versions, and routing and system-level nondeterminism move outputs. So expect an experiment's average to wobble a little between identical runs. Whether a particular delta between two experiments is meaningful is a statistics question, not a tool question; what the tool gives you is the per-example detail to go and look.
- Your target is defined with async def. What breaks if you pass it to evaluate()?Calling an async function returns a coroutine rather than a result, so each run's output is a coroutine object and every evaluator scores nonsense — or errors. Use aevaluate with the async target and await it inside an event loop. aevaluate is the async twin of evaluate and takes the same data, evaluators, max_concurrency and experiment_prefix arguments.
- Why does a rushed experiment sometimes report a suspiciously high score?Because errored rows are not the same as failed rows. If the target raises on most examples — usually a key mismatch between the dataset inputs and what the target reads — only the surviving handful get scored, and the average is computed over those. Always check the error count on the experiment before reading the aggregate.
- Should retries and output repair live inside the target function?Only if they live in production too. Everything inside the target is part of the system under test, so a hidden retry loop makes first-attempt regressions invisible, while omitting a repair step that production really performs measures something no user experiences. The safest pattern is to import the same entry point your application serves and let the target be a one-line call.
- What belongs in experiment_prefix versus metadata?The prefix is the human-readable name you scan in the experiment list, so it should say what changed: prompt-v7, with-reranker, gpt-4o-mini. metadata is the machine-readable record — model name, temperature, git sha, dataset version — that you filter and group by later. Filling both takes seconds and is the difference between a comparable history and a pile of unnamed runs.
saying these in an interview costs you the question
- Thinking the target receives the whole dataset at once
- Expecting the target to return a score
- Passing an async target to evaluate instead of aevaluate
- Reading the average without checking errored rows
- Hiding retries and repair logic inside the target