In LangSmith, what does an automation rule do to the runs in a tracing project?
answer
- continuous, not a fixed dataset
- three parts, one destination
- who is eligible, how many, what happens
- filter then sampling rate then action
- evaluator, annotation queue, or dataset
basics
~20 sA LangSmith automation rule watches a tracing project, keeps the runs matching its filter, samples a percentage of those, and sends each sampled run to an action: an online evaluator, an annotation queue, or a dataset.
solid answer
~60 sAn automation rule is LangSmith's way of doing something to production traffic continuously, instead of running an evaluation over a fixed dataset. It has three parts. **A filter** decides which runs are eligible — root runs only, a particular run name or type, tags or metadata, error status, or existing feedback. **A sampling rate** between 0 and 1 decides what fraction of the eligible runs are actually taken. **An action** decides what happens to each taken run: score it with an online evaluator, drop it into an annotation queue for a human, or add it to a dataset. The work happens on LangSmith's side, asynchronously after the run is traced, so it never sits in your request path and never adds latency for the user. Whatever the action produces — a judge score, a human verdict — is written back as feedback on that run, so it shows up in the project's monitor charts and can be filtered on in the runs table. Rules act on runs as they arrive; they are not a way to re-score last month.
go deeper
Know that LangSmith can score or queue live production runs automatically, and that this is different from running an evaluation over a dataset you prepared.
Be ready to name the three parts — filter, sampling rate, action — say which three destinations an action can be, and explain that the work happens on LangSmith asynchronously rather than in your request path.
Show that you treat the filter as the primary cost control and that you know rules act going forward only. Talk about running several rules with different filters and rates on the same project rather than one global rule.
Own the policy: which projects get rules at all, what sampling rates are standard, who is allowed to create a rule with an evaluator action, and how the resulting spend is reviewed. An unowned rule is a recurring, invisible line item.
## The problem a rule solves Offline evaluation answers "is the new prompt better on the fifty examples I curated?". It cannot answer "is the thing users are actually sending us today being handled well?", because production traffic is not a dataset — it arrives continuously, it has no expected answers, and its shape drifts. An automation rule is the mechanism that lets you attach an evaluation or a review step to that live stream. ## The three parts of a rule **Filter.** A rule is scoped to one tracing project and then narrowed by a filter over runs. Typical narrowing: only root runs (so you score the whole chain once rather than every nested LLM call inside it), a specific run name, a run type, presence of a tag or a metadata key such as a prompt version, error status, or the presence or value of existing feedback. The filter is the single most important cost lever, because everything downstream is a percentage of whatever the filter lets through. **Sampling rate.** A fraction between 0 and 1 applied to the runs the filter admitted. At 0.05 on a project taking 100,000 runs a day, 5,000 runs a day reach the action. Sampling is random rather than "first N", which matters: it means the sample stays representative as traffic composition shifts over the day, and it means two rules on the same project see different runs. **Action.** Three destinations. An *online evaluator* runs a judge over the run's inputs and outputs and writes a score back. An *annotation queue* puts the run in a worklist for a human reviewer. A *dataset* adds the run as an example, which is how interesting production traffic becomes a regression suite. Nothing stops you from having several rules on one project — that is the normal setup. A cheap judge at a low rate over everything for a trend line, plus a high-rate rule on a narrow filter (a new prompt version, the flow that handles refunds, runs that already carry a thumbs-down) into a queue or a stricter evaluator. ## Where the work happens A rule executes on the LangSmith side, after the run has been ingested. It is not code running in your process and it is not in the user's critical path, so an expensive judge model costs you money and evaluation lag, not user-facing latency. It also means that if tracing is off or a run never reaches LangSmith, no rule fires on it — the rule can only see what was traced. ## Where the output lands Every action ultimately produces feedback attached to the run. A judge produces a score under a feedback key; a human reviewer in a queue produces feedback under the keys configured for that queue, recorded as human-sourced. Because it is ordinary feedback, it behaves like the feedback your application writes through the SDK: it appears as a series on the project's monitor charts over time, it can be filtered and sorted on in the runs table, and it can itself become the filter for a second rule — for instance, "any run the judge scored below 0.5 goes to the annotation queue". ## Operational notes worth saying out loud A rule applies going forward. Creating one today does not retroactively score yesterday's traffic; if you need history covered, select those runs in the runs table and add them to a queue or dataset by hand, or run an offline experiment over a dataset built from them. A rule with an evaluator action spends money on every sampled run, silently and continuously. That is the failure mode people actually hit: a rule created during a debugging session at a 1.0 sampling rate on an unfiltered project, forgotten, and discovered on the invoice. Rules are enumerable and deletable through the SDK client as well as the UI, and reviewing which rules exist on a project is a reasonable thing to do periodically. Finally, the feedback a rule wrote does not disappear when you delete the rule. If a mis-specified judge scored two days of traffic wrongly, those scores stay on the runs and stay in the charts, so when you read a chart you should know which feedback key and which source produced the series you are looking at. ## The interview framing The expected answer is the shape — filter, sampling rate, action — plus the observation that this is asynchronous and off the request path, plus the fact that the output is just feedback and therefore feeds the same charts and filters as user feedback. Candidates who describe a rule as "the thing that runs my evals" without mentioning sampling or the filter usually have not run one against real volume.
- Does creating a rule score the runs that were already in the project?No. A rule acts on runs as they arrive after it exists, so it is not a way to re-score history. If you need past traffic covered, select those runs in the runs table and add them to an annotation queue or a dataset by hand, then evaluate that dataset offline. Expect a gap between creating a rule and having enough sampled runs to read a trend.
- You notice a misconfigured evaluator rule has been spending money for two days. What do you do?Delete or disable the rule first — it keeps firing on every new run until you do. Then deal with the residue: the feedback it already wrote stays attached to those runs and stays in the monitor charts, so anyone reading that series needs to know which feedback key and which window are contaminated. Re-scoring means a new rule going forward, or an offline run over a dataset.
- Why would you filter a rule to root runs only?Because a traced chain produces many nested runs — retriever, several LLM calls, tool calls — and an unfiltered rule treats each as a candidate. You would pay for judge calls on intermediate steps that have no meaningful answer to score, and your feedback chart would mix per-step scores with end-to-end scores. Scoring the root run once gives one score per user-visible interaction.
saying these in an interview costs you the question
- Thinking rules run inside the application request path
- Assuming a new rule retroactively scores historical runs
- Leaving the sampling rate at 1.0 on an unfiltered project
- Believing deleting a rule removes the feedback it wrote
- Describing rules as offline dataset evaluation under another name