What is a LangSmith annotation queue, and how do runs end up in one?
answer
- structured human review, not scrolling
- rule, manual, or SDK push
- rubric lives on the queue
- verdict becomes human-sourced feedback
- throughput is reviewer hours
basics
~20 sAn annotation queue is a review worklist of production runs. A human opens it, sees one run's input and output at a time, applies the queue's rubric, and their verdict is written back as human-sourced feedback on that run.
solid answer
~50 sA queue is where humans look at real traffic in a structured way instead of scrolling the runs table. Runs get in three ways: an automation rule that samples matching runs into the queue continuously, a manual selection from the project's runs table, or the SDK — `Client.create_annotation_queue(name=...)` to make it and `Client.add_runs_to_annotation_queue(queue_id, run_ids=[...])` to fill it. The queue defines the rubric: the feedback keys a reviewer must fill in, and any instructions for them. Reviewers are handed runs one at a time so two people do not grade the same one, and each verdict lands as feedback on the run, marked as human-sourced so it is distinguishable from a judge's score on the same key. The reviewer can also write a correction and send the run to a dataset, which is the point of the whole exercise: a queue is not just for measuring quality, it is the manufacturing line that turns production failures into labelled examples you can evaluate against offline.
code
python · 13 linesfrom langsmith import Client
client = Client()
queue = client.create_annotation_queue(
name="support-bot-weekly-review",
description="Runs where the judge and the user disagreed",
)
client.add_runs_to_annotation_queue(
queue_id=queue.id,
run_ids=["9f7c4a2e-1b3d-4c5e-8a90-1f2b3c4d5e6f"],
)go deeper
Know that an annotation queue is a list of production runs a person reviews, and that their verdict is saved as feedback on the run rather than as a private note.
Name the three ways runs enter — an automation rule, manual selection from the runs table, and add_runs_to_annotation_queue — and explain that the queue carries the rubric of feedback keys reviewers fill in.
Show that you feed the queue from signal rather than raw sampling, that you size it to reviewer throughput, and that you route reviewed runs with their corrections into a dataset so the failure becomes a permanent test case.
Own the review programme: who reviews, what their rubric means, how many hours a week it costs, and what happens to production content shown to a human when that traffic may be regulated.
## What it is for Automated scoring tells you a number moved. It cannot tell you what "wrong" looked like, and it cannot produce a trustworthy label. Human review does both, and it is expensive, so it needs to be structured rather than ad hoc. An annotation queue is that structure: a worklist of runs, a fixed rubric, and one reviewer at a time. Without it, human review in practice means someone filtering the runs table, reading a few traces, and forming an impression that never becomes data. The queue exists to make the reading produce feedback records. ## Getting runs into a queue Three routes, and a mature setup uses all three. **Automation rule.** The continuous route. A rule filters the project's runs and samples a fraction of them into the queue. This is how you get a steady, unbiased trickle of representative traffic to look at. **Manual selection.** From the runs table, after filtering to something interesting — errors, thumbs-downs, a specific customer's session, everything from the hour of an incident. This is the investigative route. **SDK.** `Client.create_annotation_queue(name=..., description=...)` creates one; `Client.add_runs_to_annotation_queue(queue_id, run_ids=[...])` pushes a list of runs into it. This is the programmatic route, used when your own logic decides what deserves review — for example, a nightly job that queues every run where a judge and a user disagreed. A useful pattern is chaining: a rule filters on *existing feedback* rather than raw traffic, so runs a judge scored badly, or runs the user thumbed down, are automatically escalated to a human. That gives you a queue full of probable failures instead of a queue full of ordinary successes, which is a very different use of reviewer time. ## The reviewing experience A reviewer opens the queue and is shown runs one at a time — inputs, outputs, and the trace underneath if they need to see which step went wrong. Runs are handed out individually rather than as a shared list, so two reviewers do not silently grade the same run and produce conflicting feedback on the same key; this is a property of the queue, not something you should be coordinating by convention in a chat channel. The queue carries the rubric: which feedback keys the reviewer fills in, what type each is (a number, a category, free text), and the instructions telling them what "good" means for this application. Keeping the rubric on the queue rather than in each reviewer's head is what makes two reviewers' output comparable. ## What a review produces The verdict is written as feedback on the run, under the queue's keys, recorded as human-sourced. That last part matters. If your judge writes `correctness` and your reviewers also write `correctness`, both series exist on the same runs, and you can compare them — which is how you find out whether the judge is worth trusting. If you cannot tell the two apart, you have destroyed exactly that comparison. The reviewer can also add a correction — what the answer should have been — and send the run to a dataset. That is the trace-to-dataset step: the run's inputs become the example's inputs, the correction becomes the expected output, and from then on every offline experiment you run includes this real failure. A queue whose reviewed runs never reach a dataset is generating opinions instead of assets. ## Capacity is the real constraint A judge scales with money; a queue scales with people. A reviewer might get through a few dozen substantive runs an hour, which sets a hard ceiling on the rule's sampling rate feeding it. The characteristic failure is a rule at a rate nobody sized, producing a queue with thousands of pending runs that everyone stops opening — worse than no queue, because it looks like coverage. Size the rate to your reviewers' actual weekly hours, and prefer a small queue of hard cases over a large queue of average traffic. Who reviews also matters. Domain experts produce labels worth calibrating a judge against; whoever was free that afternoon produces noise with a timestamp. If the rubric needs domain knowledge, engineers are not the right reviewers, and the queue's instructions have to be written for the people who actually are. ## Privacy A queue puts raw production inputs and outputs in front of a human, which is a data-handling decision, not a UI feature. If traffic can contain personal or regulated data, who can open the queue and what is masked before it gets there is a question you must have an answer to before you turn one on. ## The shape of a good answer Say what it is (structured human review), name the three entry routes, say the output is human-sourced feedback on the run, and finish with the two things that separate people who have run one: the queue is how corrections become dataset examples, and its throughput is bounded by reviewer hours rather than by the sampling rate you type in.
- How do you stop two reviewers grading the same run?You do not manage it yourself — the queue hands runs out one at a time rather than exposing a shared list, so a run someone is working on is not simultaneously handed to a colleague. That is a reason to use a queue rather than a filtered runs table plus a spreadsheet, where duplicate and conflicting verdicts on the same key are routine.
- What should happen to a run after a human has reviewed it?If it exposed a real failure, it should leave the queue as a dataset example: the run's inputs plus the reviewer's correction as the expected output. That closes the loop — the failure becomes a permanent case in the offline suite, so the next prompt change is tested against it. Reviews that only produce a score generate opinions, not regression coverage.
- How do you decide what goes into the queue when reviewer time is scarce?Filter on signal rather than sampling raw traffic. Runs the user thumbed down, runs a judge scored low, runs where a judge and a user disagreed, runs from a newly deployed prompt version, or errors. A queue of probable failures uses an hour of expert time far better than a random sample where most runs are unremarkable.
saying these in an interview costs you the question
- Treating a queue as a filtered view rather than something that writes feedback
- Sizing the feeding rule without regard to reviewer hours
- Letting judge and human feedback share a key indistinguishably
- Reviewing runs but never promoting corrections into a dataset
- Putting raw production content in front of reviewers with no access consideration