What is a Langfuse annotation queue, and how does human review become scores?
answer
- a work list of traces for humans
- fixed score configs, not free text
- lands as a score like any other
- source ANNOTATION versus EVAL
- humans review tens per hour
basics
~20 sAn annotation queue is a work list of traces or observations for humans to review in the Langfuse UI. Reviewers apply the project's score configs to each item, and their verdicts are stored as ordinary scores on the same trace, marked with source ANNOTATION.
solid answer
~60 sAn annotation queue turns 'someone should look at these' into a bounded, assignable task list. You create a queue in the project, pick which score configs reviewers may apply, and add items to it — individual traces or observations from the trace view, or a filtered slice of traffic. A reviewer opens the queue, sees the trace with its inputs, outputs and nested steps, and records their judgement using the configured controls; because the configs pin the data type and allowed values, two reviewers cannot invent two vocabularies. The result is a normal score attached to the same trace, distinguished only by source `ANNOTATION`. That single fact is what makes queues valuable: a human label and a judge label sit side by side on identical traces, so you can compare them, and any trace a human confirmed as wrong can be promoted into a dataset item with `source_trace_id` pointing back at it. The binding constraint is throughput — humans review tens of traces an hour, so what enters the queue must be sampled or targeted deliberately.
go deeper
Know that Langfuse lets humans review traces in a queue and that their verdicts become scores on those traces, just like scores written from code.
Explain how score configs constrain what a reviewer can enter, and that annotation scores differ from judge and SDK scores only in their source field.
Show the workflows: comparing human labels against a judge on identical traces, promoting confirmed failures into datasets with source_trace_id, and sampling what enters the queue against real reviewer throughput.
Own the labelling budget and access policy — how much human attention quality gets, split between unbiased sampling and targeted triage, and who is allowed to read trace content at all.
## What a queue is Human review is the only source of labels that does not itself need validating, and it is the scarcest. A Langfuse **annotation queue** is the machinery for spending it well: a named work list of trace or observation items, backed by a defined set of score configs, that reviewers work through in the UI. Creating one involves two choices. Which **score configs** may be applied — that is, what questions the reviewer is being asked, with their data types and allowed values fixed in advance. And what **enters** the queue: items can be added ad hoc from a trace you are looking at, or in bulk from a filtered view of traffic, so a queue can be 'everything a user thumbs-downed this week' or 'a random sample of the new prompt's output'. ## The review loop A reviewer opening a queue item sees the execution itself, not a spreadsheet row: the input, the final output, the nested observations, retrieved context, tool calls. They record their verdict through the controls the score configs define — a bounded slider for a numeric score, a fixed label list for a categorical one — and can leave a comment explaining it. Constraining the input is the point: free-text labelling is where inter-reviewer vocabularies diverge, and no amount of downstream cleaning fixes it. What lands is an ordinary score on the same trace, with source `ANNOTATION` — the same object shape your SDK writes with source `API` and an evaluator writes with source `EVAL`. Every dashboard, filter and drill-down that works for one works for all three. ## Why the shared shape matters Three workflows fall out of it. **Judge checking.** Route the same traces through an LLM judge and a human queue, then compare the two score series on identical traces. Where they disagree, you have concrete cases to look at. (How you decide whether the judge is *good enough* from that comparison — agreement statistics, calibration — is evaluation methodology, not a Langfuse feature.) **Ground-truth harvesting.** A reviewed trace that is confirmed wrong is exactly the input a regression set wants. Create a dataset item whose `source_trace_id` is that trace, and the failure becomes a permanent test case that keeps its provenance. **Triage.** Queues give a bad-output investigation a shape: a filter defines the population, the queue holds the work, and the resulting scores say how much of it was actually bad, rather than the loudest three examples anyone happened to see. ## Operating a queue The hard constraint is human throughput. A reviewer handles on the order of tens of traces an hour, which is nothing against production volume, so the real design question is *what gets queued*. Uniform random sampling gives you an unbiased quality estimate but wastes most of the effort on unremarkable traces; targeting — thumbs-down traces, high-cost traces, a specific tool's failures, a new release's first hour — finds problems faster but tells you nothing about the base rate. Most teams run both: a small unbiased stream for the trend, and targeted queues for investigations. Two more practicalities. Traces contain prompt and completion content, so who may open a queue is a data-access question, not just a workflow one; in regulated settings that pushes toward self-hosting and masking. And queue items are only as good as the trace behind them — if retrieved context is not recorded as an observation, a reviewer cannot tell a retrieval failure from a generation failure, and will label the wrong thing.
- How would you use an annotation queue to check an LLM judge you just configured?Point both at the same slice of traces: let the judge score them and queue the same traces for human review. Because both write scores onto identical traces, you can line the two series up and pull the disagreements for inspection. That gives you concrete failure cases; deciding whether the judge is trustworthy overall is an evaluation-methodology question beyond the tooling.
- A reviewed trace turns out to be a genuine failure. How does it become a regression test?Create a dataset item from it with create_dataset_item(dataset_name=..., input=..., expected_output=..., source_trace_id=<that trace>). The expected output is what the reviewer says it should have been, and source_trace_id keeps the link to the original execution, so anyone later can see why the case exists. From then on it runs in every dataset run of that dataset.
- What should decide which traces enter the queue at production volume?Human throughput is tens of traces an hour, so queueing is a sampling decision. Run a small unbiased random stream to keep an honest base-rate estimate, plus targeted queues for investigations — thumbs-down traces, a new release's first hour, a failing tool. Targeted-only queues find problems fast but tell you nothing about overall quality.
saying these in an interview costs you the question
- Thinks annotations are stored separately from scores
- Lets reviewers type free-text labels instead of using configs
- Assumes humans can review production volume
- Queues only bad traces and then reports a quality rate
- Ignores that trace content shown to reviewers may be sensitive