How do you set up a Langfuse LLM-as-judge evaluator to score production traces?
answer
- configured in the project, not in code
- a prompt template with variables
- binding variables to trace fields
- filter plus sampling rate
- the judge is its own model bill
basics
~20 sYou configure an evaluator in the Langfuse project: pick or write an evaluator template whose prompt has variables, map those variables onto fields of the trace, choose a judge model, then target it with a filter over traces plus a sampling rate. Matching traces get a score whose source is EVAL.
solid answer
~60 sA Langfuse judge evaluator is configured in the project rather than in your application code. Four decisions make it up. First the **template**: either one from Langfuse's managed library (hallucination, helpfulness, relevance, toxicity and similar) or your own prompt, containing `{{variables}}` and asking for a score plus a reason. Second the **variable mapping**: each template variable is bound to a field of the trace or observation — the input, the output, retrieved context in metadata — which is the step that most often goes wrong, because a variable mapped to a field your traces do not populate yields empty judgements. Third the **judge model**, which needs its own provider credentials in the project and is normally a different, cheaper or stronger model than the one under test. Fourth the **target**: a filter selecting which traces qualify, plus a sampling rate, since one judge call per sampled trace is real money at production volume. Results land as scores with source `EVAL`, carrying the judge's reasoning in the comment, so you can filter traces by a bad judge score and read why.
go deeper
Know that Langfuse can score traces automatically with an LLM judge configured in the project, and that the result appears as a score on the trace with the judge's reason.
Walk through the setup: template with variables, mapping those variables to trace fields, judge model with its own credentials, target filter and sampling rate.
Show the operational side — verifying what the judge actually receives, backfilling over historical traces, and treating sampling rate as the cost dial rather than a default.
Own which dimensions get judged online at all, what the judge budget is against the inference budget, and where a deterministic self-written score should replace an LLM judge entirely.
## Where the evaluator lives Unlike a metric you call in code, a Langfuse LLM-as-judge evaluator is a piece of **project configuration**. It watches traces as they arrive and writes scores back onto them. Nothing changes in the application; the same traces you were already sending become the evaluator's input. That is the whole appeal, and also the source of its failure modes, because the evaluator can only see what your instrumentation happened to record. ## The four decisions **Template.** A template is the judge prompt: an instruction, one or more `{{variables}}` to be filled from the trace, and an output contract (a score plus a short reason). Langfuse ships a managed library of templates for common dimensions, and you can copy one and edit it or write your own. Templates are versioned in the project, so a judge prompt change is a visible event rather than a silent drift. **Variable mapping.** Each variable must be bound to something concrete on the trace or on a chosen observation type: the trace input, the trace output, a field inside metadata, the output of a particular kind of observation. This is the configuration step people get wrong. If your RAG traces put retrieved chunks in a retriever observation but the mapping points at trace metadata, the judge receives an empty context and confidently scores hallucination on nothing. After creating an evaluator, always inspect the first few resulting scores and read the judge's input on one trace. **Judge model.** The judge needs credentials for a model provider configured in the project. Choose it deliberately: judging is usually a cheaper task than generation, but too weak a judge produces noise, and using the exact model under test invites the model to like its own style. The judge's cost is separate from your application's inference bill and is easy to forget when you turn the evaluator on across all traffic. **Target and sampling.** An evaluator is pointed at a filtered slice of traces — an environment, a tag, a trace name, a time window — and given a sampling rate. Langfuse can also run an evaluator over historical traces, not just new ones, which is how you backfill a new dimension over last month's traffic. Sampling is the cost dial: at a few thousand traces an hour, one judge call each is a bill you did not have yesterday. A few percent of traffic is usually enough to move a trend line, and you can always raise it for a suspicious cohort. ## Reading the result Each run produces a score whose source is `EVAL`, with the judge's explanation in the comment. Two things follow. First, aggregate views over that score name give you the trend, sliced by any trace attribute you recorded — user, release, prompt version. Second, the individual bad scores are clickable: you land on the trace, see the prompt, the retrieved context and the model call, and read the judge's stated reason next to them. That round trip from a metric to the offending execution is what makes an online judge worth configuring at all. ## The limits An evaluator on live traffic can only judge what needs no ground truth — groundedness against the retrieved context, relevance to the question, tone, refusal. Anything that compares against a correct answer has nothing to compare with on production traffic, and belongs to an offline dataset run where the expected output exists. Whether the resulting score is *valid* — how the judge is biased, whether its numbers track human labels — is a separate discipline from configuring it, and the configuration screen will happily produce a well-formed number for a badly conceived judge. If you need judging logic Langfuse's templates cannot express, the escape hatch is to compute the score yourself in your own code or job and write it with the score API; a self-computed score is indistinguishable to every dashboard except for its source field.
- A new evaluator returns the same middling score on almost every trace. Where do you look first?At the variable mapping. Open one scored trace and inspect what the judge actually received: a variable bound to a field your traces do not populate arrives empty, and a judge given no context falls back to a bland mid-range verdict. Fix the mapping, then re-run the evaluator over historical traces to confirm the distribution spreads out.
- How do you keep a judge evaluator's cost under control at production volume?Lower the sampling rate — it is the main dial, since cost is one judge call per sampled trace — and narrow the target filter to the environment and trace names you actually care about. Pick a cheaper judge model where the dimension is easy, and raise sampling temporarily on a cohort you are investigating rather than leaving it high everywhere.
- When would you write the score yourself instead of configuring an evaluator?When the judgement is deterministic or needs your own code: schema validity, a regex or tool-call check, a business rule, or a score computed from a system you own. Write it with the score API and it behaves like any other score in dashboards, differing only in its source field. Reserve the LLM judge for judgements that genuinely need language understanding.
saying these in an interview costs you the question
- Thinks evaluators run inside the application SDK
- Skips the variable mapping and trusts defaults
- Runs the judge on 100% of traffic without costing it
- Uses the model under test as its own judge without noting the risk
- Expects an online judge to compare against a correct answer