In Langfuse, how do you attach a score to a trace or an observation?
answer
- a named value hung off a trace
- scores are their own records
- inside the span versus after the fact
- keep the trace id to score later
- score_trace(trace_id=...) from another process
basics
~20 sA Langfuse score is a named value hung off a trace or off one observation inside it. From inside an active observation call score_current_span() or score_current_trace(); from anywhere else call score_trace(trace_id=...) with a trace id you stored earlier.
solid answer
~50 sIn Langfuse a score is its own record: a `name`, a `value`, an optional `data_type` and `comment`, pointed at a trace id and optionally at one observation id. While your code is inside an instrumented block, the client methods `score_current_span()` and `score_current_trace()` attach a score to that observation or to the whole trace; if you are holding the span object, `span.score()` and `span.score_trace()` do the same two things. Because scores are separate records, they do not have to be written while the trace is open. For an end-user thumbs-up that arrives minutes later, capture the trace id during the request (`span.trace_id`, or `langfuse.get_current_trace_id()` inside an `@observe` function), persist it next to the message, and later call `langfuse.score_trace(trace_id=..., name='user_feedback', value=1)`. One trace can carry many differently-named scores at once, from code, from a judge evaluator and from a human reviewer.
code
python · 13 linesfrom langfuse import get_client
langfuse = get_client()
with langfuse.start_as_current_span(name="handle-request") as span:
answer = "Paris"
span.score(name="retrieval_hit", value=1, data_type="NUMERIC")
span.score_trace(name="answered", value="yes", data_type="CATEGORICAL")
trace_id = span.trace_id
# later, from a different request handler
langfuse.score_trace(trace_id=trace_id, name="user_feedback", value=1)
langfuse.flush()go deeper
Be able to say that a score is a name plus a value attached to a trace, and show the two cases: scoring while inside the traced code, and scoring later with a stored trace id.
Explain why scores are separate objects from traces, when to score an observation instead of the trace, and how deferred user feedback is joined back by trace id.
Show the production wiring: persist the trace id with the outgoing message, flush before short-lived processes exit, and keep score names stable so dashboards and alerts stay meaningful.
Own the policy question of which signals are worth capturing at all — user feedback, implicit signals, judge scores — and how they are named so that different teams' scores compose in one project.
## What a score is Langfuse keeps *what happened* and *how good it was* in different objects. A trace (with its nested observations) records the execution. A **score** is a separate record attached to it, carrying a `name` such as `user_feedback` or `resolved`, a `value`, an optional `data_type` (NUMERIC, BOOLEAN or CATEGORICAL), an optional `comment` explaining the value, and a pointer: always a trace id, plus an observation id when the score is about one step rather than the whole request. Every score also records where it came from. The API stores a source of `API` (written by your code through the SDK or the HTTP API), `EVAL` (written by a Langfuse judge evaluator) or `ANNOTATION` (written by a human in the UI). That field is why one trace can hold a model-written and a human-written opinion of the same thing without them being confused for each other. ## Scoring from inside the trace When you are executing inside an instrumented block, Langfuse already knows which observation is current, so you only supply name and value: - `langfuse.score_current_span(name=..., value=...)` attaches the score to the currently active observation. - `langfuse.score_current_trace(name=..., value=...)` attaches it to the enclosing trace instead. - If you hold the span object returned by `start_as_current_span()` or `start_as_current_observation()`, `span.score(...)` and `span.score_trace(...)` are the same two operations. The choice matters for analysis. A retrieval-quality score belongs on the retriever observation, because that is the step it judges; an overall satisfaction score belongs on the trace, because no single step owns it. Trace-level scores are what the dashboards aggregate per user, per session and per release. ## Scoring from outside the trace The common production case is deferred: the request finished, the answer went out, and the user reacted later. The score API is built for this. During the request, capture the trace id and store it with the message you rendered. When feedback arrives, call `langfuse.score_trace(trace_id=stored_id, name='user_feedback', value=1)` from an entirely different process. There is no need to reopen or update the trace itself. The only hard requirement is that the trace id must be the real one. `span.trace_id` and `langfuse.get_current_trace_id()` give it to you; inventing an id later produces a score that points at nothing and silently disappears from every chart. ## Delivery and timing The SDK queues score events and ships them in the background, so a score written just before a short-lived process exits can be lost. In a script, a serverless handler or a test, call `langfuse.flush()` before the process ends. In a long-running server this is handled for you at shutdown. ## Mistakes worth avoiding Do not reuse one score name for two different meanings; the name is the aggregation key, and mixing meanings makes the chart meaningless. Do not assume a score has to be numeric — categorical labels and booleans are first-class. And do not build your own side table of feedback keyed by request id when Langfuse already joins feedback to the exact execution that produced it: the value of scoring in Langfuse is that from a bad score you can click straight through to the prompt, the retrieved context and the model call that earned it.
- A user clicks thumbs-up an hour after the response — how does that score reach the right trace?Capture the trace id during the request with span.trace_id or langfuse.get_current_trace_id(), store it alongside the message you rendered, and when the click arrives call langfuse.score_trace(trace_id=stored_id, name='user_feedback', value=1). The trace is long since closed; scores are independent records, so nothing needs reopening.
- Would you put a retrieval-relevance score on the trace or on the retriever observation?On the retriever observation, because that is the step it judges — that way you can rank retrieval quality independently of the final answer. Reserve trace-level scores for judgements no single step owns, such as overall satisfaction or task success, since those are what per-user and per-release aggregations chart.
- Your scores never appear when you write them from a short script. What is the likely cause?The SDK batches score events in a background queue, and the process exits before the queue is flushed. Call langfuse.flush() before the script ends, or make sure the client shuts down cleanly. A second possibility is a trace id that never existed, which produces a score pointing at nothing.
saying these in an interview costs you the question
- Thinks a score must be written before the trace closes
- Says you must update the trace to add feedback
- Invents a fresh trace id when late feedback arrives
- Assumes every score has to be a number
- Treats observation-level and trace-level scores as identical