In Langfuse, what is a trace and what are the observations inside it?
answer
- two levels, one tree
- the outer one is per request
- steps inside are typed
- GENERATION is the model call
- ten ObservationType members, not three
basics
~20 sA Langfuse trace is one end-to-end unit of work, typically one request or agent run. Observations are the timed steps inside it: nested and typed, with GENERATION recording an LLM call and SPAN any other step.
solid answer
~50 sLangfuse's data model has two levels. A **trace** is the top-level container for one unit of work — an API request, a chat turn, an agent run — and carries trace-level attributes such as name, input, output, user id, session id, tags and metadata. Inside it sit **observations**: individually timed, nestable records of each step, forming a tree under the trace. Every observation has a type. `ObservationType` in the v4 Python SDK has ten members — SPAN, GENERATION, EVENT, AGENT, TOOL, CHAIN, RETRIEVER, EVALUATOR, EMBEDDING and GUARDRAIL — and the type is what tells the UI whether a node is a generic step, a model call, an instantaneous marker or a retrieval. GENERATION is the one that carries model name, parameters, token usage and cost, which is why the token and cost dashboards depend on it. Scores can attach to a whole trace or to a single observation.
go deeper
Be able to say plainly that one trace equals one request or agent run, and that observations are the steps inside it. Name SPAN, GENERATION and EVENT and say what each is for.
Explain how nesting happens automatically from a currently-active observation, and why the observation type matters: only GENERATION carries model, tokens and cost into the analytics.
Show judgement about trace boundaries in real systems — one trace per user-visible turn, sessions to stitch turns, and enough observation granularity to localise a regression without exploding the tree.
Own the convention: which unit of work is a trace across all services, a house style for observation names and types, and required trace attributes, so cost and quality can be sliced consistently org-wide.
## The two levels Langfuse stores what your application did as a **trace** containing a tree of **observations**. A trace is the outermost unit: one request, one chat turn, one agent run, one background job. It is the row you see in the Traces list, and it holds the attributes you filter and group by — a name, trace-level input and output, a user id, a session id, tags, arbitrary metadata, a version, and an environment. If you cannot answer "which single thing did the caller ask for here?", your trace boundary is probably drawn in the wrong place. An observation is one timed step inside that trace. Observations nest: a parent observation for the whole pipeline, children for retrieval, for the model call, for a tool invocation, and so on. Each one records a start and end time, its own input and output, its own metadata, and — depending on type — extra structured fields. The waterfall view in the UI is exactly this tree, drawn on a time axis. ## Observation types In the v4 Python SDK, `ObservationType` has ten members: - **SPAN** — the generic timed step. Anything that takes time and is not something more specific. - **GENERATION** — a call to a model. Carries model name, model parameters, prompt and completion, token usage and cost. - **EVENT** — a point in time, not a duration. Use it to mark that something happened (a cache miss, a fallback, a validation failure). - **AGENT**, **TOOL**, **CHAIN**, **RETRIEVER**, **EMBEDDING** — semantic types for the parts of an LLM application, so the UI can render and filter them meaningfully. - **EVALUATOR** — a step whose job is to judge an output. - **GUARDRAIL** — a safety or policy check. Older write-ups describe Langfuse as having "three kinds of observation" (span, generation, event). That was true of much older SDKs; it is not the current model, and an interviewer working with a recent deployment will expect you to know the typed set exists even if you only ever use three of them in anger. The practical consequence of type: **only GENERATION aggregates into token and cost analytics.** Record a model call as a plain SPAN and you will still see its latency in the waterfall, but its tokens and its money will be missing from every chart. That mistake is common enough that it is worth saying out loud in an interview. ## How the tree gets built You rarely construct the tree by hand. The three usual routes are: 1. The `@observe` decorator on a function, which makes each call one observation and nests it under whatever observation is already active. 2. `start_as_current_observation(as_type=...)` as a context manager, for explicit control inside a function. 3. Integrations — most notably importing the OpenAI client from `langfuse.openai`, which turns every model call into a GENERATION automatically. All three share one mechanism: there is a notion of the *currently active* observation, and anything created while it is active becomes its child. That is why nesting works without you passing parent ids around. ## Sessions and users sit above the trace A trace is one turn; a conversation is many turns. Langfuse expresses that with `session_id` on the trace — the Sessions view stitches every trace sharing an id into one timeline, which is how you answer "what did this customer's whole conversation look like?". `user_id` does the same for per-user volume, cost and quality. Both are trace-level attributes, not observation-level ones. ## Where the data comes from In v4 the SDK's ingestion is built on OpenTelemetry: observations travel as spans carrying Langfuse-specific attributes, and the server reconstructs the trace tree from them. For day-to-day use you do not need to think about that layer — but it explains two visible behaviours. First, export is asynchronous and batched, so data appears after a short delay and a process that exits abruptly can lose it. Second, a trace id and observation ids are generated client-side, which is what lets you reference a trace (for example to attach a score to it) before the server has ever seen it. ## What to say in an interview Define the two levels, name the types you actually use and mention that the set is larger than three, state that GENERATION is what makes cost and token analytics work, and finish with session and user ids as the trace-level attributes that turn a pile of requests into something you can slice by conversation and by customer.
- What happens to your cost dashboard if an LLM call is recorded as a SPAN rather than a GENERATION?You lose it. The span still shows up in the waterfall with correct latency, but model name, token usage and cost live only on GENERATION observations, so that call contributes nothing to token or spend aggregations. The usual symptom is a trace that looks complete while the project's cost chart under-reports — and the fix is to change the observation type, not to add metadata.
- Where do scores attach — to a trace or to an observation?Either. A score can target the whole trace, which is what you want for an end-user thumbs-up or an overall quality judgement, or a single observation, which is what you want when you are judging one retrieval step or one generation inside a longer run. Choosing the narrower target is what later lets you say which step regressed rather than only that the run did.
- How do many traces get grouped into one conversation?By setting the same session id on each trace. Each turn stays its own trace with its own tree, and the Sessions view replays them in order as a single timeline. Grouping is trace-level, so you set it on the trace rather than on individual observations, and you set it at the start of the turn so the whole trace carries it.
saying these in an interview costs you the question
- Says Langfuse observations come in exactly three kinds
- Calls a whole conversation one trace
- Records model calls as generic spans and expects cost data
- Thinks session id is set per observation
- Confuses a trace with a Langfuse project or dataset