In Langfuse, when do you record a GENERATION instead of a plain SPAN?
answer
- both are timed, one is billed
- model, parameters, tokens, cost
- the analytics only see one type
- under-reporting, not an error
- ten types, not three
basics
~20 sUse GENERATION for a call to a model. It carries model name, model parameters, prompt and completion, token usage and cost, which is what feeds Langfuse's token and spend analytics. SPAN is a generic timed step with none of those fields.
solid answer
~50 sSPAN and GENERATION are both timed observations, but GENERATION is the type reserved for a model invocation and it has extra structured fields: `model`, the model parameters, input messages, the completion, and usage details in tokens from which Langfuse derives cost using its model pricing definitions. Record a model call as a SPAN and everything still looks fine in the waterfall — the latency is right, the payloads are there — but the call contributes nothing to token or cost aggregations, so your spend chart quietly under-reports. The other typed observations exist for the same reason at a different level: RETRIEVER, TOOL, AGENT, CHAIN, EMBEDDING, EVALUATOR and GUARDRAIL let the UI and your filters distinguish a retrieval from a guardrail check rather than showing a wall of anonymous spans. Type is metadata that the platform actually acts on, not decoration.
code
python · 16 linesfrom langfuse import get_client
langfuse = get_client()
with langfuse.start_as_current_observation(
as_type="generation",
name="summarise",
model="gpt-4o-mini",
) as gen:
gen.update(
input=[{"role": "user", "content": "summarise this"}],
output="a summary",
usage_details={"input": 120, "output": 35},
)
langfuse.flush()go deeper
Know that a call to a model should be recorded as a GENERATION, and that a SPAN is for any other timed step in your code.
List what GENERATION carries beyond a span — model, parameters, prompt and completion, token usage and cost — and explain that Langfuse's token and spend analytics are built from generations only.
Diagnose the under-reporting failure: calls filed as spans, streamed calls with usage never requested, retries folded together, and models with no pricing definition.
Set the accounting convention — one generation per provider call including retries, usage always from the provider response, pricing defined for in-house models — so that platform spend reconciles with the invoice.
## The distinction Every Langfuse observation has a start, an end, a name, input, output and metadata. What differs by type is the *extra structure* the platform understands. **SPAN** is the generic case: a timed step, nothing more. Use it for retrieval orchestration, parsing, post-processing, a database call, a chunk of business logic you want to see on the waterfall. **GENERATION** is a model invocation, and it adds: - `model` — the model identifier actually used. - **model parameters** — temperature, max tokens, top_p, stop sequences and so on. - **input** — the messages or prompt sent. - **output** — the completion returned, including tool calls. - **usage details** — input and output token counts, ideally taken from the provider's response rather than estimated. - **cost** — derived from the model plus usage against Langfuse's pricing definitions, or supplied explicitly when you know better than the table does. ## Why the choice has teeth Token and cost analytics in Langfuse are built on generations. If a model call is filed as a SPAN: - it appears in the trace with correct timing, so nothing looks broken; - its tokens are absent from token charts; - its cost is absent from spend charts and from per-user and per-session cost; - model-level breakdowns ("how much are we spending on the reranker model?") silently omit it. The failure mode is under-reporting rather than an error, which is why it survives so long in real deployments. The tell is a Langfuse spend figure that is comfortably below the provider invoice while trace volume looks right. Going the other way is less harmful but still wrong: filing a database call as a GENERATION gives you an observation with an empty model and no usage, which pollutes model breakdowns with a phantom entry. ## The rest of the type set The v4 SDK's `ObservationType` has ten members. Beyond SPAN, GENERATION and EVENT: - **RETRIEVER** — a retrieval step. Marking it as such is what lets you look at retrieval latency and outputs as a class across traces. - **TOOL** — a tool or function invocation by an agent. - **AGENT** — an agent's turn or loop. - **CHAIN** — a composite step made of other steps. - **EMBEDDING** — an embedding call, which is a model call with a different shape from a chat completion. - **EVALUATOR** — a step whose job is judging an output. - **GUARDRAIL** — a safety or policy check. None of these change billing the way GENERATION does, but they change what you can find. In an agent trace with forty nodes, being able to filter to tool calls or to guardrail checks is the difference between reading a trace and hunting through one. ## Getting usage right Two rules keep the cost numbers honest: 1. **Take usage from the provider's response.** Integrations do this for you. Hand-instrumented calls often do not, and a locally tokenised estimate will diverge from the invoice — differently per model, and worst on the models where it matters. 2. **Watch streaming.** Several providers omit usage from streamed responses unless you explicitly request it, so streamed generations are where missing usage concentrates. When a model's pricing is not in Langfuse's definitions — a fine-tune, a self-hosted model, a negotiated rate — the derived cost will be missing or wrong, and the answer is to define the model's pricing or to supply cost details on the generation rather than to accept the gap. ## Practical guidance A reasonable house rule: exactly one GENERATION per provider call, no more and no fewer; a SPAN for each logical step of your own; specific types where they exist for what the step really is; EVENT for things that happened at an instant and have no duration. Retries deserve their own generations rather than being folded into one, because two attempts cost two calls' worth of tokens and you want both on the bill. ## Interview framing Don't just say "generation is for LLM calls". Say what it carries that a span does not, and name the consequence — cost and token analytics are built on generations, so the wrong type produces a plausible-looking trace and a wrong spend number. Then mention that the type set is broader than three, and that the other types are about findability rather than billing.
- Langfuse reports far less spend than the provider invoice, yet trace volume looks right. Where do you look?Start with model calls filed as plain spans — they show latency but contribute no tokens or cost. Then check streamed generations, which lose usage when the provider was not asked to return it, and models whose pricing Langfuse has no definition for, where usage is present but cost cannot be derived. Retries folded into a single generation are a fourth, smaller source.
- How should a retried model call be recorded?As separate generations, one per attempt. Two attempts consume two calls' worth of tokens and both appear on the invoice, so collapsing them into one observation understates spend and hides the retry rate entirely. Keeping them distinct also lets you see how often the first attempt fails and what the failures cost, which is usually the more actionable number.
- What do the non-billing observation types like RETRIEVER and GUARDRAIL actually buy you?Findability and consistent rendering. In a large agent trace they let you filter to "all tool calls" or "all guardrail checks" across traces instead of scanning anonymous spans, and they let the UI display each step in the shape it expects. They do not change cost accounting — that remains specific to generations.
saying these in an interview costs you the question
- Files LLM calls as generic spans and trusts the cost chart
- Estimates tokens locally instead of using the provider's response
- Collapses retries into a single generation
- Thinks observation type is only a display label
- Uses GENERATION for non-model steps like database calls