In Langfuse, how do you link the prompt version used to the generation it produced?
answer
- Pass the object, not the rendered text
- One keyword argument at generation time
- Turns a version into a metrics dimension
- Labels move; the recorded number does not
- Fallback traffic has nothing to record
basics
~20 sPass the fetched prompt client object to the generation — for example langfuse.start_as_current_generation(name=..., prompt=prompt) or langfuse.update_current_generation(prompt=prompt). Langfuse then records the prompt name and version on that generation, so latency, cost and scores can be broken down per version.
solid answer
~40 sYou hand the prompt object itself, not its text, to the generation you create. In the v4 Python SDK that is `langfuse.start_as_current_generation(name="answer", model=..., prompt=prompt)`, or `langfuse.update_current_generation(prompt=prompt)` when the generation already exists; the OpenAI drop-in integration takes it as the `langfuse_prompt=` keyword on the completion call. Langfuse stores the prompt's name and version number on the generation, which is what turns a prompt into something measurable: you can see, per version, how many generations it produced, what it cost, how slow it was and how it scored. Without the link, moving the `production` label is an untraceable change — a quality regression shows up in aggregate with nothing tying it to the prompt that caused it. It is also the only reliable record of what actually served a request, since labels move and caches lag.
code
python · 13 linesfrom langfuse import Langfuse
langfuse = Langfuse()
prompt = langfuse.get_prompt("support-reply")
with langfuse.start_as_current_generation(
name="answer",
model=prompt.config.get("model", "gpt-4o-mini"),
prompt=prompt, # links name + version to this generation
) as generation:
text = prompt.compile(question="Where is my order?")
# ... call the model with `text` ...
generation.update(output="...")go deeper
Know that Langfuse can record which prompt version produced a generation, and that you enable it by passing the fetched prompt object to the generation rather than only its text.
Name the mechanism: prompt= on start_as_current_generation, update_current_generation(prompt=...), or langfuse_prompt= on the OpenAI integration, and explain that Langfuse stores the prompt name and version on the generation.
Explain why it matters operationally — per-version cost, latency and score breakdowns, regression attribution after a label move, and the fact that the recorded version is the only reliable answer to what actually served a request given caching and fallbacks.
Treat the prompt version as a deployment identifier that changes outside your build, and require it as a telemetry dimension wherever prompts ship without a deploy. Also own the interpretation risk: version comparisons on live traffic are unrandomised slices, not experiments.
## The gap the link closes Deploy-free prompt changes create an attribution problem. A label move is a release that leaves no trace in your version control, your deploy log or your change management system. A week later somebody notices answers have got worse. Which change? There were four prompt promotions and two code deploys. Without a link between the prompt version and the generations it produced, that question is answered by guesswork. Linking makes the prompt a first-class dimension of your telemetry, the same way a service version or a build SHA is. ## How to do it You pass the **prompt client object** returned by `get_prompt`, not the compiled string. The object carries `.name` and `.version`, which is what Langfuse persists on the generation. - Creating the generation with the link in place: `with langfuse.start_as_current_generation(name="answer", model="gpt-4o-mini", prompt=prompt) as gen: ...` - Attaching it to a generation that already exists in the current context: `langfuse.update_current_generation(prompt=prompt)` - Through the OpenAI drop-in integration, as a keyword on the call itself: `openai.chat.completions.create(..., langfuse_prompt=prompt)` The important discipline is to link the object you actually used for this call. Re-fetching the prompt to link it, or linking a differently-labelled copy, produces a record that is confidently wrong — which is worse than no record. ## What the link buys you 1. **Per-version metrics.** Generations roll up by prompt version: volume, token usage, cost and latency. "Version 12 is 30% more expensive per call" is a question you can only ask if the dimension exists. 2. **Per-version quality.** When evaluation scores land on traces, they inherit the version dimension, so you can compare how versions scored on live traffic rather than only in an offline experiment. 3. **Regression attribution.** A step change in a metric can be aligned with the moment the version changed, which is exactly the correlation you cannot make from a label alone. 4. **Ground truth during rollout.** Because caching means a promotion rolls across replicas over roughly a TTL, and because a fallback may have served some requests, the only reliable answer to "what prompt served this request?" is the version recorded on that generation. 5. **Auditability.** For a regulated surface, being able to show which exact prompt text produced a given output is often a requirement, not a nicety. ## Failure modes - **Not linking at all.** The most common state. Everything works, nothing is measurable per version, and the cost only becomes visible during an incident. - **Linking the compiled string instead of the object.** The compiled text has no version number; it is just a string. The API expects the client object. - **Fallback traffic.** A prompt served from the `fallback` argument has no server-side version, so those generations carry no version attribution. When comparing versions across a period, remember an outage can leave a hole. - **Fan-out.** One request may compile several prompts across several generations. Each generation should be linked to the prompt that produced it; linking the whole trace to one prompt loses the resolution you wanted. - **Version churn in a comparison window.** If a label moved mid-window, a per-version comparison is comparing unequal, non-random slices of traffic — the versions saw different times of day and possibly different user populations. That is a caveat about interpreting the numbers, not a reason to skip the link. ## The mental model Think of the prompt version as a deployment identifier that happens to change independently of your build. Everything you would normally want to slice by build SHA — error rate, latency, cost, quality — you want to slice by prompt version too, and the link is what makes that possible. It is one keyword argument, added at the moment you already have both objects in hand, and it is the difference between prompt management being an operational tool and being a shared text editor.
- You linked the prompt but see no version on some generations. What is the likely cause?Those calls were served from the `fallback` argument during a period when Langfuse was unreachable. A fallback is a local literal with no server-side version, so there is nothing to record. Counting fallback use as a metric makes this explainable rather than mysterious, and it tells you which windows to exclude from a version comparison.
- Why link the prompt object rather than just putting the version number in metadata?You could stuff a number into metadata, but the first-class link is what makes the prompt a queryable dimension in Langfuse's prompt views — per-version volume, cost, latency and scores — and it keeps name and version consistent rather than depending on every call site to format a string the same way.
- A request compiles three different prompts across three model calls. How should linking work?Link each generation to the prompt that produced it. A trace-level link to one prompt would attribute all three calls to it and destroy the resolution you wanted; the point of the dimension is being able to say which specific prompt version drove which specific call's cost and quality.
- Two prompt versions ran in the same week and one scores better. What should you check before promoting it?That the comparison is not confounded. Because a label move rolls out over a cache TTL and then serves everything, the two versions saw different time windows and possibly different traffic mixes rather than a randomised split. Confirm volumes are comparable, look for time-of-day or cohort effects, and prefer a controlled offline run on a shared dataset before treating the delta as real.
saying these in an interview costs you the question
- Passes the compiled prompt string instead of the prompt object
- Never links, then cannot attribute a quality regression
- Assumes the current production label reveals what served an old request
- Links one prompt at trace level despite several generations
- Compares version metrics without noticing a mid-window label move