skip to content

Langfuse

The open-source, self-hostable option for LLM observability: nested traces, prompt management with versioning, evaluation scores, and cost and latency analytics. It comes up whenever the requirement is keeping prompts and traces inside your own infrastructure.

on this pageshow

explore

questions

25

In Langfuse, what does compile() do to a fetched prompt, and what placeholder syntax does it use?

level: juniorimportance: must knowfreq 60%

answer

  1. Fetching is not the same as rendering
  2. Double braces, not single
  3. Text in, string out; chat in, messages out
  4. Missing variable does not raise
  5. Substitution is local and offline

basics

~20 s

compile() substitutes values into a Langfuse prompt's double-curly-brace placeholders, such as {{question}}. A text prompt compiles to a plain string; a chat prompt compiles to a list of role/content message dicts, ready to pass straight to a model SDK.

solid answer

~40 s

`langfuse.get_prompt(...)` returns a prompt client, not a ready-to-send string — the stored template still contains placeholders written as `{{variable}}`. Calling `prompt.compile(question="Where is my order?")` replaces them with the values you pass. For a text prompt the result is a plain string; for a chat prompt (`get_prompt(name, type="chat")`) the result is a list of `{"role": ..., "content": ...}` dicts you can hand directly to a chat completions call. The double-brace syntax is deliberate: it does not collide with Python f-strings or `str.format`, and it survives user content containing single braces. The failure mode to know is that `compile()` is forgiving — forgetting a variable does not raise, so a typo'd keyword ships a prompt with an unfilled slot to the model instead of failing the request. Validate inputs yourself if the variable is load-bearing.

code

python · 13 lines
python
from langfuse import Langfuse

langfuse = Langfuse()

# Text prompt -> plain string
prompt = langfuse.get_prompt("support-reply")
text = prompt.compile(question="Where is my order?")

# Chat prompt -> list of role/content dicts
chat = langfuse.get_prompt("support-chat", type="chat")
messages = chat.compile(question="Where is my order?")
# [{"role": "system", "content": "..."},
#  {"role": "user", "content": "Where is my order?"}]

go deeper

for a junior

Remember the two steps: fetch the prompt, then call compile() with the variables. Know that placeholders are written {{like_this}} and that a chat prompt compiles to a list of messages.

for a middle

Explain why double braces are used, what each prompt type compiles to, and the forgiving-substitution behaviour — a missing variable does not raise, it ships an unfilled slot to the model.

for a senior

Talk about the operational consequence: the variable set is an interface that can now change without a deploy, so you validate required keys at the call site and record the compiled prompt in the trace so silent gaps are diagnosable.

for a principal

Frame it as contract management between prompt authors and application code — who is allowed to introduce a new variable, how that change is reviewed, and whether the organisation wants a schema check enforced before a version can be promoted.

## What you get back from get_prompt `langfuse.get_prompt("support-reply")` does not return a string. It returns a prompt client object holding the stored template on `.prompt`, plus metadata: `.name`, `.version`, `.labels` and `.config`. The template is what a human authored in the Langfuse UI or via `create_prompt`, and it usually contains slots for the runtime data. ## The placeholder syntax Langfuse uses **double curly braces**: `{{question}}`, `{{customer_name}}`, `{{retrieved_context}}`. This choice is not cosmetic: - Single braces are extremely common inside prompt bodies — JSON examples, code samples, format specifications. A single-brace template would break the moment a prompt included `{"role": "user"}` as an illustration. - Python's own f-strings and `str.format` use single braces, so a double-brace template can be stored, logged and moved around without any Python string machinery accidentally interpreting it. ## Compiling a text prompt ``` prompt = langfuse.get_prompt("support-reply") text = prompt.compile(question="Where is my order?") ``` `compile()` takes keyword arguments named after the placeholders and returns a plain `str` with each `{{name}}` replaced by the corresponding value. That string is what you send to the model. ## Compiling a chat prompt A chat prompt is authored as an ordered list of messages, each with a `role` and `content`. Fetch it with `type="chat"`: ``` chat = langfuse.get_prompt("support-chat", type="chat") messages = chat.compile(question="Where is my order?") ``` The result is a list of `{"role": ..., "content": ...}` dicts, with substitution applied inside each message's content. This shape matches what the major chat completion APIs expect, so it usually goes straight into the `messages` argument with no adapter in between. Asking for the wrong `type` is a real mistake: fetching a chat prompt without `type="chat"` gives you the wrong client shape and the compile result will not be what your call site expects. ## The forgiving-substitution trap `compile()` does not enforce that you supplied every variable the template declares. Forgetting one — or misspelling the keyword — does not raise a `KeyError`; it produces a prompt with an unfilled or empty slot, which then goes to the model and comes back as a confidently wrong answer. Nothing crashes, so nothing pages anyone. Mitigations, in rough order of cost: - Build the variable dict in one place and assert on its keys before the call. - Add a thin wrapper around `compile()` in your own code that checks the required keys for that prompt. - Include the compiled prompt in the trace so the unfilled slot is visible when you investigate a bad answer, rather than invisible. The general lesson: because the prompt now changes without a deploy, the set of variables it expects can also change without a deploy. An author who adds `{{tone}}` to the template has silently added a new contract that your code does not fulfil. Treat the variable list as an interface between the prompt author and the application, and review changes to it as you would review a function signature. ## Where compilation sits in the flow The full runtime path is: fetch by name (usually from the local cache) → compile with this request's variables → send to the model → record the generation with a link back to the prompt version used. Compilation is deliberately the cheap, local, synchronous step in the middle; it does no network I/O and does not depend on the Langfuse backend being reachable. ## Framework interop When the downstream consumer is a framework whose own template objects use single braces, the prompt client exposes `get_langchain_prompt()`, which converts the double-brace form into the single-brace form that framework expects. The important point for an interview is the direction of responsibility: Langfuse owns the storage and the `{{...}}` convention, and offers a conversion at the boundary rather than adopting the framework's syntax internally.

  • What is stored in a Langfuse prompt's config field, and why keep it there rather than in your code?
    `config` is an arbitrary JSON blob saved and versioned alongside the prompt text — typically the model name, temperature, max tokens, or a JSON schema. Keeping it there means a prompt rewritten for a different model ships with that model setting attached, so the two cannot drift apart or half-apply. You read it as `prompt.config` at call time and pass the values into your model call.
  • Why does Langfuse use {{variable}} rather than single braces?
    Prompt bodies routinely contain single braces — JSON examples, code snippets, format instructions — and Python's f-strings and str.format also claim single braces. Double braces let a template hold literal `{}` characters and survive being passed through ordinary Python string handling without accidental interpretation or escaping.
  • A prompt author adds a new {{tone}} placeholder but the service never passes it. What happens?
    Nothing fails loudly. `compile()` does not raise for a variable you did not supply, so the model receives a prompt with an unfilled slot and answers anyway — usually plausibly, sometimes wrongly. That is why the variable list should be treated as an interface: assert on required keys in your wrapper, and make the compiled prompt visible in the trace so the gap is diagnosable after the fact.

saying these in an interview costs you the question

  • Thinks get_prompt already returns a ready-to-send string
  • Uses single braces for Langfuse prompt variables
  • Assumes compile() raises when a variable is missing
  • Expects a chat prompt to compile down to one string
  • Believes compile() makes a network call to Langfuse

context

open as a page

In Langfuse, how do prompt versions and labels decide what get_prompt() returns?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Every save of a Langfuse prompt creates a new immutable numbered version. Labels such as production and latest are movable pointers to one of those versions. langfuse.get_prompt("name") resolves the version labelled production; pass label= or version= to target a different one.

open as a page

In Langfuse, how do you attach a score to a trace or an observation?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A Langfuse score is a named value hung off a trace or off one observation inside it. From inside an active observation call score_current_span() or score_current_trace(); from anywhere else call score_trace(trace_id=...) with a trace id you stored earlier.

open as a page

How does Langfuse's @observe decorator trace a Python function?

level: juniorimportance: must knowfreq 70%

basics

~20 s

@observe wraps a Python function so each call becomes one Langfuse observation, named after the function, timed, with the arguments captured as input and the return value as output, nested automatically under whatever observation is already active.

open as a page

In Langfuse, what is a trace and what are the observations inside it?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A Langfuse trace is one end-to-end unit of work, typically one request or agent run. Observations are the timed steps inside it: nested and typed, with GENERATION recording an LLM call and SPAN any other step.

open as a page

How does the Langfuse SDK cache prompts fetched with get_prompt, and what is the default TTL?

level: middleimportance: must knowfreq 66%

basics

~20 s

The Langfuse SDK keeps a client-side, in-process cache of fetched prompts with a default TTL of 60 seconds. After expiry the next get_prompt call serves the cached copy and refreshes in the background, so only the very first fetch in a process blocks on the network.

open as a page

In Langfuse, how do you build a dataset and run your app over it as a dataset run?

level: middleimportance: must knowfreq 58%

basics

~20 s

Create the dataset with create_dataset(name=...) and fill it using create_dataset_item(dataset_name=..., input=..., expected_output=...), optionally pointing source_trace_id at the production trace an item came from. Then fetch it with get_dataset(name) and execute a run — dataset.run_experiment(name=..., task=..., evaluators=[...]) — which traces every item and attaches its scores.

open as a page

How do you set up a Langfuse LLM-as-judge evaluator to score production traces?

level: middleimportance: must knowfreq 55%

basics

~20 s

You configure an evaluator in the Langfuse project: pick or write an evaluator template whose prompt has variables, map those variables onto fields of the trace, choose a judge model, then target it with a filter over traces plus a sampling rate. Matching traces get a score whose source is EVAL.

open as a page

How do you keep prompt and completion text out of a Langfuse trace store?

level: middleimportance: must knowfreq 55%

basics

~20 s

Pass a mask function to the Langfuse client. It runs inside your process on every observation's input and output before anything is sent, so redacted content never leaves the application. Combine it with not passing sensitive fields at all, plus self-hosting and short retention for what remains.

open as a page

What infrastructure does a self-hosted Langfuse v4 deployment need, and what does each store?

level: middleimportance: must knowfreq 62%

basics

~20 s

Langfuse v4 runs two application containers, web and worker, over four stores: Postgres for configuration, users and prompts; ClickHouse for traces, observations and scores; Redis or Valkey for the ingestion queue and cache; and S3-compatible blob storage for raw events and media.

open as a page

In Langfuse, when do you record a GENERATION instead of a plain SPAN?

level: middleimportance: must knowfreq 60%

basics

~20 s

Use GENERATION for a call to a model. It carries model name, model parameters, prompt and completion, token usage and cost, which is what feeds Langfuse's token and spend analytics. SPAN is a generic timed step with none of those fields.

open as a page

In Langfuse v4, how do you attach session and user IDs to a trace?

level: middleimportance: must knowfreq 57%

basics

~20 s

Wrap the work in Langfuse's propagate_attributes context manager, passing session_id, user_id and tags. Enter it as early as possible: it is not retroactive, so only the active span and spans created inside it carry the attributes.

open as a page

What happens to langfuse.get_prompt() when the Langfuse API is unreachable?

level: seniorimportance: must knowfreq 54%

basics

~20 s

A process that already holds a cached copy keeps serving it, even past the TTL, so a warm process rides out the outage. A process with a cold cache fails unless you passed fallback=, which returns a prompt client built from the literal you supplied.

open as a page

What does Langfuse's drop-in OpenAI wrapper capture automatically?

level: juniorimportance: should knowfreq 56%

basics

~20 s

Importing the OpenAI client from langfuse.openai instead of the openai package makes every call emit a GENERATION observation on its own — model, parameters, messages, completion, token usage, latency and inferred cost — with no other code change.

open as a page

What do Langfuse's NUMERIC, CATEGORICAL and BOOLEAN score data types change?

level: middleimportance: should knowfreq 48%

basics

~20 s

The data_type decides what a Langfuse score value may be and how it aggregates: NUMERIC takes a number and averages, CATEGORICAL takes a string label and counts per label, BOOLEAN takes a pass/fail recorded numerically so it charts as a rate. A score config pins the type and allowed values for a score name.

open as a page

Why is Langfuse's docker compose stack not a production self-hosted deployment?

level: middleimportance: should knowfreq 44%

basics

~20 s

The compose file runs web and worker next to single-node Postgres, ClickHouse, Redis and MinIO on one host with local volumes and example secrets. It is a working evaluation stack with no replication, no backups, no TLS and a single point of failure.

open as a page

How do you create nested Langfuse observations without the @observe decorator?

level: middleimportance: should knowfreq 46%

basics

~20 s

Get the client with get_client(), then use start_as_current_observation(as_type=...) as a context manager: it opens an observation, makes it current so nested work attaches beneath it, and ends it on exit. start_observation() does not make it current, so you call end() yourself.

open as a page

What is a Langfuse annotation queue, and how does human review become scores?

level: seniorimportance: should knowfreq 40%

basics

~20 s

An annotation queue is a work list of traces or observations for humans to review in the Langfuse UI. Reviewers apply the project's score configs to each item, and their verdicts are stored as ordinary scores on the same trace, marked with source ANNOTATION.

open as a page

In self-hosted Langfuse, how do you limit how long trace data is retained?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Set a data retention window in days per project. The worker runs the deletion job, removing aged traces, observations and scores from ClickHouse along with the associated objects in blob storage. Reduce inflow separately with SDK sampling, since retention only trims what you already stored.

open as a page

In self-hosted Langfuse v4, what does the worker container do that the web container does not?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The web container accepts ingestion batches, writes the raw payloads to blob storage and enqueues them, then serves the UI and API. The worker consumes that queue asynchronously: it writes traces, observations and scores into ClickHouse and runs background jobs such as evaluators and retention cleanup.

open as a page

Langfuse traces never appear from your serverless function. How do you debug it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The SDK batches observations and exports them from a background thread, so a handler that returns or freezes before the export loses them. Call flush() before returning, then check credentials, host and project before suspecting the platform.

open as a page

How would you govern Langfuse prompt changes that ship without a deploy?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat moving the production label as a release with none of the usual safeguards. Restrict who can move it, require the candidate version to pass an offline run first, record the version on every generation so regressions are attributable, and pin versions on surfaces where silent change is unacceptable.

open as a page

How would you govern score names and configs across teams in one Langfuse project?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat the score name as a schema, not a label: it is the key every chart, filter and run comparison groups by. Register each name once as a score config with a fixed type and range, keep human, judge and code-written signals distinguishable, and make a changed meaning mean a new name.

open as a page

When would you self-host Langfuse rather than use the managed cloud, and what do you take on?

level: principalimportance: should knowfreq 38%

basics

~20 s

Self-host when prompt and trace content cannot leave your network for regulatory or contractual reasons, or when volume makes managed pricing worse than running it. In exchange you own a replicated OLAP store, an object store, backups, upgrades, schema migrations and on-call for an internal tool.

open as a page