skip to content

Self-Hosting & Data Controls

What it actually takes to run Langfuse in your own infrastructure, and how to keep prompt content out of the trace store when data cannot leave. This is the deciding question for regulated teams.

on this pageshow

questions

6

How do you keep prompt and completion text out of a Langfuse trace store?

level: middleimportance: must knowfreq 55%

answer

  1. redact before the network call
  2. a callback on the client, not a server rule
  3. recurse over dicts and lists
  4. a raising mask fails closed

basics

~20 s

Pass a mask function to the Langfuse client. It runs inside your process on every observation's input and output before anything is sent, so redacted content never leaves the application. Combine it with not passing sensitive fields at all, plus self-hosting and short retention for what remains.

solid answer

~50 s

Langfuse records inputs and outputs by default, which for an LLM application means full prompts and completions. The primary control is the SDK's `mask` parameter: you pass a callable to the `Langfuse` client and the SDK applies it to the input and output of every observation before serialisation, so redaction happens client-side and the sensitive text is never transmitted. That is the important property. A server-side rule would still require you to ship the data first. Write the function defensively, handling strings, dicts and lists, because inputs are arbitrary structures; if it raises, the SDK falls back to a fully masked placeholder rather than sending raw content. Masking is one layer, not the answer on its own: also avoid putting secrets into traced arguments, self-host so traces stay in your network, set a short retention window, and restrict who can read the project.

code

python · 15 lines
python
import re
from langfuse import Langfuse

CARD = re.compile(r"\b\d{13,16}\b")

def mask(data):
    if isinstance(data, str):
        return CARD.sub("[REDACTED]", data)
    if isinstance(data, dict):
        return {k: mask(v) for k, v in data.items()}
    if isinstance(data, list):
        return [mask(v) for v in data]
    return data

langfuse = Langfuse(mask=mask)

go deeper

for a junior

Know that Langfuse records prompts and completions by default, and that this is a deliberate decision to review rather than something that just happens.

for a middle

Be able to describe the mask callback on the client, that it runs in your process before data is sent, and that it must handle nested dicts and lists rather than plain strings only.

for a senior

Show the layered answer: masking plus not passing sensitive data, plus self-hosting, retention limits and project-level access separation, and name the tradeoff against debuggability.

for a principal

Own the position you would defend to a compliance reviewer: which controls you rely on, why masking alone is a single point of failure, and what evidence you keep that it works.

## The default is the problem Every tool in this category records the model call in full, because a trace without the prompt and the completion is nearly useless for debugging. For a regulated team that default is exactly the thing they cannot accept: the trace store becomes a second copy of customer data, health records, card numbers or personal identifiers, sitting in a system that was procured as a debugging tool and is often readable by the whole engineering team. ## Client-side masking Langfuse's mechanism is a mask callback registered on the client. You give the `Langfuse` constructor a `mask` function, and the SDK applies it to the input and output of every observation it produces, before the event is serialised and sent. The crucial property is where it runs: in your process. Redaction happens before the network call, so the raw value never leaves the application boundary. This is a different guarantee from a server-side redaction rule, which requires you to transmit and store the data before it is scrubbed. When someone asks whether you can use a hosted trace tool under a data-residency constraint, this distinction is the whole answer. ## Writing a mask function that survives production Three practical rules: **Recurse over structures.** Inputs and outputs are arbitrary Python objects: strings, dicts of message objects, lists of tool arguments. A function that only handles `str` silently passes structured payloads through untouched, which is the most common way masking gives false comfort. **Fail safe, and know what failure does.** If your function raises, the SDK does not fall back to sending the raw value; it substitutes a fully masked placeholder. That is the right default, but it also means a buggy mask can wipe out the observability you deployed the tool for. Test the function directly, on real payload shapes. **Prefer allow-listing over pattern hunting where you can.** Regexes for card numbers and emails catch the obvious cases and miss free text. Where the schema is known, redact by field name and keep only what you deliberately decided is safe. ## Masking is one layer of several A complete answer names the layers around it, because masking alone is a single point of failure: - **Do not pass it in the first place.** The strongest control is structuring calls so the sensitive value never reaches an instrumented argument, for example passing a customer reference rather than the record. - **Self-host.** With Langfuse running on your own infrastructure, whatever does get recorded stays inside your network and under your access controls. This is the reason self-hosting and data controls belong in the same conversation. - **Retention.** Shorten how long recorded content is kept, so an imperfect mask has a bounded blast radius. - **Access control and separation.** Restrict project membership, and consider a separate project for a workload with stricter obligations so its data is not co-mingled with general debugging traffic. - **Sampling.** Recording a fraction of traffic reduces exposure as well as cost, at the price of missing the specific trace you wanted. ## The conversation to expect An interviewer will typically push on whether masking makes a hosted deployment acceptable. The defensible position is that client-side masking genuinely prevents transmission of what it catches, but you are betting compliance on the correctness of one function over untyped payloads. For a workload where being wrong is a reportable event, self-hosting plus masking is the combination people actually approve, because the two failure modes do not overlap: masking protects content, self-hosting bounds where an unmasked leak can land.

  • If masking runs client-side anyway, why does self-hosting still matter?
    Because masking is code you wrote over untyped payloads, and it will eventually miss something. Self-hosting bounds the damage when it does: the unmasked value lands in a store inside your network, under your access controls and your retention policy, rather than in a third party's system that you must now notify and negotiate deletion with. The two controls fail independently, which is why regulated teams use both.
  • What are the practical downsides of aggressive masking?
    You lose the debugging value the tool was bought for. A trace whose prompt reads as a row of redaction markers tells you the call happened and what it cost, but not why the model answered badly, and it makes LLM-as-judge evaluation on production traces impossible because the judge sees nothing to score. Teams usually converge on field-level redaction of known-sensitive attributes, keeping surrounding structure intact.
  • Where should the mask function be defined in a large codebase?
    Once, next to client initialisation, and applied to the single shared client so no code path can construct an unmasked one. Scattering per-call redaction guarantees that the newest integration forgets it. Treat the function as security-relevant code: unit test it against real payload shapes, review changes to it, and fail the build if it stops handling nested structures.

saying these in an interview costs you the question

  • Believing redaction happens on the server after upload
  • Masking only top-level strings and ignoring nested dicts
  • Assuming a failing mask function sends the raw value
  • Treating masking as sufficient compliance on its own

context

open as a page

What infrastructure does a self-hosted Langfuse v4 deployment need, and what does each store?

level: middleimportance: must knowfreq 62%

basics

~20 s

Langfuse v4 runs two application containers, web and worker, over four stores: Postgres for configuration, users and prompts; ClickHouse for traces, observations and scores; Redis or Valkey for the ingestion queue and cache; and S3-compatible blob storage for raw events and media.

open as a page

Why is Langfuse's docker compose stack not a production self-hosted deployment?

level: middleimportance: should knowfreq 44%

basics

~20 s

The compose file runs web and worker next to single-node Postgres, ClickHouse, Redis and MinIO on one host with local volumes and example secrets. It is a working evaluation stack with no replication, no backups, no TLS and a single point of failure.

open as a page

In self-hosted Langfuse, how do you limit how long trace data is retained?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Set a data retention window in days per project. The worker runs the deletion job, removing aged traces, observations and scores from ClickHouse along with the associated objects in blob storage. Reduce inflow separately with SDK sampling, since retention only trims what you already stored.

open as a page

In self-hosted Langfuse v4, what does the worker container do that the web container does not?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The web container accepts ingestion batches, writes the raw payloads to blob storage and enqueues them, then serves the UI and API. The worker consumes that queue asynchronously: it writes traces, observations and scores into ClickHouse and runs background jobs such as evaluators and retention cleanup.

open as a page

When would you self-host Langfuse rather than use the managed cloud, and what do you take on?

level: principalimportance: should knowfreq 38%

basics

~20 s

Self-host when prompt and trace content cannot leave your network for regulatory or contractual reasons, or when volume makes managed pricing worse than running it. In exchange you own a replicated OLAP store, an object store, backups, upgrades, schema migrations and on-call for an internal tool.

open as a page