In self-hosted Langfuse, how do you limit how long trace data is retained?
answer
- per-project window, in days
- the worker performs the deletion
- deletes ClickHouse rows and stored objects too
- sampling cuts inflow, retention cuts age
basics
~20 sSet a data retention window in days per project. The worker runs the deletion job, removing aged traces, observations and scores from ClickHouse along with the associated objects in blob storage. Reduce inflow separately with SDK sampling, since retention only trims what you already stored.
solid answer
~50 sRetention in Langfuse is a per-project setting expressed in days, configured in project settings and enforced asynchronously by the worker container, which deletes aged traces, observations and scores from ClickHouse and the corresponding objects from blob storage. Two things follow. First, deletion is a background job, not an instant guarantee, so "data older than N days is gone" is true within the job's cadence rather than to the second. Second, retention is a downstream control: it bounds how long you hold data, not how much arrives. To control inflow you sample at the SDK, using the client's `sample_rate` or the `LANGFUSE_SAMPLE_RATE` environment variable, which drops whole traces so the ones you keep stay complete and readable. In a regulated setting a short retention window is what makes an imperfect masking rule survivable, because it bounds the exposure of anything that slipped through.
go deeper
Know that trace data does not live forever by default configuration and that retention is something a project sets deliberately, in days.
Be able to separate retention from sampling: one bounds how long data survives on the server, the other bounds how much is ever recorded by the client.
Show the operational detail: deletion is an asynchronous worker job across ClickHouse and blob storage, disk does not shrink instantly, and retention is what makes an imperfect mask survivable.
Own the policy. Derive the window from regulatory commitments and the real debugging horizon, and design the promote-to-dataset path so shortening retention does not destroy evaluation history.
## Two different levers Teams conflate them constantly, so separate them explicitly: - **Retention** decides how long recorded data survives. It is a property of the server, set per project. - **Sampling** decides how much data is recorded at all. It is a property of the client SDK. Only the second reduces ingestion load, ClickHouse write volume and blob storage growth in real time. Only the first bounds how long a leak sits in your store. You usually need both, and for different reasons: sampling for cost and throughput, retention for compliance. ## How retention behaves You configure a retention window in days on the project. The worker container runs the deletion job, and it deletes across the stores that hold the data: rows in the ClickHouse trace, observation and score tables, and the corresponding raw event and media objects in blob storage. Deleting from only one of those would leave the payloads behind, which is why this is a server-side job rather than something you can approximate with a bucket lifecycle rule alone. Because it is asynchronous and batched, express the guarantee honestly: data is removed within the job's cadence after it ages out, not at the instant the clock passes N days. If someone needs a hard contractual deletion deadline, that gap is worth naming rather than glossing over. A practical note for self-hosters: retention deletes reclaim logical space, but reclaiming disk in an OLAP store and in object storage has its own mechanics and lag. Size ClickHouse for steady-state volume at your retention window, plus headroom, rather than assuming the delete job keeps disk flat automatically. ## Sampling, and why trace-level sampling matters The SDK's sample rate is a fraction between 0 and 1, set as `sample_rate` on the client or as `LANGFUSE_SAMPLE_RATE` in the environment. It decides whether a trace is recorded, and the decision applies to the whole trace, not per observation. That is the important behaviour: a half-recorded trace, with the retriever span but not the generation it fed, is worse than no trace, because it misleads you during a debugging session. Sampling whole traces keeps every kept trace complete. The cost of sampling is specific and worth stating: the incident you are called about may be in the 90% you dropped. Teams often run a high sample rate on low-volume, high-stakes flows and a low rate on chatty background traffic, rather than one global number. ## Choosing the window The inputs are not primarily technical: - **Regulatory and contractual obligations.** If a data-protection commitment says customer content is not retained beyond a period, that number is the ceiling, and it is not yours to negotiate. - **Debugging horizon.** How far back do people actually look? For most LLM applications, a bug report arrives within days. Long tails matter mainly for slow-burn quality regressions. - **Evaluation needs.** If you promote production traces into datasets, you need them alive long enough to curate. The usual pattern is a short trace retention window plus datasets, which are configuration-tier data in Postgres and are not swept by trace retention, so a curated example outlives the trace it came from. - **Storage cost.** Real but usually secondary once ClickHouse compression is accounted for. ## Saying it well in an interview The strong answer connects retention to the rest of the data-control story: mask at the client so sensitive content is never sent, self-host so what is sent stays in your network, retain briefly so anything that slipped through has a short life, and sample to keep the volume, and therefore the exposure, proportionate to what you actually debug. Each layer covers a different failure of the others.
- If you shorten retention, do your curated evaluation datasets disappear too?No. Datasets and their items are configuration-tier data in Postgres, created when you promote an example, so they are not swept by the trace retention job. That is exactly why the promote-then-expire pattern works: keep raw traces briefly for debugging, and lift the handful worth keeping into a dataset that outlives them. Verify it in your own deployment before relying on it for a compliance claim.
- Why sample whole traces rather than individual observations?Because partial traces mislead. If the retriever span is kept and the generation it fed is dropped, the surviving trace suggests a retrieval-only call and hides the causal chain you were trying to follow. Trace-level sampling means every kept trace is internally complete, so a smaller sample is still trustworthy, at the cost of missing whole requests rather than fragments of them.
saying these in an interview costs you the question
- Thinking retention reduces ingestion cost or write load
- Expecting deletion to be instant rather than a background job
- Assuming a bucket lifecycle rule alone satisfies retention
- Setting a sample rate and calling it a compliance control