When would you self-host Langfuse rather than use the managed cloud, and what do you take on?
answer
- the driver is data, not features
- cost crossover only bites at volume
- you are adopting a stateful platform
- managed stores, self-hosted app
basics
~20 sSelf-host when prompt and trace content cannot leave your network for regulatory or contractual reasons, or when volume makes managed pricing worse than running it. In exchange you own a replicated OLAP store, an object store, backups, upgrades, schema migrations and on-call for an internal tool.
solid answer
~50 sThe decision is rarely about features, because the open-source core is the same product. It turns on data and on operating capacity. Self-host when trace content is subject to residency, contractual or regulatory constraints, when network policy forbids egress to a third party, or when trace volume is high enough that running the stack is cheaper than paying per event. What you take on is real: Postgres, ClickHouse, Redis and object storage, plus backups, restore drills, version upgrades that carry database migrations, TLS and SSO wiring, capacity for ingestion bursts, and an on-call rotation for a tool whose outage blocks debugging rather than serving customers. The middle grounds usually win. Run the stateless containers yourself and buy the stateful services; or stay managed and rely on client-side masking plus short retention. Decide by asking what happens when unmasked customer content lands in the store, and where you need it to land when it does.
go deeper
Know that Langfuse can run either as a managed service or on your own infrastructure, and that the usual reason to choose self-hosting is keeping data in-house.
Be able to name what self-hosting actually adds operationally: four stateful dependencies, backups, upgrades with migrations, and TLS plus access configuration.
Argue the tradeoff with real signals: the specific obligation, measured trace volume, who operates the stateful services, and what an outage of the trace store costs during an incident.
Own the decision and its review condition. Name the middle option of self-hosted app with managed stores, state what would flip the choice, and be willing to refuse self-hosting when no team will operate it.
## Why this is a leadership question Nothing here is technically hard to state. What makes it a principal-level conversation is that the inputs are organisational: a regulatory obligation you did not write, a procurement process, a platform team's appetite for another stateful system, and a cost curve that only bites at volume. The failure mode is a team self-hosting on instinct, discovering six months later that nobody owns ClickHouse upgrades, and quietly running an unpatched, unbacked-up copy of every customer prompt. ## The arguments for self-hosting **Data residency and regulation.** The strongest and most common driver. If prompts contain health, financial or personal data and your obligations say it stays in your infrastructure or your region, the conversation ends there. This is exactly why the open-source, self-hostable option exists. **Network policy.** Some environments simply do not allow outbound traffic to a third-party endpoint from the services that would emit traces. Self-hosting turns an egress exception into an internal service call. **Cost at volume.** Per-event pricing is attractive at low volume and less so when an agent application emits many observations per user request. At some crossover point, running the stack yourself is cheaper, though be honest that engineer time on-call is part of the comparison and is usually the term people omit. **Control.** Your own retention policy, your own SSO and access model, your own upgrade timing, no dependency on a vendor's availability for your debugging workflow. ## The arguments against **You are adopting a stateful platform.** Not one container: a transactional database, an analytical database, a queue and an object store. ClickHouse in particular is not something most application teams operate today, and its failure modes are unfamiliar. **Upgrades carry migrations.** Version upgrades touch database schemas across two engines. That is a maintenance stream someone must own, and skipping it for a year turns a routine upgrade into a project. **Backups only count if restores are rehearsed.** Losing Postgres means losing users, projects, API keys, prompt versions and dataset definitions, which is far worse than losing trace history. **On-call for an internal tool.** When the trace store is down during an incident, you have lost the instrument you were using to diagnose the incident. Teams under-invest here precisely because it is not customer-facing. **Feature and licence nuance.** The core is open source and self-hostable, while some enterprise-oriented capabilities are covered by a commercial licence. Check what your requirements actually depend on before assuming self-hosting is free of vendor conversation. ## The options in between The realistic space is not binary: - **Self-host the application, buy the stores.** Run the stateless web and worker containers on your existing platform; use managed Postgres, hosted or managed ClickHouse, managed Redis and S3. Data stays in accounts you control, and you operate an app rather than a database fleet. This is where most regulated teams land. - **Managed plus client-side controls.** Stay on the managed service and rely on masking so sensitive content never leaves the process, plus short retention and tight access. Viable when the constraint is about content rather than about jurisdiction. - **Split by workload.** Regulated products self-hosted, everything else managed. It costs you a second deployment and a fragmented view, so only do it when the obligation genuinely differs. ## How to reason out loud Start from the requirement, not the preference: what specifically may not leave, and who says so. Then ask what happens the day an unmasked payload is recorded, because that is the scenario the decision is actually protecting against. Then be candid about capacity: who operates ClickHouse, who is paged, who runs the upgrade. A principal answer names the middle option and the conditions under which you would revisit the choice, such as trace volume crossing the point where managed pricing dominates, or a new obligation removing the option entirely.
- A team says they must self-host because prompts contain personal data. What would you probe before agreeing?Whether the obligation is about jurisdiction or about content. If it is jurisdictional, self-hosting or a region-bound deployment is the answer. If it is about content, client-side masking may satisfy it, because redaction happens before transmission, and then the question becomes how confident you are in that function. I would also ask what retention they need, since a short window plus masking is often the proportionate control rather than adopting a four-store platform.
- How would you decide when trace volume justifies moving from managed to self-hosted on cost?Model both sides honestly. Managed cost scales with events, so project it from measured observations per request, not per user request. Self-hosted cost is infrastructure plus the engineer time to operate four stateful services, run upgrades and carry on-call, and that second term is the one teams omit. I would set the trigger as a sustained volume level rather than a spike, and revisit it at planned intervals so the decision is not made during a bill shock.
- What would make you refuse to self-host even under a residency requirement?No owner. If no team will hold ClickHouse and Postgres operations, backups and upgrades on their roadmap, self-hosting produces something worse than the managed service: an unpatched, unmonitored copy of every prompt with no restore path. In that case I would push for a region-bound managed deployment plus aggressive masking and short retention, and escalate the residency requirement as a resourcing decision rather than quietly under-delivering it.
saying these in an interview costs you the question
- Self-hosting on principle without naming the actual requirement
- Comparing costs while ignoring on-call and upgrade effort
- Assuming self-hosting alone satisfies data-protection obligations
- Treating it as binary and missing managed-stores middle ground