Why is Langfuse's docker compose stack not a production self-hosted deployment?
answer
- one host, local volumes, no replication
- example secrets ship in the file
- MinIO stands in for object storage
- stateless containers, managed stateful services
basics
~20 sThe compose file runs web and worker next to single-node Postgres, ClickHouse, Redis and MinIO on one host with local volumes and example secrets. It is a working evaluation stack with no replication, no backups, no TLS and a single point of failure.
solid answer
~50 sLangfuse ships a compose file that brings the whole stack up in one command, which is excellent for evaluating the product and misleading as a production model. It runs the web and worker containers alongside single-node Postgres, single-node ClickHouse, Redis and MinIO as the S3-compatible store, all on one host with local volumes. There is no replication, no backup, no TLS termination, and the security values ship as placeholders that you must replace with generated ones: `NEXTAUTH_SECRET`, `SALT` and a 32-byte hex `ENCRYPTION_KEY`. One host failure loses everything, and upgrades mean downtime. A production deployment keeps the two stateless containers, on Kubernetes via the official Helm chart or on VMs, and points them at managed Postgres, a managed or replicated ClickHouse, managed Redis and real object storage, with TLS in front, generated secrets, backups per store, and every component running in UTC.
go deeper
Know that the compose file exists and is the fast way to try Langfuse locally, and that it starts the databases for you rather than expecting you to provide them.
Be able to list concretely what production adds: generated secrets, TLS, backups, replication for the stateful services, and more than one replica of the worker.
Show the migration shape and the ordering: keep the stateless containers, move Postgres and ClickHouse to managed services first, and carry over UTC, retention and queue-depth monitoring.
Own the total cost of the decision. Weigh the team's ability to operate a replicated OLAP store against the residency requirement that motivated self-hosting, and decide where the managed boundary sits.
## What the compose stack actually is The Langfuse repository includes a compose file that starts the application and every dependency: the **web** container, the **worker** container, **Postgres**, **ClickHouse**, **Redis** and **MinIO** standing in for S3-compatible object storage. It is genuinely useful. Someone can evaluate nested traces, prompt management and evaluators in minutes without provisioning anything, and it is the right first step before a self-hosting decision. The trap is treating it as a deployment template. It is a demonstration of the topology, not an implementation of it. ## What is missing **Single point of failure.** Every component is one process on one host with local volumes. Lose the host and you lose the application, the configuration database, the trace history and the object store at once. Nothing in the file replicates anything. **No backups.** Local volumes are not a backup strategy. Production needs a real backup for Postgres, which holds users, projects, API keys, prompt versions and dataset definitions, and a recovery plan for ClickHouse appropriate to the volume you are willing to lose. **Placeholder secrets.** `NEXTAUTH_SECRET`, `SALT` and `ENCRYPTION_KEY` come as example values so the stack starts. Shipping those to production means anyone with the public repository knows the values protecting your session layer, your hashed API keys, and the encryption of stored provider credentials. Generate real ones, for example `openssl rand -hex 32` for `ENCRYPTION_KEY`, and store them where you keep other secrets. Treat `ENCRYPTION_KEY` as unrecoverable: lose it and the secrets it protects become unreadable. **No TLS or exposure model.** The compose stack listens in plaintext and assumes a trusted local network. Production needs TLS termination, a real hostname in `NEXTAUTH_URL`, and a decision about who can reach the UI and the ingestion endpoint. **No capacity story.** One replica of each container, default resources, no scaling. A real ingestion load needs worker replicas sized to event volume and ClickHouse sized for the retention window. **Upgrades are downtime.** Bringing the stack down to pull new images is fine on a laptop. In production you want the stateless containers to roll while the stateful services stay up, which is a different deployment shape entirely. ## What production looks like instead The useful mental model: **keep the stateless part, outsource the stateful part.** The web and worker containers are stateless and horizontally scalable, so they belong wherever your team already runs services, commonly Kubernetes using the official Langfuse Helm chart, or plain containers on VMs behind a load balancer. Behind them, most teams point at managed services: managed Postgres, a managed or properly replicated ClickHouse, managed Redis or Valkey, and real S3 or an equivalent object store. That leaves you operating an application rather than a four-way stateful platform, while keeping every byte inside accounts you control, which was the point of self-hosting in the first place. Also carry over the operational details the compose file hides: run everything in UTC to avoid time ranges that disagree with the UI, monitor queue depth so a stalled worker is visible, set project retention deliberately, and back up Postgres first. ## How to answer it Do not just say "compose is not production". Name what is missing, in order of what would actually hurt: single host and no replication, no backups, example secrets, no TLS, no scaling, downtime upgrades. Then give the migration shape, stateless containers plus managed stateful services. That is the answer of someone who has run it rather than read about it.
- Which single component would you move off the compose host first, and why?Postgres. It holds the identity of everything else: organisations, users, projects, hashed API keys, prompt versions and dataset definitions, and losing it means the trace history in ClickHouse refers to entities that no longer exist. It is also the smallest and cheapest to hand to a managed service, so it gives the largest reduction in blast radius per unit of effort. ClickHouse usually follows, once trace volume makes single-node storage a real risk.
- What breaks if you change ENCRYPTION_KEY on an existing deployment?Anything encrypted with the old key becomes unreadable, which in practice means stored secrets such as the LLM provider credentials configured for server-side evaluators, and integration settings. It is not a rotation you can perform casually; you would need to re-enter the affected secrets afterwards. Generate it once at install time and store it with the same care as a database credential.
saying these in an interview costs you the question
- Running the shipped example secrets in production
- Treating local docker volumes as a backup
- Assuming MinIO on one host equals durable object storage
- Expecting to scale ingestion without adding worker replicas