What infrastructure does a self-hosted Langfuse v4 deployment need, and what does each store?
answer
- two containers, four stores
- one transactional store, one analytical
- the queue sits between web and worker
- raw event bodies land in a bucket first
basics
~20 sLangfuse v4 runs two application containers, web and worker, over four stores: Postgres for configuration, users and prompts; ClickHouse for traces, observations and scores; Redis or Valkey for the ingestion queue and cache; and S3-compatible blob storage for raw events and media.
solid answer
~40 sA self-hosted Langfuse v4 deployment is two stateless application containers plus four stateful dependencies. The **web** container serves the UI, the public API and the ingestion endpoint; the **worker** drains queued events asynchronously. Behind them, **Postgres** is the transactional store for organisations, users, projects, API keys, prompts and dataset definitions; **ClickHouse** is the analytical store holding traces, observations and scores, the high-volume tables every dashboard and filter queries; **Redis or Valkey** carries the ingestion queue and caches; an **S3-compatible bucket** holds raw ingestion event bodies and multi-modal media attachments. Wiring is environment-driven: `DATABASE_URL`, `CLICKHOUSE_URL`, `CLICKHOUSE_MIGRATION_URL`, `REDIS_CONNECTION_STRING`, `LANGFUSE_S3_EVENT_UPLOAD_BUCKET`, plus the security values `NEXTAUTH_SECRET`, `SALT` and `ENCRYPTION_KEY`. Every component should run in UTC. The v2 generation needed only Postgres, which is why older self-hosting write-ups look far simpler than what v4 asks of you.
code
bash · 11 linesexport DATABASE_URL=postgresql://langfuse:secret@postgres:5432/langfuse
export CLICKHOUSE_URL=http://clickhouse:8123
export CLICKHOUSE_MIGRATION_URL=clickhouse://clickhouse:9000
export CLICKHOUSE_USER=default
export CLICKHOUSE_PASSWORD=secret
export REDIS_CONNECTION_STRING=redis://redis:6379
export LANGFUSE_S3_EVENT_UPLOAD_BUCKET=langfuse-events
export NEXTAUTH_URL=https://langfuse.internal.example.com
export NEXTAUTH_SECRET=$(openssl rand -base64 32)
export SALT=$(openssl rand -base64 32)
export ENCRYPTION_KEY=$(openssl rand -hex 32)go deeper
Know that Langfuse is open source and can run on your own infrastructure, and be able to name that it needs a database rather than just a container.
Be able to list the two containers and the four stores and say what each one holds, especially why traces go to ClickHouse while projects and prompts go to Postgres.
Show that you have thought about operating this: backups per store, sizing ClickHouse for trace volume, running everything in UTC, and treating ENCRYPTION_KEY as unrecoverable.
Own the build-versus-buy line inside self-hosting: which of the four stores your team runs and which you take as managed services, and what that choice costs in on-call load.
## Why this is the first question Langfuse's differentiator against hosted-only observability products is that you can run the whole thing inside your own network. So "can we self-host it?" is really a question about the operational bill, and the honest answer is that Langfuse v4 is not one container with a database next to it. It is a small distributed system, and a candidate who says "just docker compose up" has answered the evaluation question, not the production one. ## The two application containers Both ship from the same image family and both are stateless, so you scale them horizontally. - **web** serves the Next.js UI, the public REST API, and the ingestion endpoint that every SDK posts batches to. It reads from ClickHouse and Postgres to render traces and dashboards. - **worker** is the asynchronous half. It consumes the queue and does the writes and background jobs the request path deliberately does not do. Splitting them means a burst of ingestion traffic does not make the UI unusable, and you can size the two independently. ## The four stateful dependencies **Postgres** is the OLTP store: organisations, users and memberships, projects, hashed API keys, prompt versions and labels, dataset definitions, configuration. Low volume, high value, transactional. This is the database whose backup you would actually be asked about after an incident. **ClickHouse** is the OLAP store, added in the v3 generation and still central in v4. Traces, observations and scores live here. This is where the volume is: a chatty agent application writes far more observations than anything else in the system, and the product's core interactions are analytical scans, filters and aggregations over long time ranges. A row-oriented store would not hold up. **Redis or Valkey** provides the ingestion queue between web and worker plus application caching. It is not optional in v4; without it the asynchronous pipeline has nowhere to hand work off. **S3-compatible blob storage** takes raw event bodies at ingestion time and multi-modal media attachments such as images and audio. Writing the raw payload to object storage first is what lets the API accept a batch quickly and durably before any database work happens. ## Configuration Everything is environment variables. The connection set is `DATABASE_URL`, `CLICKHOUSE_URL` and `CLICKHOUSE_MIGRATION_URL` (the HTTP interface and the native-protocol URL that migrations use), `CLICKHOUSE_USER` and `CLICKHOUSE_PASSWORD`, `REDIS_CONNECTION_STRING`, and the bucket settings led by `LANGFUSE_S3_EVENT_UPLOAD_BUCKET`, with a separate `LANGFUSE_S3_MEDIA_UPLOAD_BUCKET` for attachments. On the security side, `NEXTAUTH_URL` and `NEXTAUTH_SECRET` back the session layer, `SALT` is used when hashing API keys, and `ENCRYPTION_KEY` (32 bytes of hex) encrypts stored secrets such as the LLM provider credentials you configure for evaluators. Generate those three properly and treat them as unrecoverable: rotate `ENCRYPTION_KEY` and the values it protects become unreadable. One easily-missed operational rule: run the infrastructure in UTC. Mixed timezones between the application and ClickHouse produce time ranges that quietly disagree with what the UI shows. ## What this means for you Self-hosting Langfuse means you now operate a replicated OLAP database and an object store, not just "an app". Most teams reduce the burden by keeping the two containers themselves in Kubernetes or on a VM and pointing them at managed Postgres, a managed or hosted ClickHouse, a managed Redis, and real S3, rather than running all four themselves.
- Which of those stores would you back up, and how would you prioritise them?Postgres first: it holds users, projects, API keys, prompt versions and dataset definitions, and losing it means losing the identity of everything else. ClickHouse next, because it holds the observability history, though most teams accept a coarser recovery point there given the volume. The blob bucket needs a lifecycle policy and versioning rather than classic backups, and Redis is a queue and cache, so it is rebuildable, at the cost of whatever was still in flight.
- Do you have to run ClickHouse yourself, or can you point Langfuse at a managed one?You can point it at a managed ClickHouse. `CLICKHOUSE_URL` and `CLICKHOUSE_MIGRATION_URL` accept any reachable endpoint with credentials, so a hosted ClickHouse plus managed Postgres, managed Redis and real S3 is a common middle ground: you keep the data in accounts you control while handing the hardest stateful operations to a provider. The Langfuse containers stay the only thing your team actually runs.
saying these in an interview costs you the question
- Saying Langfuse self-hosts as a single container with Postgres
- Thinking ClickHouse is optional or just a cache
- Assuming traces are stored in Postgres like configuration is
- Skipping ENCRYPTION_KEY and SALT as optional extras