skip to content

In self-hosted Langfuse v4, what does the worker container do that the web container does not?

level: seniorimportance: should knowfreq 45%

answer

  1. fast accept path, slow write path
  2. the queue is the seam between them
  3. success on ingest does not mean stored
  4. watch queue depth, not API errors

basics

~20 s

The web container accepts ingestion batches, writes the raw payloads to blob storage and enqueues them, then serves the UI and API. The worker consumes that queue asynchronously: it writes traces, observations and scores into ClickHouse and runs background jobs such as evaluators and retention cleanup.

solid answer

~50 s

Langfuse v4 splits ingestion into a fast accept path and a slow write path. The **web** container receives an SDK batch on the ingestion endpoint, persists the raw event bodies to the S3-compatible bucket, pushes references onto the Redis queue and returns immediately. The **worker** container drains that queue and does the expensive work: transforming events into the ClickHouse trace, observation and score tables, and running background jobs such as LLM-as-judge evaluators, exports and data-retention deletion. The practical consequence is the failure mode you get paged for: if the worker is down, undersized or wedged, the SDK still reports success and no application error appears, yet traces never show up in the UI. You diagnose it by watching queue depth in Redis and worker logs, not by looking at the API. Scale the two independently: web with request volume, worker with event volume.

go deeper

for a junior

Know that Langfuse sends traces in the background and that they appear in the UI a moment later rather than instantly, so a script that exits immediately may show nothing.

for a middle

Be able to trace one event through the pipeline: SDK batch to the web ingestion endpoint, raw payload to blob storage, reference to the queue, worker writes into ClickHouse.

for a senior

Demonstrate the operational half: name the silent-failure mode when the worker lags, say which signal you alert on, and explain how you size web and worker on different metrics.

for a principal

Own the capacity model. Decide what ingestion lag your organisation tolerates, whether evaluator jobs share worker capacity with ingestion, and how a traffic spike degrades rather than drops data.

## The shape of the pipeline An LLM application produces a lot of small writes, in bursts, from processes that must not be slowed down by observability. Langfuse v4 answers that with an asynchronous pipeline and a two-container topology, both stateless and both horizontally scalable. The path of a single event: 1. The SDK batches events in the background and posts them to the ingestion endpoint on the **web** container. 2. Web authenticates the project keys, writes the raw event bodies into the S3-compatible bucket, and enqueues references on Redis or Valkey. 3. Web returns to the SDK. From the application's point of view, the trace is delivered. 4. The **worker** consumes the queue, reads the raw payloads, and writes rows into the ClickHouse trace, observation and score tables. 5. The UI, served by web, reads back from ClickHouse and Postgres. Step 3 is why the split exists. The instrumented application never waits on an OLAP insert, and a burst of traffic becomes queue depth rather than request latency or dropped data. ## What else lives on the worker The worker is not only an ingestion consumer. Background work that must not run in a request goes there: server-side evaluator runs, which call an LLM per sampled trace and therefore take seconds each; batch exports and integrations; and the retention job that deletes aged data from ClickHouse and blob storage. That matters for sizing, because an evaluator sampling a busy project competes for the same worker capacity as ingestion. ## The failure mode you must recognise This architecture produces one signature incident, and it is the reason the question is asked at senior level. When the worker is stopped, crash-looping, or simply out-scaled by traffic: - The SDK is happy. Ingestion returned success, so the application logs nothing. - The UI shows no new traces, or traces that lag by minutes and then hours. - Redis queue depth climbs steadily. - The blob bucket keeps growing, because raw events are still being accepted and stored. Nothing in the instrumented application surfaces this. You find it by monitoring queue depth and worker throughput, and by alerting on ingestion-to-visible lag rather than on API errors. "Traces are missing" from a developer is far more often a worker problem than an SDK problem, and the second thing to check, after confirming the SDK flushed, is the queue. ## Scaling and sizing Because both containers are stateless, both scale out. They scale on different signals: - **web** scales with UI users and ingestion request rate. - **worker** scales with event volume and background job load. A common misconfiguration is running one replica of each and assuming symmetry. A chatty agent application generates many observations per user request, so worker capacity is usually the binding constraint long before web is. Equally, a short-lived job that exits immediately after work can lose events that never left the SDK buffer, which is a client-side flush problem rather than a worker one, and the two get confused constantly. ## Interview framing Say the split explicitly: web is the synchronous accept path plus the read path, worker is the asynchronous write path plus background jobs, Redis is the seam, blob storage is the durable landing zone that makes the fast accept safe. Then name the silent-failure consequence. That combination, mechanism plus operational signature, is what distinguishes a real answer from a diagram recital.

  • A developer says traces are missing. How do you decide whether it is the SDK or the backend?
    Split at the ingestion boundary. On the client, confirm the process lived long enough to flush and that the batch was actually posted, since short-lived scripts and serverless handlers commonly exit before the background exporter sends. If the batch was accepted, the problem is behind the boundary: check worker health, queue depth and worker logs. Rising queue depth with a healthy API is the backend signature; nothing ever posted is the client signature.
  • Why does this architecture make blob storage part of the write path rather than an optional extra?
    Because the fast accept is only safe if the payload is already durable. Web writes raw event bodies to the bucket before acknowledging, so a worker crash or a queue restart loses at most references, not content, and processing can be retried against the stored payload. It also keeps large payloads, long prompts and media, out of the queue itself.

saying these in an interview costs you the question

  • Assuming a successful ingestion response means the trace is queryable
  • Running one worker replica regardless of event volume
  • Thinking the worker only exists to run evaluators
  • Diagnosing missing traces by only checking SDK code

context