skip to content

Langfuse traces never appear from your serverless function. How do you debug it?

level: seniorimportance: should knowfreq 45%

answer

  1. export is asynchronous by design
  2. the container freezes on return
  3. flush is the first fix
  4. then keys, host, project
  5. filters and clocks hide traces too

basics

~20 s

The SDK batches observations and exports them from a background thread, so a handler that returns or freezes before the export loses them. Call flush() before returning, then check credentials, host and project before suspecting the platform.

solid answer

~50 s

Start with the lifecycle, because it explains most cases. Langfuse's SDK does not send an observation when you create it: it buffers and exports asynchronously in the background so tracing never sits in your request path. A serverless container is frozen or killed the moment the handler returns, which kills the exporter mid-flight — so the fix is `get_client().flush()` before returning, or `shutdown()` at the end of a genuinely short-lived process. Only then look at configuration: are `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY` a matched pair, is `LANGFUSE_HOST` pointing at the right cloud region or self-hosted URL, and are you looking at the project those keys belong to? After that, check whether the traces exist but are invisible: wrong environment filter, a time window skewed by a non-UTC clock, or a self-hosted deployment whose ingestion worker is not consuming. Turn on the SDK's debug logging to see whether the exporter is even attempting to send.

code

python · 10 lines
python
from langfuse import get_client, observe

langfuse = get_client()


@observe(name="lambda-handler")
def handler(event, context):
    result = {"answer": "..."}
    langfuse.flush()  # block until buffered observations are sent
    return result

go deeper

for a junior

Know that Langfuse sends traces in the background rather than immediately, so a script or function that ends quickly must call flush() before it exits or the data is lost.

for a middle

Walk the diagnosis in order: flush and process lifetime first, then keys, host and project, then filters and time range, using the SDK's debug logging to see whether an export was even attempted.

for a senior

Explain why the failure is silent — tracing is deliberately fail-safe and out of the request path — and cover the non-serverless variant: flush on SIGTERM so a deploy does not drop the last seconds of traffic.

for a principal

Make it structural: flushing in the shared handler wrapper, a post-deploy smoke check that a trace actually lands, and an explicit position on how much tail latency tracing may add to a request.

## Why this happens at all Langfuse's SDK is deliberately fire-and-forget. Creating an observation writes into an in-process buffer; a background exporter batches and ships it. That design is right — you do not want a tracing backend adding latency to, or failing, a user request — and it is exactly what makes traces disappear in short-lived processes. A serverless function is the pathological case. The platform freezes or reclaims the execution environment as soon as the handler returns. Anything still in the buffer, or a request in flight, dies with it. Nothing errors: your function returns 200 and the trace simply never exists. The same applies to CLI scripts, one-off jobs, notebook processes that get killed, and containers that exit immediately after their work. ## Work the list in order **1. Flush before the process can stop.** Call `get_client().flush()` as the last thing before returning from the handler. `flush()` blocks until the buffer is sent, so it belongs at the end, not sprinkled through hot paths. For a process that is definitely ending, `shutdown()` flushes and stops the exporter cleanly. In long-running servers, do neither per request — rely on background batching and flush on SIGTERM during graceful shutdown, so the last few seconds of traffic is not lost on every deploy. **2. Confirm credentials and destination.** A public/secret key pair belongs to one project; a mismatched pair, or keys from a different project, gives you either rejected ingestion or traces that arrive somewhere you are not looking. `LANGFUSE_HOST` must be set for self-hosted deployments and for any cloud region other than the default — an unset host silently sends to the default endpoint, where your keys are meaningless. In a serverless deployment, also confirm the variables actually reached the function's runtime environment rather than only the build. **3. Ask whether the trace exists but is hidden.** Filters lie convincingly. Check the environment selector, any saved filter on the project, and the time range — Langfuse's own infrastructure is expected to run in UTC, and a host clock that is wrong or interpreted in local time can place your trace outside the window you are looking at. **4. Turn on debug logging.** The SDK is fail-safe: it logs tracing problems rather than raising them into your application, which is desirable in production and unhelpful when nothing appears. Enable its debug logging and you will see whether the exporter is being invoked, what it is sending, and what the server said. "No export attempted" and "export attempted and rejected" are completely different bugs, and this is how you tell them apart. **5. Check ingestion on a self-hosted deployment.** A self-hosted Langfuse accepts events at the web tier and processes them asynchronously. If that processing is not running or is backed up, events are accepted and simply do not become visible traces. If you own the deployment, that is where to look next; if you do not, this is the point to hand it to whoever does, with the evidence that the client got a successful response. ## Related symptoms worth recognising - **Some spans appear, others do not.** That is usually not delivery — it is context. Work handed to a thread pool or a detached task can run outside the active observation context and land in a separate trace rather than being lost. - **The trace exists but has no user or session.** Attribution is set by a context manager that is not retroactive, so entering it late leaves early spans unattributed. - **Everything arrives but late.** Batching plus a queued ingestion pipeline means a delay of seconds is normal; treating a five-second lag as data loss sends people debugging the wrong thing. ## Designing the problem away For serverless, put the flush in the wrapper every handler already uses rather than trusting each author to remember it. Accept that flush adds tail latency to the invocation and measure it — usually milliseconds, occasionally not, and if it is not, that is an argument for shipping traces out-of-process rather than for skipping the flush. And add a smoke check that asserts a trace lands after deploy, because a silent tracing outage is discovered weeks later, in the middle of an incident, when someone finally needs the data. ## Interview framing The grade comes from ordering. Lead with the asynchronous export and the frozen container, because that is the cause in most real cases; then configuration; then visibility; then ingestion. Mentioning that the SDK deliberately does not raise tracing errors into the request path shows you understand why the failure is silent, and mentioning graceful-shutdown flush for long-running servers shows you have thought about the non-serverless version of the same bug.

  • Should a long-running web server call flush() on every request?
    No. Flushing blocks until the buffer is sent, so per-request flushing pushes tracing latency straight into the user's request — the very thing background batching exists to avoid. Rely on the background exporter, and flush once during graceful shutdown on SIGTERM so the final seconds of traffic survive a deploy or a scale-in.
  • How do you tell "never sent" apart from "sent and rejected"?
    Turn on the SDK's debug logging and watch the exporter. If nothing is attempted, the process died before export or tracing was never enabled; if an attempt is made and refused, you have a credential, host or payload problem instead. They look identical from the UI — an empty project — but the fixes have nothing in common, so distinguishing them first saves most of the debugging time.
  • Traces arrive but consistently five seconds late. Is that a bug?
    Usually not. Observations are buffered and exported in batches, and the server processes ingested events asynchronously, so a delay of seconds between the call and a visible trace is normal operation. Treat it as a bug only if the lag grows without bound, which points at a backed-up ingestion pipeline rather than at the client.
  • Only some observations from one request reach Langfuse. What does that suggest?
    Not a delivery problem — partial loss within one request usually means context, not transport. Work handed to a thread pool or a detached task can run without the active observation context and start its own trace instead, so the spans exist but are somewhere else. Look for orphan traces with the missing names before assuming they were dropped.

saying these in an interview costs you the question

  • Assumes each observation is sent synchronously when created
  • Calls flush() on every request in a long-running server
  • Never sets the host for a self-hosted or non-default region
  • Blames the platform before checking keys and project
  • Treats a few seconds of ingestion lag as data loss

context