When should you pass disable_batch=True to OpenLLMetry's Traceloop.init(), and what does it cost?
answer
- how long does the process live?
- buffered spans die with the process
- serverless freezes right after return
- immediate export costs latency per span
- default is right for long-running services
basics
~20 sPass disable_batch=True in short-lived processes — serverless handlers, CLI scripts, notebooks, tests — where the process can end before buffered spans are sent. It exports each span as it finishes, which is safe but adds export work on the calling path, so leave it off in long-running services.
solid answer
~40 sBy default `Traceloop.init()` buffers finished spans and exports them in batches, which is right for a long-running service: fewer network round trips and no export latency on the request path. The catch is that anything still in the buffer when the process dies is lost. In a Lambda handler, a one-shot script, a notebook cell or a unit test, the process finishes right after the model call, so the common outcome is a run that produced no traces at all. `Traceloop.init(app_name="job", disable_batch=True)` switches to exporting each span the moment it ends, so nothing depends on a later flush. The price is a synchronous export per span, which adds latency and a network call per span — acceptable for a handful of spans per invocation, wasteful at thousands of requests per second.
code
python · 8 linesfrom traceloop.sdk import Traceloop
# Short-lived process: export each span as it ends so nothing
# is left in a buffer when the runtime freezes on return.
Traceloop.init(app_name="nightly-summarizer", disable_batch=True)
def handler(event, context):
return {"status": "ok"}go deeper
Remember that disable_batch=True makes each span export immediately, and that you want it in scripts, notebooks and serverless handlers where the process ends quickly.
Explain the tradeoff both ways: batching amortizes network calls but loses whatever is buffered when the process dies, while immediate export is safe but adds synchronous work per span.
Distinguish a flush problem from a shutdown problem. Know that a long-running service should keep batching and fix its termination grace period instead, and be able to describe how missing telemetry at deploy time is diagnosed.
Set the environment policy rather than leaving it per-service: immediate export in jobs and CI, batching plus enforced graceful shutdown in services, and a local collector hop so export cost never lands on user-facing latency.
## Two export rhythms, one flag `Traceloop.init()` accepts `disable_batch`, a boolean that is `False` by default. It selects between the two OpenTelemetry export rhythms: batching (accumulate finished spans, ship them periodically or when the buffer fills) and immediate export (ship each span when it ends). OpenLLMetry exposes it as a single flag because in practice the choice follows from one property of your process — how long it lives. ## Why batching is the default A long-running service handles many requests, each producing several spans. Exporting each one individually would mean a network call per span, on a thread the request is waiting on. Batching amortizes that: spans go into a buffer, a background worker ships them in groups, and the request path pays almost nothing. For a service, that is unambiguously the right default. ## Why batching loses data in short-lived processes Buffering only works if someone drains the buffer. A process that exits — or a serverless runtime that *freezes the execution environment* the moment the handler returns — may never reach the point where the pending batch is sent. The spans exist, they were recorded correctly, and they never leave the machine. The symptom is maddeningly specific: the code clearly ran, the model was clearly called, and the backend shows nothing. The situations where this bites are all variations of "the process is about to end": - **Serverless functions.** After the handler returns, the environment can be frozen instantly; background threads do not get scheduled again until the next invocation, if there is one. - **CLI tools and batch jobs.** The script calls the model, prints a result, exits. - **Notebooks.** You run a cell, look for the trace, and it is not there yet. - **Test suites.** Especially ones that assert on emitted telemetry, where determinism matters more than throughput. Setting `disable_batch=True` removes the dependency on a later flush entirely: by the time the span ends, it has already been sent. ## What it costs Immediate export is synchronous work attached to span end. Each finished span becomes its own export call, on the thread that finished it. Three consequences follow: 1. **Added latency per span.** The caller waits for the export, so a slow or unreachable collector shows up as application slowness rather than as missing telemetry. 2. **More network calls.** One per span rather than one per batch, which matters to the collector and to any per-request billing on the receiving end. 3. **No smoothing.** A burst of spans becomes a burst of exports rather than a steady drip. At a few spans per invocation this is invisible. At production request volume it is a self-inflicted performance problem, and it is a genuine interview red flag to propose it as a blanket setting "so we never lose traces". ## Choosing correctly The decision rule is short: **if the process outlives many requests, batch; if the process is about to exit, do not.** In a container running a web server, leave the default. In a Lambda, a Cloud Run job, a Celery one-shot, a script or a notebook, set `disable_batch=True`. A useful nuance for interviews: a long-running service still loses its final buffer on an abrupt kill. Graceful shutdown — letting the process run its shutdown path rather than sending an immediate hard kill — is what preserves those last spans, not switching the whole service to immediate export. If your orchestrator's termination grace period is too short, you lose the tail of your telemetry during every deploy, and the fix is the grace period, not the flag. ## Where the boundary is `disable_batch` is the OpenLLMetry-level convenience for this choice. The detailed tuning knobs of the underlying batching processor — queue size, schedule delay, maximum batch size, export timeout — belong to OpenTelemetry SDK configuration rather than to OpenLLMetry, and you reach them by supplying your own processor rather than through this flag. Knowing where that line sits is itself a good signal in an interview: OpenLLMetry chooses sane defaults and exposes the one decision that most applications actually need to make.
- A long-running service is redeployed and the last few seconds of traces are always missing. Is disable_batch the fix?No. That is the buffer being lost on shutdown, and switching the entire service to per-span export to protect a few seconds of data every deploy is a bad trade. The fix is graceful shutdown — give the process enough termination grace to run its shutdown path and flush pending spans — plus checking that your orchestrator sends a polite signal before a hard kill.
- You set disable_batch=True in a serverless handler and now cold-start latency is noticeably worse. What is happening?Every span now performs its own synchronous export, so per-invocation wall time includes those round trips, and a cold start adds SDK initialization on top. Options are to reduce span volume by decorating fewer functions, to send to a local collector so the export hop is short, or to accept the cost as the price of not losing traces in a process that cannot flush later.
- Does disable_batch change what is recorded in each span?No — it only changes when spans leave the process. The same attributes, timings and parent/child relationships are captured either way. It is purely an export-rhythm choice, which is why it is safe to flip it per environment: immediate in jobs and tests, batched in services.
saying these in an interview costs you the question
- Setting disable_batch=True everywhere to be safe
- Assuming batching drops spans rather than delaying them
- Believing the flag changes what attributes are captured
- Blaming the exporter when a script exits before the buffer flushes
- Thinking per-span export is free because it happens on span end