skip to content

Compare the simple and batching span processors in the OpenTelemetry SDK. What do the queue size, schedule delay, maximum export batch size and export timeout settings control, what happens when the queue fills, and what must a short-lived process do before exiting?

level: seniorimportance: must knowfreq 45%

answer

  1. simple = sync per span, debug only
  2. batch = bounded queue + timer + batch cap
  3. queue full ⇒ drop, never block the app
  4. force flush before a short-lived process exits
  5. metrics use a periodic reader, not a processor

basics

~20 s

The simple processor exports each span synchronously as it ends — correct but slow, for debugging. The batching processor enqueues ended spans and a background thread exports batches on a timer or when full; a full queue drops spans silently. Short-lived processes must force-flush and shut down or lose buffered spans.

solid answer

~60 s

A span processor sits between span end and the exporter. **Simple** hands each ended span straight to the exporter, blocking the application thread on a network call — acceptable only in tests or with a console exporter. **Batching** puts ended spans on a bounded queue; a background worker drains it into export batches. Four knobs, with the usual spelling of their environment variables: **max queue size** (`OTEL_BSP_MAX_QUEUE_SIZE`, commonly 2048) is the in-memory buffer and your memory ceiling; **schedule delay** (`OTEL_BSP_SCHEDULE_DELAY`, commonly 5s) is how long a span can wait before export; **max export batch size** (`OTEL_BSP_MAX_EXPORT_BATCH_SIZE`, commonly 512) caps one request; **export timeout** (`OTEL_BSP_EXPORT_TIMEOUT`) bounds a single export attempt. When the queue is full, new spans are **dropped, not blocked** — deliberately, so telemetry never stalls the application. The loss is silent unless you watch the SDK's own dropped-span metric. Only sampled spans reach the processor's export path, so processor tuning follows the sampling decision. And a batch of buffered spans dies with the process: CLI jobs and function invocations must force-flush and shut down.

go deeper

for a junior

Know that batching exports spans in the background while the simple processor exports each span immediately, and that the batching one is what production uses.

for a middle

Explain the four settings and their trade-offs, and that a full queue drops spans rather than blocking.

for a senior

Diagnose silent drops from self-metrics, relate processor sizing to the sampled volume, and handle flush and shutdown for short-lived processes.

for a principal

Set fleet defaults for queue and delay against a memory budget, mandate self-telemetry on drops and export failures, and decide where durability lives — in-process buffering or a local collector.

## Where processors sit Span creation → sampler decision → the span runs → span end → **span processor** → exporter → the wire. The processor's job is to decouple the application from the exporter's latency and failure modes. ## Simple processor On every span end, export immediately and synchronously. Its virtue is zero delay and zero buffering: what ends is on the wire. Its vice is that the thread that finished your request now waits on a network call, and every span costs one request. Use it with a console exporter while debugging, or in tests where determinism beats throughput. It is not a production configuration. ## Batching processor Ended spans go onto a bounded queue; a background worker drains them into batches and calls the exporter. The four settings interact: - **Max queue size** — how many ended spans may wait. This is your memory bound: queue size × average span footprint. Raising it buys tolerance for bursts and backend hiccups, and costs heap. - **Schedule delay** — the maximum time a span waits before the worker exports what it has. Lowering it reduces the lag between an event happening and being visible, at the cost of smaller, more frequent requests. - **Max export batch size** — the ceiling on one export request. Bounds payload size and per-request latency; a batch this large triggers an export before the timer fires. - **Export timeout** — how long one export attempt may take before being abandoned. Too generous and a stalled backend blocks the worker while the queue fills behind it. ## The drop behaviour, and why it is right If the queue is full when a span ends, the span is discarded. The alternative — blocking the application thread until space frees — would let a slow or unreachable backend degrade the very service you are observing. Telemetry must fail before the application does. The cost is that loss is **silent**: the trace simply has holes, or entire traces vanish, with no error in your logs. This is why SDK self-diagnostics matter. SDKs expose their own metrics for queued and dropped spans and for export failures; those should be on a dashboard, because 'my traces are incomplete' is otherwise indistinguishable from a sampling or instrumentation problem. The common root causes are a genuinely slow backend, an export timeout too large relative to the schedule delay, or a burst far beyond the queue size. ## Interaction with sampling Spans that the sampler dropped never reach the export path, and record-only spans are recorded but not exported. So the processor is sized for the *sampled* volume, not the total. If you change the sampling rate, the batching settings need revisiting — a tenfold increase in kept spans against an unchanged queue means drops. ## Shutdown, flush and short-lived processes Batching implies that at any moment some spans exist only in memory. Two operations matter: - **force flush** — export everything queued now and wait for completion. - **shutdown** — flush and then release the exporter, after which the processor accepts nothing more. A long-running server should hook shutdown into its termination path so the final seconds of telemetry survive a rolling deploy. A **short-lived process** — a CLI, a batch job, a serverless invocation — is far more exposed: it can easily die with its entire telemetry still in the queue, so it must force-flush before exit. In freeze-the-process execution environments the flush must additionally happen before the runtime is suspended, not in a background thread that is never scheduled again. Where flush latency per invocation is unacceptable, the usual answer is to export to a local collector instead so the handoff is cheap and durability moves out of the process. ## The sibling mechanisms Metrics do not use span processors. They use a **metric reader**, typically a periodic exporting reader on its own interval (`OTEL_METRIC_EXPORT_INTERVAL`, commonly 60s), because metrics are aggregated state read on a cadence rather than a stream of finished items. Logs mirror traces with their own batching log-record processor and its own `OTEL_BLRP_*` variables. Same shape, separate knobs — do not assume tuning one signal tuned the others.

  • Traces are arriving with random spans missing and nothing is logged as an error. Where do you look first?
    The batching processor's drop counter in the SDK's self-metrics. A full queue discards spans silently by design, and the pattern of randomly missing spans across otherwise healthy traces is its signature. Check export failure and duration metrics alongside it: a backend slow enough to make exports time out will back the queue up even at modest volume.
  • A function-style workload emits traces that only sometimes arrive. Why, and what would you change?
    The invocation ends while spans are still queued and the environment freezes or destroys the process before the background worker exports them. Force-flush before returning, so export completes within the invocation. If that latency is unacceptable, send to a local collector process or extension so the in-process handoff is fast and durability is handled outside the invocation.

saying these in an interview costs you the question

  • Running the simple processor in production because it is 'more reliable'
  • Assuming a full queue blocks or back-pressures the application instead of dropping spans
  • Never checking SDK self-metrics, so silent drops look like an instrumentation or sampling problem
  • Letting a batch job or function exit without force-flushing
  • Tuning the batch processor after raising the sampling rate without revisiting queue size

context