skip to content

An OpenTelemetry Collector is dropping telemetry under load and occasionally restarting after running out of memory. Explain the roles of the memory_limiter processor, the batch processor and the exporter's sending queue and retry, and how you would locate where data is being lost.

level: seniorimportance: should knowfreq 40%

answer

  1. memory_limiter first: soft limit refuses, hard limit forces GC
  2. Refusal keeps the process alive; it does not add capacity
  3. batch after filtering/sampling, before export
  4. Queue full = drop, not backpressure; persistence needs file_storage
  5. receiver_refused / exporter_enqueue_failed / queue_size vs capacity

basics

~20 s

memory_limiter refuses new data when heap crosses a soft limit so the process survives; batch amortises export cost; the exporter's sending_queue plus retry absorb backend outages and drop when full. Locate loss with the Collector's own receiver/processor/exporter counters.

solid answer

~50 s

Loss happens at one of three places and the Collector's self-telemetry tells you which. **At the receiver**, `memory_limiter` (placed first) checks heap on an interval; past its soft limit it refuses incoming data, which surfaces to the sender as an error — with OTLP/gRPC a `RESOURCE_EXHAUSTED` the client may retry — and past the hard limit it forces garbage collection. It protects the process; it does not create capacity. **In the middle**, `batch` groups records by `send_batch_size`/`timeout` to amortise export overhead; put it after filtering and sampling. **At the exporter**, `sending_queue` buffers batches for `num_consumers` workers and `retry_on_failure` retries with exponential backoff; when the queue is full new batches are dropped, and only the `file_storage`-backed persistent queue survives a restart. Diagnose with the internal metrics: `otelcol_receiver_refused_*` means the limiter is refusing, `otelcol_exporter_enqueue_failed_*` and a `queue_size` pinned at `queue_capacity` mean the backend cannot keep up, `otelcol_exporter_send_failed_*` means it is erroring. Then fix the actual cause: scale out, shed volume, or extend the queue.

code

text · 5 lines
text
receiver_accepted 100k/s | receiver_refused 0     -> door is open
exporter_queue_size == exporter_queue_capacity     -> backend too slow
exporter_enqueue_failed 12k/s                      -> drops happen HERE
exporter_send_failed ~0, exporter_sent 40k/s       -> backend accepting, just slower than ingress
=> scale out / raise consumers / shed volume; a bigger queue only delays the same drop

go deeper

for a junior

Know that batch groups records for efficient export, that memory_limiter protects against OOM, and that the exporter has a queue that can fill up.

for a middle

Explain limiter soft/hard limits and correct ordering, and that a full queue drops rather than blocks.

for a senior

Diagnose by reading the internal counters end to end, distinguish burst absorption from a rate mismatch, and know when persistence and more consumers help.

for a principal

Set the policy: how much telemetry loss is acceptable and where it should occur, whether the fix is capacity or volume reduction, and what alerts make silent drops visible before an incident does.

## Three places data dies A Collector accepts, transforms and forwards. Loss therefore occurs at the door, in the chain, or on the way out — and each has its own control and its own counter. Debugging without naming which one you are in is guesswork. ## memory_limiter — self-preservation, not capacity The `memory_limiter` processor periodically (`check_interval`, typically 1s) samples heap usage against a configured `limit_mib` and `spike_limit_mib`. The soft limit is `limit - spike`. Above the soft limit the processor **refuses** incoming data, returning an error to the component upstream; above the hard limit it also forces garbage collection. Refusal is the point: an unbounded Collector under a traffic spike simply dies, losing everything buffered, whereas a limiter sheds the newest arrivals and stays alive to deliver what it already holds. It belongs **first** in the processor list so refusal happens before the pipeline allocates for the record. It must be sized against the actual memory available to the process — in a container, against the container limit, leaving headroom for the runtime — and, critically, it is not a capacity solution. Persistent refusal means the Collector is undersized or the backend is too slow, and the limiter is only choosing how you fail. Whether refusal turns into real backpressure depends on the sender. Over OTLP/gRPC the Collector returns `RESOURCE_EXHAUSTED`, which a well-behaved client retries with backoff; but the client's own queue is bounded too, so sustained refusal becomes loss at the sender rather than at the Collector. ## batch — throughput economics The `batch` processor accumulates records and emits them when `send_batch_size` is reached or `timeout` elapses, with `send_batch_max_size` capping the emitted batch so downstream message-size limits are respected. Batching is what makes export efficient: fewer requests, better compression, less per-request overhead. Two placement rules follow. Put it **after** anything that filters, samples or drops, so you do not assemble batches you throw away. Put it **before** the exporter, obviously, and understand that a large `timeout` adds delivery latency while a large `send_batch_size` adds memory and risks exceeding the receiving side's maximum message size. ## The exporter queue and retry Each exporter has its own `sending_queue` — a bounded queue of batches drained by `num_consumers` concurrent workers — and `retry_on_failure`, exponential backoff with a maximum elapsed time. Together they absorb transient backend failures: a thirty-second backend blip is invisible if the queue can hold thirty seconds of data. The important semantics: when the queue is **full**, enqueue fails and the batch is **dropped**, immediately and quietly, unless the build and configuration support a blocking queue that propagates backpressure instead. And when `retry_on_failure` exhausts its maximum elapsed time, the batch is dropped too. By default the queue is in memory, so a restart loses it; adding the `file_storage` extension and enabling persistence writes it to disk so a Collector restart or crash does not vaporise the buffer. Persistence costs disk I/O and adds a failure mode of its own (a full disk), so it is a deliberate choice for gateway tiers, rarely for agents. Because exporters fan out with independent queues, one slow backend fills only its own queue — a genuinely useful isolation property when dual-shipping. ## Locating the loss The Collector emits telemetry about itself, and this is the whole diagnosis. Reading the pipeline left to right: - `otelcol_receiver_accepted_spans` / `otelcol_receiver_refused_spans` — refused climbing means the limiter (or a receiver-level limit) is rejecting at the door. - processor-level dropped/refused counters — deliberate drops from filtering or sampling versus pressure. - `otelcol_exporter_queue_size` versus `otelcol_exporter_queue_capacity` — a queue pinned at capacity means the backend cannot absorb your rate. - `otelcol_exporter_enqueue_failed_spans` — the queue was full and batches were discarded. - `otelcol_exporter_send_failed_spans` / `otelcol_exporter_sent_spans` — the backend is erroring or throttling. Accepted high with sent near zero and no failures at all points at a wiring mistake rather than pressure — data entering a pipeline whose exporter is not referenced. Process-level heap and GC metrics, plus the `pprof` extension, complete the picture for the OOM side. ## Fixing it Order the interventions by cause. Backend throttling: raise `num_consumers` cautiously, lengthen retry, or negotiate quota — a bigger queue only delays the same drop. Genuine volume growth: scale the tier out, and for gateways ensure the load balancer spreads evenly. Bursty traffic against a steady backend: enlarge the queue, ideally persistent. Memory pressure from oversized batches: shrink `send_batch_size` and cap `send_batch_max_size`. And often the correct answer is volume reduction — sample, filter noisy spans, or prune high-cardinality attributes — because a pipeline sized for telemetry nobody reads is a bad trade regardless of tuning. ## The habit to build Alert on refused and enqueue-failed counters and on queue occupancy as a fraction of capacity. Silent loss is the characteristic Collector failure: the process is up, the dashboards are populated, and a fraction of your traces simply never existed.

  • Why is a bigger sending queue often the wrong answer to sustained drops?
    A queue absorbs bursts, not a rate mismatch. If ingress exceeds the rate the backend accepts, any finite queue fills at a predictable pace and then drops exactly as before, only later and after consuming more memory. Sustained pressure needs more export throughput, more replicas, or less data — the queue is a shock absorber, not capacity.
  • Does the memory_limiter propagate backpressure to the application?
    Only partially. Refusal returns an error to the caller — for OTLP/gRPC typically RESOURCE_EXHAUSTED — and a compliant client retries with backoff. But SDK and agent queues are small and bounded by design so they never threaten the application, so sustained refusal simply relocates the drop upstream. The limiter's real job is keeping the Collector alive, not creating end-to-end flow control.
  • When is a persistent sending queue worth the cost?
    On a gateway tier where the buffered data is expensive to re-create and backend outages are measured in minutes: persistence lets a restart or crash resume rather than lose the buffer. It costs disk throughput and introduces a full-disk failure mode, and it makes little sense on thousands of node agents, where the same durability would multiply I/O across the serving fleet for data that is cheap to lose from one node.

saying these in an interview costs you the question

  • Treating memory_limiter as a way to handle more traffic rather than to survive it
  • Assuming a full sending queue blocks the sender instead of dropping
  • Expecting an in-memory queue to survive a Collector restart
  • Placing batch before sampling or filtering and paying to assemble discarded batches
  • Concluding "no errors in the log, so no data loss" — drops are counted, not logged

context