skip to content

Collector & Pipelines

The pipeline process between your apps and your backends: receivers in, processors like batch, filter, transform and tail sampling in the middle, exporters out. Expect to compare agent-per-host and gateway topologies, and to explain why tail sampling can only happen where whole traces meet.

on this pageshow

questions

6

An application can export OTLP telemetry straight to a vendor backend. What does running an OpenTelemetry Collector in between buy you, and when is direct export from the SDK good enough?

level: juniorimportance: must knowfreq 55%

answer

  1. Receive → process → export, in a separate process
  2. Late binding: backend swap = config, not redeploy
  3. Queue and retry outside the app heap
  4. k8sattributes / resourcedetection see what the app cannot
  5. Tail sampling and span metrics need a meeting point

basics

~20 s

The Collector is a separate process that receives, processes and re-exports telemetry. It decouples apps from backends: batching, retry, enrichment, redaction, translation and routing become config in one place instead of code in every service.

solid answer

~50 s

The Collector is a standalone, vendor-neutral binary with receivers, processors and exporters. It is optional — an SDK can speak OTLP to any OTLP backend — but it buys four things. **Late binding**: the destination, credentials, redaction rules and sampling live in Collector config, so switching or dual-shipping to a second backend is a config change, not a redeploy of fifty services. **Buffering outside the app**: bounded queues, exponential-backoff retry, and optionally a disk-backed queue, so a backend outage burns Collector memory rather than application heap. **Enrichment**: processors like `k8sattributes` and `resourcedetection` attach pod, node, cloud-region and instance identity the app cannot see. **Translation and global functions**: receivers for Prometheus scrape, Jaeger, Zipkin, statsd, fluentforward and log files converge on one egress format, and whole-trace work like tail sampling or span-derived metrics is only possible where many traces meet. Direct export is fine for prototypes, single-service systems, and local development.

go deeper

for a junior

Say what it is (a separate receive-process-export process), that it is optional, and give two concrete wins: swap backends without redeploying, and enrich with Kubernetes metadata the app cannot see.

for a middle

Add the buffering/retry argument and the translation role (Prometheus, Jaeger, Zipkin, file logs in; OTLP out), and name the cost of running another process.

for a senior

Frame it as where policy lives: destination, credentials, redaction, sampling and cost shaping become one configuration surface, and whole-trace functions become possible at all.

for a principal

Argue the organisational case — a Collector tier is the control point for telemetry spend, PII egress and vendor lock-in — and state the availability contract: telemetry failure must never degrade the serving path.

## What the Collector actually is The OpenTelemetry Collector is an ordinary process (a Go binary, usually shipped as the `otelcol` core distribution, the larger `otelcol-contrib`, or a custom build produced by the OpenTelemetry Collector Builder). It receives telemetry over the network or from files, runs it through processors, and exports it onward. It is **not** part of the protocol: an SDK that speaks OTLP can talk directly to any backend that accepts OTLP. So the question is always "what does this extra hop earn?" ## Job 1 — decoupling the app from the backend Without a Collector, every service embeds the destination endpoint, the vendor API key, the retry policy and any redaction logic. Changing vendors, adding a second destination during a migration, or lowering export volume means rebuilding and redeploying every service. With a Collector, applications export to one local, unauthenticated endpoint (typically OTLP/gRPC on `4317` or OTLP/HTTP on `4318` on localhost or a cluster service), and everything downstream is Collector configuration. Dual-shipping to old and new backends for a migration window is two exporters in one pipeline. It also contains credentials: the backend token lives in one deployment's environment, not in every service image. ## Job 2 — buffering and retry outside the application process SDK exporters have small in-process queues by design; they must never consume the heap that serves user traffic, so under a backend outage they drop. A Collector can hold a much larger bounded queue, retry with exponential backoff, and — with the `file_storage` extension enabling a persistent sending queue — survive its own restart with data on disk. The failure mode moves out of the application, which is the point: telemetry loss should never become an application incident. ## Job 3 — enrichment the application cannot do A process does not know its Kubernetes pod name, namespace, node, deployment, or its cloud region and instance id — and usually should not hold the API credentials required to look them up. A node-local Collector can, and processors such as `k8sattributes` and `resourcedetection` attach those as resource attributes to every record passing through. This is why the agent tier exists even in shops that also run a gateway. ## Job 4 — translation and whole-fleet functions Receivers exist for Prometheus scrape, Jaeger, Zipkin, statsd, fluentforward, host metrics, container logs and Kafka. That lets a Collector be the bridge from a legacy telemetry estate into OTLP without touching legacy apps. And some processing is *only* possible where many streams converge: tail-based sampling needs every span of a trace, span-derived RED metrics need aggregation across requests, cardinality limiting needs a view of the whole series population. No single application process can do these. ## The costs, stated honestly It is one more thing to deploy, version, monitor, capacity-plan and debug. It has its own failure modes — out-of-memory under a traffic spike, a full sending queue silently dropping, a misconfigured pipeline that starts but exports nowhere. It adds latency to *telemetry delivery*, though not to user requests, since export is asynchronous. And it cannot fix what was never captured: missing instrumentation, a wrong `service.name` at the source, or a head-sampling decision already made in the SDK that discarded the span before it ever left the process. ## When direct export is fine A prototype, a local dev loop, a single service, or a small system with one backend and no compliance requirement on egress. Short-lived serverless functions are a special case: there may be no node-local agent to send to and no time to drain a queue, so many teams point them at a gateway Collector rather than an agent, or export directly and accept the coupling. ## Availability Because the Collector sits in the path of telemetry only, its outage must degrade observability, never the application. SDK exporters are expected to fail quietly and keep serving. Reduce blast radius by running the agent tier per node (one node's telemetry at risk) and replicating the gateway tier behind a load balancer, and watch the Collector's own internal metrics rather than assuming silence means health.

  • If the Collector goes down, what happens to the application?
    Nothing user-visible should happen. SDK exporters are required to fail without propagating errors into application code; they log, retry within a bounded queue, and drop when it fills. You lose telemetry for the outage window, which is why the agent tier is deployed per node to limit blast radius and the gateway tier is replicated. A disk-backed persistent queue can preserve data across a Collector restart, but nothing preserves data the SDK already dropped.
  • Name something a Collector cannot fix that you must get right in the application.
    Anything decided before export: instrumentation that was never added, a head-sampling decision that discarded a trace at creation time, span names and attributes with the wrong semantics, and the identity of the service itself. You can rewrite `service.name` in a transform processor, but then two teams' config disagree about who owns the data — better to set it correctly at the source.

A mail room. Every desk could stamp and post its own parcels, but you put a mail room in the middle so postage accounts, address changes, redaction of sensitive contents and courier switches happen once, not at every desk.

saying these in an interview costs you the question

  • Believing OTLP requires a Collector, or that the Collector is part of the SDK
  • Claiming the Collector instruments the application — it never sees code, only exported records
  • Assuming a Collector guarantees no data loss; its queue is bounded and drops when full unless disk-backed
  • Treating the Collector as a storage backend rather than a pipeline
  • Saying it removes all latency concerns — export is async either way; the win is operational, not latency

context

open as a page

Walk through how an OpenTelemetry Collector configuration is structured — receivers, processors, exporters, connectors, extensions and the `service` section — and explain what determines the order data flows through it.

level: middleimportance: must knowfreq 56%

basics

~20 s

Top-level blocks declare components by id; the service.pipelines block wires them into per-signal pipelines. Data enters a receiver, passes processors in the declared order, then fans out to all exporters in parallel. Declared but unreferenced components are never started.

open as a page

Compare deploying the OpenTelemetry Collector as a per-host or sidecar agent versus a central gateway cluster. What work belongs in each tier, and how do you scale the gateway tier without breaking stateful processing?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Agents run next to workloads for local-only work: a cheap local endpoint, host and pod enrichment, log/host-metric collection. Gateways are a shared cluster for egress control, credentials, routing and whole-trace work. Scale gateways horizontally, but stateful processors need trace-affinity routing.

open as a page

Explain how tail-based sampling works in the OpenTelemetry Collector — the decision buffer, policies, and late-arriving spans — and what it costs compared with deciding at span creation time.

level: seniorimportance: must knowfreq 44%

basics

~20 s

The tail_sampling processor buffers spans by trace id for a wait window, then evaluates policies (latency, error status, attribute match, probabilistic) against the assembled trace and keeps or drops the whole trace. It costs memory and requires trace affinity.

open as a page

An OpenTelemetry Collector is dropping telemetry under load and occasionally restarting after running out of memory. Explain the roles of the memory_limiter processor, the batch processor and the exporter's sending queue and retry, and how you would locate where data is being lost.

level: seniorimportance: should knowfreq 40%

basics

~20 s

memory_limiter refuses new data when heap crosses a soft limit so the process survives; batch amortises export cost; the exporter's sending_queue plus retry absorb backend outages and drop when full. Locate loss with the Collector's own receiver/processor/exporter counters.

open as a page

Telemetry spend is doubling every quarter and legal wants user email addresses removed from span attributes. How would you use OpenTelemetry Collector processors — filter, transform (OTTL), attributes — and connectors to shape volume and scrub data, and what breaks if you get it wrong?

level: principalimportance: should knowfreq 29%

basics

~20 s

Scrub PII at the agent with transform/OTTL or redaction so raw values never leave the node; shape volume centrally with filter, attribute pruning and aggregation connectors. Measure before dropping, and never break trace continuity or metric series continuity.

open as a page