An application can export OTLP telemetry straight to a vendor backend. What does running an OpenTelemetry Collector in between buy you, and when is direct export from the SDK good enough?
answer
- Receive → process → export, in a separate process
- Late binding: backend swap = config, not redeploy
- Queue and retry outside the app heap
- k8sattributes / resourcedetection see what the app cannot
- Tail sampling and span metrics need a meeting point
basics
~20 sThe Collector is a separate process that receives, processes and re-exports telemetry. It decouples apps from backends: batching, retry, enrichment, redaction, translation and routing become config in one place instead of code in every service.
solid answer
~50 sThe Collector is a standalone, vendor-neutral binary with receivers, processors and exporters. It is optional — an SDK can speak OTLP to any OTLP backend — but it buys four things. **Late binding**: the destination, credentials, redaction rules and sampling live in Collector config, so switching or dual-shipping to a second backend is a config change, not a redeploy of fifty services. **Buffering outside the app**: bounded queues, exponential-backoff retry, and optionally a disk-backed queue, so a backend outage burns Collector memory rather than application heap. **Enrichment**: processors like `k8sattributes` and `resourcedetection` attach pod, node, cloud-region and instance identity the app cannot see. **Translation and global functions**: receivers for Prometheus scrape, Jaeger, Zipkin, statsd, fluentforward and log files converge on one egress format, and whole-trace work like tail sampling or span-derived metrics is only possible where many traces meet. Direct export is fine for prototypes, single-service systems, and local development.
go deeper
Say what it is (a separate receive-process-export process), that it is optional, and give two concrete wins: swap backends without redeploying, and enrich with Kubernetes metadata the app cannot see.
Add the buffering/retry argument and the translation role (Prometheus, Jaeger, Zipkin, file logs in; OTLP out), and name the cost of running another process.
Frame it as where policy lives: destination, credentials, redaction, sampling and cost shaping become one configuration surface, and whole-trace functions become possible at all.
Argue the organisational case — a Collector tier is the control point for telemetry spend, PII egress and vendor lock-in — and state the availability contract: telemetry failure must never degrade the serving path.
## What the Collector actually is The OpenTelemetry Collector is an ordinary process (a Go binary, usually shipped as the `otelcol` core distribution, the larger `otelcol-contrib`, or a custom build produced by the OpenTelemetry Collector Builder). It receives telemetry over the network or from files, runs it through processors, and exports it onward. It is **not** part of the protocol: an SDK that speaks OTLP can talk directly to any backend that accepts OTLP. So the question is always "what does this extra hop earn?" ## Job 1 — decoupling the app from the backend Without a Collector, every service embeds the destination endpoint, the vendor API key, the retry policy and any redaction logic. Changing vendors, adding a second destination during a migration, or lowering export volume means rebuilding and redeploying every service. With a Collector, applications export to one local, unauthenticated endpoint (typically OTLP/gRPC on `4317` or OTLP/HTTP on `4318` on localhost or a cluster service), and everything downstream is Collector configuration. Dual-shipping to old and new backends for a migration window is two exporters in one pipeline. It also contains credentials: the backend token lives in one deployment's environment, not in every service image. ## Job 2 — buffering and retry outside the application process SDK exporters have small in-process queues by design; they must never consume the heap that serves user traffic, so under a backend outage they drop. A Collector can hold a much larger bounded queue, retry with exponential backoff, and — with the `file_storage` extension enabling a persistent sending queue — survive its own restart with data on disk. The failure mode moves out of the application, which is the point: telemetry loss should never become an application incident. ## Job 3 — enrichment the application cannot do A process does not know its Kubernetes pod name, namespace, node, deployment, or its cloud region and instance id — and usually should not hold the API credentials required to look them up. A node-local Collector can, and processors such as `k8sattributes` and `resourcedetection` attach those as resource attributes to every record passing through. This is why the agent tier exists even in shops that also run a gateway. ## Job 4 — translation and whole-fleet functions Receivers exist for Prometheus scrape, Jaeger, Zipkin, statsd, fluentforward, host metrics, container logs and Kafka. That lets a Collector be the bridge from a legacy telemetry estate into OTLP without touching legacy apps. And some processing is *only* possible where many streams converge: tail-based sampling needs every span of a trace, span-derived RED metrics need aggregation across requests, cardinality limiting needs a view of the whole series population. No single application process can do these. ## The costs, stated honestly It is one more thing to deploy, version, monitor, capacity-plan and debug. It has its own failure modes — out-of-memory under a traffic spike, a full sending queue silently dropping, a misconfigured pipeline that starts but exports nowhere. It adds latency to *telemetry delivery*, though not to user requests, since export is asynchronous. And it cannot fix what was never captured: missing instrumentation, a wrong `service.name` at the source, or a head-sampling decision already made in the SDK that discarded the span before it ever left the process. ## When direct export is fine A prototype, a local dev loop, a single service, or a small system with one backend and no compliance requirement on egress. Short-lived serverless functions are a special case: there may be no node-local agent to send to and no time to drain a queue, so many teams point them at a gateway Collector rather than an agent, or export directly and accept the coupling. ## Availability Because the Collector sits in the path of telemetry only, its outage must degrade observability, never the application. SDK exporters are expected to fail quietly and keep serving. Reduce blast radius by running the agent tier per node (one node's telemetry at risk) and replicating the gateway tier behind a load balancer, and watch the Collector's own internal metrics rather than assuming silence means health.
- If the Collector goes down, what happens to the application?Nothing user-visible should happen. SDK exporters are required to fail without propagating errors into application code; they log, retry within a bounded queue, and drop when it fills. You lose telemetry for the outage window, which is why the agent tier is deployed per node to limit blast radius and the gateway tier is replicated. A disk-backed persistent queue can preserve data across a Collector restart, but nothing preserves data the SDK already dropped.
- Name something a Collector cannot fix that you must get right in the application.Anything decided before export: instrumentation that was never added, a head-sampling decision that discarded a trace at creation time, span names and attributes with the wrong semantics, and the identity of the service itself. You can rewrite `service.name` in a transform processor, but then two teams' config disagree about who owns the data — better to set it correctly at the source.
A mail room. Every desk could stamp and post its own parcels, but you put a mail room in the middle so postage accounts, address changes, redaction of sensitive contents and courier switches happen once, not at every desk.
saying these in an interview costs you the question
- Believing OTLP requires a Collector, or that the Collector is part of the SDK
- Claiming the Collector instruments the application — it never sees code, only exported records
- Assuming a Collector guarantees no data loss; its queue is bounded and drops when full unless disk-backed
- Treating the Collector as a storage backend rather than a pipeline
- Saying it removes all latency concerns — export is async either way; the win is operational, not latency