skip to content

Walk through how an OpenTelemetry Collector configuration is structured — receivers, processors, exporters, connectors, extensions and the `service` section — and explain what determines the order data flows through it.

level: middleimportance: must knowfreq 56%

answer

  1. Declare at top level, wire in `service.pipelines`
  2. Pipeline type = signal: traces / metrics / logs
  3. Receivers fan in, processors chain in order, exporters fan out
  4. Connector = exporter of one pipeline + receiver of another
  5. Extensions live outside pipelines; unreferenced = never started

basics

~20 s

Top-level blocks declare components by id; the service.pipelines block wires them into per-signal pipelines. Data enters a receiver, passes processors in the declared order, then fans out to all exporters in parallel. Declared but unreferenced components are never started.

solid answer

~50 s

A Collector config has declaration blocks — `receivers`, `processors`, `exporters`, `connectors`, `extensions` — where each component gets an id of the form `type` or `type/name`, and a `service` block that actually wires things up. Pipelines under `service.pipelines` are **typed by signal**: `traces`, `metrics`, `logs`, each listing receivers, an ordered processor list and exporters. Declaration alone does nothing; a component not referenced by `service` is never instantiated. Within a pipeline the processor list is a chain and **order is significant** — `memory_limiter` first, enrichment early, `batch` last — while exporters receive a fan-out copy in parallel, so one slow exporter does not reorder another. Connectors are the seam between pipelines: a connector is an exporter in one pipeline and a receiver in another, which is how `spanmetrics` turns a trace pipeline into a metrics pipeline. Extensions such as `health_check`, `pprof`, `zpages` and `file_storage` sit outside pipelines entirely and are listed under `service.extensions`.

code

text · 10 lines
text
service:
  pipelines:
    traces:
      receivers:  [otlp]
      processors: [memory_limiter, k8sattributes, batch]
      exporters:  [spanmetrics, otlp/vendor]   # connector acts as an exporter here
    metrics/from-spans:
      receivers:  [spanmetrics]                # ...and as a receiver here
      processors: [batch]
      exporters:  [otlp/vendor]

go deeper

for a junior

Recall the five declaration blocks plus service, and that data goes receiver → processors in order → exporters.

for a middle

Be precise about per-signal pipeline typing, type/name ids, ordering rules for memory_limiter/enrichment/batch, and that unreferenced components never start.

for a senior

Add fan-in/fan-out semantics, the shared-instance rule for receivers and exporters versus per-pipeline processor instances, and how connectors let you cross signals.

for a principal

Talk about config as a governed artifact: composed from fragments, environment-substituted, validated in CI, with self-telemetry that proves each pipeline is actually carrying data.

## Two halves: declaration and wiring Every Collector configuration file has the same shape. The top-level maps `receivers:`, `processors:`, `exporters:`, `connectors:` and `extensions:` **declare and configure** components. The `service:` map **activates** them. This split trips people up constantly: a perfectly configured exporter that no pipeline references is dead configuration — the Collector will not start it and will not warn you loudly. Conversely, referencing an id in `service` that was never declared is a startup failure. ## Component ids A component id is either its type (`otlp`) or `type/name` (`otlp/backend-b`, `filter/drop-healthchecks`). The suffix exists so you can run two differently configured instances of the same component type — for example an `otlp` receiver on the standard ports and another on an internal-only port, or two `otlp` exporters pointing at different vendors during a migration. ## Pipelines are per-signal Under `service.pipelines` you declare pipelines whose *type* is the signal: `traces`, `metrics`, `logs` (and `profiles` in newer, still-developing builds). A pipeline id follows the same `type/name` convention, so `traces/tail-sampled` and `traces/raw` can coexist. A component may only appear in a pipeline of a signal it supports — the `tail_sampling` processor is traces-only, the `prometheus` receiver is metrics-only. ``` receivers: otlp: protocols: grpc: { endpoint: 0.0.0.0:4317 } http: { endpoint: 0.0.0.0:4318 } processors: memory_limiter: { check_interval: 1s, limit_mib: 4000, spike_limit_mib: 800 } batch: { timeout: 5s, send_batch_size: 8192 } exporters: otlp/vendor: { endpoint: ingest.example:4317 } debug: { verbosity: normal } extensions: [health_check, pprof, zpages] service: extensions: [health_check, pprof, zpages] pipelines: traces: receivers: [otlp] processors: [memory_limiter, batch] exporters: [otlp/vendor, debug] ``` ## Flow and ordering rules Data enters through **any** receiver in the list (receivers fan **in**), traverses the processor list **in the written order**, and is then handed to **every** exporter in the list (exporters fan **out**, concurrently). Three consequences matter in practice: - Processor order is semantic, not cosmetic. `memory_limiter` belongs first so it can refuse data before the rest of the chain allocates for it. Enrichment (`k8sattributes`, `resourcedetection`) belongs early so later filters can match on the attributes it adds. Anything that drops or samples belongs before `batch`, so you are not paying to assemble batches you then discard, and `batch` belongs last so batches are not re-cut afterwards. - Exporter fan-out means each exporter gets its own copy and its own queue and retry state. A backend that is slow does not stall a healthy one, but it also does not get "the same" data if a processor mutates in place — which is why mutation belongs in the processor chain, and per-destination differences belong in separate pipelines joined by a `forward` connector. - Instantiation is not uniform: **receivers and exporters referenced from several pipelines are shared single instances** (one listening socket, one connection pool), while **a processor referenced from several pipelines gets a separate instance per pipeline**, because processors may carry per-pipeline state. ## Connectors A connector is simultaneously the exporter of one pipeline and the receiver of another, which lets you cross signal types or split traffic. `spanmetrics` consumes spans and emits request/duration metrics; `count` emits counts of records; `routing` sends records down different pipelines based on attributes; `forward` simply chains one pipeline into another so you can share a common tail of processors. Because a connector links two pipelines, both must exist in `service.pipelines` — the producing pipeline lists it under `exporters`, the consuming one under `receivers`. ## Extensions Extensions provide capability that is not per-record: `health_check` exposes liveness/readiness, `pprof` exposes Go profiling, `zpages` exposes live in-process debug pages, `file_storage` backs persistent queues, and auth extensions (`basicauth`, `oauth2client`, `headers_setter`) supply credentials that receivers and exporters reference by id. They never appear inside a pipeline; they are listed once under `service.extensions`. ## The rest of `service` `service.telemetry` configures the Collector's own logs and metrics — its self-observability, which is the first thing you need when the pipeline misbehaves. Configuration can be assembled from multiple `--config` sources and environment variable substitution (`${env:VAR}`), and validated without running traffic using the Collector's `validate` subcommand. ## The failure to avoid The most common production surprise is the silently inert pipeline: components declared, `service` not updated, telemetry accepted at the receiver and going nowhere. Check the Collector's own metrics — accepted at the receiver but zero sent at the exporter is the signature.

  • You add a filter processor to the config, restart, and nothing changes. What do you check first?
    Whether the processor id is listed in the relevant pipeline under `service.pipelines`, and in a pipeline of the right signal. Declaring a component under `processors:` only configures it; it is inert until referenced. Next check the position in the chain — a filter placed after a component that already dropped or rewrote the matching attribute will match nothing.
  • Why is `batch` recommended after sampling and filtering rather than before?
    Batching is about amortising export overhead, so it should operate on the data you actually intend to send. Batching first means building and holding batches of records that a later processor throws away, wasting memory and CPU. Some processors also need whole groupings — tail sampling buffers by trace — so putting `batch` in front of them just adds latency without helping.

saying these in an interview costs you the question

  • Thinking declaring a component under `exporters:` makes it active without touching `service`
  • Believing processor order is cosmetic
  • Assuming exporters run in sequence, so "the last one gets modified data"
  • Putting `memory_limiter` late or omitting it entirely
  • Trying to put an extension such as `health_check` inside a pipeline's processor list

context