skip to content

Compare deploying the OpenTelemetry Collector as a per-host or sidecar agent versus a central gateway cluster. What work belongs in each tier, and how do you scale the gateway tier without breaking stateful processing?

level: seniorimportance: must knowfreq 46%

answer

  1. Agent = local endpoint, host/pod enrichment, log + host metrics, PII stripping
  2. Gateway = egress, credentials, routing, whole-trace work
  3. Stateful processors need all spans of a trace on one replica
  4. loadbalancing exporter, routing_key: traceID → two-layer gateway
  5. Rescaling reshuffles the hash ring; traces in flight can split

basics

~20 s

Agents run next to workloads for local-only work: a cheap local endpoint, host and pod enrichment, log/host-metric collection. Gateways are a shared cluster for egress control, credentials, routing and whole-trace work. Scale gateways horizontally, but stateful processors need trace-affinity routing.

solid answer

~50 s

The **agent** tier is one Collector per node (DaemonSet) or per pod (sidecar). It gives applications a local, low-latency, unauthenticated OTLP endpoint, and it does the work that requires being *there*: `k8sattributes` and `resourcedetection` enrichment, scraping host metrics, tailing container logs, and stripping PII before data leaves the node. Its blast radius is one node. The **gateway** tier is a horizontally scaled deployment behind a service address: it owns egress — credentials, TLS, quotas, tenant routing, backend fan-out — plus everything requiring a fleet-wide view: tail sampling, span-derived metrics, service graphs, cardinality limiting. The scaling catch: those processors are **stateful per trace**. A plain load balancer spreads a trace's spans across replicas and every replica sees a partial trace. The fix is a two-layer gateway — a stateless front layer using the `loadbalancing` exporter with `routing_key: traceID` to hash spans onto a consistent backend replica, and a second layer that actually runs `tail_sampling` or `spanmetrics`. Small estates can run agent-only until a global function is needed.

code

text · 13 lines
text
apps ──▶ agent (DaemonSet)            enrich, small batches
              │
              ▼
        gateway L1 (stateless, N pods)
          exporter: loadbalancing
            routing_key: traceID   ── hash(trace_id) ──▶ pick L2 pod
              │
              ▼
        gateway L2 (stateful, M pods)
          processors: [tail_sampling]   sees ALL spans of a trace
              │
              ▼
           backend

go deeper

for a junior

Know the three shapes — sidecar, node agent, central gateway — and that the agent adds local metadata while the gateway talks to the backend.

for a middle

Assign work correctly: enrichment and local collection at the agent, credentials, routing and fan-out at the gateway, and explain why the agent must stay lightweight.

for a senior

Own the affinity problem: name the loadbalancing exporter with routing_key: traceID, the two-layer design, and the sizing model for the stateful tier.

for a principal

Decide when the complexity is justified at all, weigh blast radius (agent config change hits every node) against cost, and set the governance boundary for where PII stripping and spend control live.

## Three deployment shapes **Sidecar**: a Collector container in every pod. Strongest isolation and per-workload configuration, and useful when tenants must not share a telemetry process. Cost is linear in pods — memory and CPU reservations multiply, and config rollout touches every workload. **Agent (node-local)**: one Collector per host or Kubernetes node, usually a DaemonSet. This is the common default. Applications export to `localhost` or the node IP, so the hop is cheap and needs no authentication, and the failure blast radius is one node's telemetry. **Gateway**: a standalone Deployment of several replicas behind a service address, receiving from agents (or directly from workloads that have no agent, such as serverless functions). This is a shared, scalable pool. Agent and gateway are not alternatives; most real deployments run both, and the interesting question is which work lives where. ## What only the agent can do Anything that depends on being co-located. `resourcedetection` reads the host/cloud instance metadata endpoint; `k8sattributes` associates records with the pod that sent them, which is far easier and cheaper when the sender is on the same node; `hostmetrics` and `filelog` receivers read node-local resources and container log files. Redaction of personal data is also better here on a defensibility argument: scrubbed at the agent, the raw value never crosses the node boundary at all. Agents should stay cheap and stateless. Small batches, short timeouts, tight memory limits — a thousand agents each holding a large buffer is a large hidden memory bill on the serving fleet. ## What belongs in the gateway - **Egress and credentials.** One place holds the backend token, the TLS material, the proxy route. One network path to firewall and meter. - **Routing and multi-tenancy.** A `routing` connector or per-tenant pipelines send each team's data to its own destination or account. - **Fan-out and migration.** Dual-shipping to two vendors is one pipeline with two exporters, changed once. - **Whole-trace and whole-fleet processing.** Tail sampling, `spanmetrics` and `servicegraph` connectors, cardinality limiting, quota enforcement. None of these can work correctly on a partial view. - **Big buffers.** A disk-backed persistent queue on a modest number of gateway replicas is far more practical than on every node. ## Scaling the gateway — the affinity problem Gateway replicas are trivially scalable *if* every processor is stateless per record. They are not, the moment you add a processor that reasons about a whole trace. `tail_sampling`, `groupbytrace`, `spanmetrics` and `servicegraph` all accumulate state keyed by trace id. Put N replicas behind a normal round-robin load balancer and each replica sees a random subset of each trace: tail-sampling policies evaluate against fragments, service graphs miss edges, and derived metrics under-count. The standard answer is a **two-layer gateway**. Layer one is stateless and does nothing but route: its pipeline exports through the `loadbalancing` exporter configured with `routing_key: traceID`, which hashes on the trace id so all spans of a trace land on the same layer-two replica; backend membership is resolved via DNS or Kubernetes service resolution so the ring updates as replicas change. Layer two runs the stateful processors and exports to the backend. For span-derived metrics that must aggregate per service rather than per trace, the same exporter supports `routing_key: service`. Two caveats worth voicing. Rescaling layer two reshuffles the hash ring, so traces in flight during a scale event can be split — expect a small, transient sampling inaccuracy, and prefer scaling deliberately rather than aggressively autoscaling that tier. And layer two is a *stateful* tier: sized by concurrent-traces-in-flight × average trace size, not by request rate alone. ## Choosing a topology Start with agents only if all you need is enrichment and shipping; that is one fewer tier to run. Introduce a gateway when any of these becomes true: you need a single controlled egress point, you need whole-trace decisions, you need per-tenant routing, or you have workloads (functions, external partners, non-Kubernetes hosts) with nowhere to put an agent. Add the two-layer split only when a stateful processor arrives — it is real operational complexity and should be paid for by a real requirement. ## Operating both tiers Each tier needs its own capacity model and its own alerts on the Collector's internal telemetry: refused records at the receiver, queue occupancy versus capacity, failed sends at the exporter. Version them independently and roll the agent tier gradually — an agent config bug fails a thousand nodes at once, which makes the agent tier, not the gateway, the higher-risk change surface.

  • Why can't you just put the tail-sampling processor on every agent instead?
    Because an agent only ever sees the spans produced on its own node. A distributed trace crosses many services on many nodes, so each agent would hold a fragment and would have to decide on incomplete evidence — a latency or error policy would fire or not fire essentially at random. Tail sampling needs a point where a whole trace can be assembled, which is what the gateway tier provides.
  • What is the operational risk of running only sidecars and no shared tier?
    Cost and rollout risk. Every pod carries a Collector's memory and CPU reservation, which is a large multiplier at scale, and a config change touches every workload's deployment. You also still lack any place where whole-trace or cross-service aggregation can happen, and every sidecar needs backend credentials, widening the secret's exposure.
  • How do you size the stateful gateway layer?
    By concurrent traces in flight rather than by span rate alone: roughly the number of traces started per second multiplied by the decision wait window, multiplied by the average bytes per trace, plus headroom for spikes and for the sampling processor's own bookkeeping. Then verify empirically against queue occupancy and heap, because attribute-heavy spans move the average sharply.

saying these in an interview costs you the question

  • Putting tail sampling or spanmetrics behind a round-robin load balancer and expecting correct results
  • Claiming agents and gateways are alternatives rather than complementary tiers
  • Running fat buffers on every agent, multiplying memory cost across the serving fleet
  • Believing a Kubernetes Service with session affinity solves trace affinity (affinity is per client connection, not per trace)
  • Treating the gateway as the risky tier while rolling agent config to every node at once

context