skip to content

What does the OpenTelemetry Operator for Kubernetes actually do? Name the custom resources it introduces and explain what each one buys you compared with deploying the same thing by hand.

level: seniorimportance: should knowfreq 30%

answer

  1. CRDs: OpenTelemetryCollector, Instrumentation (+ OpAMPBridge)
  2. modes: deployment / daemonset / statefulset / sidecar
  3. Target Allocator shards scrape targets across a StatefulSet
  4. annotation inject-<lang> → webhook → init container copies agent
  5. injection at pod creation only; node-local collects, gateway tail-samples

basics

~20 s

It manages telemetry infrastructure declaratively. An OpenTelemetryCollector resource deploys and configures collectors in one of four modes — deployment, daemonset, statefulset or sidecar. An Instrumentation resource describes language agents and endpoints, and an admission webhook injects them into annotated pods via init containers, so applications get auto-instrumentation without changing their images.

solid answer

~50 s

Two custom resources carry the weight. **`OpenTelemetryCollector`** — you write the collector configuration in the spec and the operator reconciles the workload: a **Deployment** (a shared gateway), a **DaemonSet** (one per node, for node-local signals), a **StatefulSet** (when instances need stable identity, notably for scrape sharding), or a **Sidecar** injected into annotated pods. It handles the ConfigMap, service, ports derived from the receivers, and rollout on config change. It also manages the **Target Allocator**, which shards Prometheus scrape targets across a StatefulSet of collectors so scraping scales horizontally without duplicate scrapes. **`Instrumentation`** — declares, per language, which auto-instrumentation image to use plus the exporter endpoint, propagators and sampler defaults. A mutating admission webhook watches for pod annotations such as `instrumentation.opentelemetry.io/inject-java: "true"` and injects an init container that copies the agent onto a shared volume, then sets the environment so the runtime loads it. The value is that instrumentation and telemetry routing become cluster policy rather than per-image build steps.

go deeper

for a junior

Name the two main custom resources and say the operator deploys collectors and can inject agents into pods automatically.

for a middle

List the four collector modes with a use case each and describe the init-container injection mechanism triggered by pod annotations.

for a senior

Add the Target Allocator, the node-local-plus-gateway topology and why tail sampling needs the gateway, plus restart semantics and agent/runtime compatibility.

for a principal

Treat it as fleet policy: which tiers exist, webhook failure policy and blast radius, agent version rollout strategy, and the migration path from an existing scrape-based metrics stack.

## What problem the operator solves Running OpenTelemetry in Kubernetes without an operator means: build collector config into ConfigMaps by hand, keep Deployments/DaemonSets in sync with it, restart on change, and — the painful part — get a language agent into every application image or entrypoint. The operator turns all of that into declarative resources reconciled by a controller, and it does the agent delivery with an admission webhook so application images stay untouched. ## `OpenTelemetryCollector` The spec contains the collector configuration itself plus deployment intent. The **mode** is the key field: - **deployment** — a scalable, centrally reachable gateway. Applications push OTLP to its service. This is where expensive processing (tail sampling, attribute policy, routing to several backends) belongs, because it sees traffic from many pods. - **daemonset** — one instance per node. Correct for anything node-local: reading container logs off the node filesystem, host metrics, or acting as a nearby first hop that batches before a longer network trip. - **statefulset** — stable network identity per replica. The motivating case is Prometheus-style scraping, where each replica must own a deterministic slice of the targets. - **sidecar** — injected into pods annotated for it, giving each application its own local collector. Lowest-latency handoff and per-pod isolation, at the cost of one more container per pod and N copies of the config. The operator reconciles the ConfigMap, the workload, and a Service whose ports it derives from the receivers you declared. Changing the config triggers a managed rollout, which is exactly the fiddly part when hand-rolled. It can also manage autoscaling and expose the collector's own telemetry. ### Target Allocator A separate component the operator can run alongside a StatefulSet collector. It performs service discovery for Prometheus scrape configuration and then *assigns* targets to collector replicas, keeping the assignment stable as replicas and targets change. Without it, every replica scrapes everything (duplicate samples) or you shard by hand. It also supports discovering scrape configuration from Prometheus-ecosystem CRDs, which is what makes migrating an existing Prometheus-operator setup onto collectors realistic. ## `Instrumentation` This resource is the declarative form of 'how should applications in this namespace be instrumented': the auto-instrumentation image per language, the exporter endpoint, the propagator list, sampler defaults, and extra resource attributes. Delivery is by mutating admission webhook. A pod (or its namespace) carries an annotation naming the language — `instrumentation.opentelemetry.io/inject-java`, `...inject-python`, `...inject-nodejs`, `...inject-dotnet`, `...inject-go` (the Go case works differently, via eBPF, and carries stricter requirements) — plus optionally which `Instrumentation` resource to use. The webhook adds an **init container** carrying the agent, mounts a shared volume, copies the agent files onto it, and injects the environment variables that make the runtime load the agent and point it at the right endpoint. The application image never changes. The consequences are worth naming: - Instrumentation becomes **cluster policy**, changeable centrally — a new exporter endpoint or a propagator change is one resource edit plus a pod restart, not a fleet-wide rebuild. - It only applies **at pod creation**, so existing pods must be restarted to pick it up. - Agent version and application runtime version must be compatible; a pinned agent image across a heterogeneous fleet will eventually meet a runtime it does not support. - The webhook is in the pod-admission path. If it is unhealthy and configured to fail closed, pod creation stalls cluster-wide — a genuine blast-radius consideration. - Auto-injection cannot instrument what the agent does not know about, and it cannot see custom business spans; it is a floor, not a ceiling. ## `OpAMPBridge` A third resource that connects managed collectors to an OpAMP server for remote configuration and health reporting. Niche, but worth naming as evidence you have read the operator's surface rather than one blog post. ## Choosing a topology The common production shape is **daemonset or sidecar for collection, deployment for the gateway**: node-local or pod-local collectors do cheap work (receive, batch, add k8s attributes) and forward to a gateway tier that does the expensive, whole-picture work such as tail sampling — which cannot be done correctly at a node that only sees part of each trace. A StatefulSet plus Target Allocator is added when the cluster is also doing scrape-based metrics collection. ## Interview framing Name both primary CRDs, give the four collector modes with a reason for each, explain the init-container injection mechanism concretely, then spend the rest on the operational consequences: restart-to-apply, webhook blast radius, agent/runtime compatibility, and why tail sampling forces a gateway tier.

  • Why does tail sampling force a gateway collector rather than a per-node one?
    A tail sampling decision needs the whole trace, and a trace's spans are produced by pods spread across many nodes. A node-local collector sees only its fragment, so it cannot decide. The usual shape is node-local or sidecar collectors doing cheap enrichment and batching, forwarding to a gateway tier where spans of a trace converge — often with load balancing by trace id so all spans of one trace reach the same gateway instance.
  • An application was annotated for auto-instrumentation but emits nothing. What do you check?
    First whether the pod was recreated after the annotation was added — the mutating webhook only acts at admission. Then inspect the pod spec for the injected init container, the shared volume and the injected environment; if they are absent the webhook did not match the pod or the Instrumentation resource is not resolvable from that namespace. If they are present, the problem has moved downstream to agent/runtime compatibility or exporter connectivity.
  • What is the risk of the operator's mutating admission webhook in a large cluster?
    It sits in the pod-creation path. If it is unavailable and the webhook is configured to fail closed, pod creation is blocked for everything it matches, which can prevent recovery during an incident. The mitigations are the standard ones — narrow the match by namespace or label, run the webhook highly available, and make a deliberate decision about the failure policy rather than accepting the default blindly.

saying these in an interview costs you the question

  • Thinking the operator instruments already-running pods without a restart.
  • Believing auto-injection replaces manual instrumentation of business logic.
  • Running tail sampling on a DaemonSet collector that only sees part of each trace.
  • Not knowing the Target Allocator exists, so every scraping replica duplicates every target.
  • Ignoring that the injecting webhook is in the pod-admission path and has a failure policy.

context