skip to content

OpenTelemetry

The vendor-neutral standard for producing telemetry: one set of SDKs and one wire protocol for traces, metrics and logs, plus a collector that reshapes and routes them. Interviewers care because OTel is how you avoid rewriting instrumentation every time the observability vendor changes.

on this pageshow

questions

page 1 of 2

An application can export OTLP telemetry straight to a vendor backend. What does running an OpenTelemetry Collector in between buy you, and when is direct export from the SDK good enough?

level: juniorimportance: must knowfreq 55%

answer

  1. Receive → process → export, in a separate process
  2. Late binding: backend swap = config, not redeploy
  3. Queue and retry outside the app heap
  4. k8sattributes / resourcedetection see what the app cannot
  5. Tail sampling and span metrics need a meeting point

basics

~20 s

The Collector is a separate process that receives, processes and re-exports telemetry. It decouples apps from backends: batching, retry, enrichment, redaction, translation and routing become config in one place instead of code in every service.

solid answer

~50 s

The Collector is a standalone, vendor-neutral binary with receivers, processors and exporters. It is optional — an SDK can speak OTLP to any OTLP backend — but it buys four things. **Late binding**: the destination, credentials, redaction rules and sampling live in Collector config, so switching or dual-shipping to a second backend is a config change, not a redeploy of fifty services. **Buffering outside the app**: bounded queues, exponential-backoff retry, and optionally a disk-backed queue, so a backend outage burns Collector memory rather than application heap. **Enrichment**: processors like `k8sattributes` and `resourcedetection` attach pod, node, cloud-region and instance identity the app cannot see. **Translation and global functions**: receivers for Prometheus scrape, Jaeger, Zipkin, statsd, fluentforward and log files converge on one egress format, and whole-trace work like tail sampling or span-derived metrics is only possible where many traces meet. Direct export is fine for prototypes, single-service systems, and local development.

go deeper

for a junior

Say what it is (a separate receive-process-export process), that it is optional, and give two concrete wins: swap backends without redeploying, and enrich with Kubernetes metadata the app cannot see.

for a middle

Add the buffering/retry argument and the translation role (Prometheus, Jaeger, Zipkin, file logs in; OTLP out), and name the cost of running another process.

for a senior

Frame it as where policy lives: destination, credentials, redaction, sampling and cost shaping become one configuration surface, and whole-trace functions become possible at all.

for a principal

Argue the organisational case — a Collector tier is the control point for telemetry spend, PII egress and vendor lock-in — and state the availability contract: telemetry failure must never degrade the serving path.

## What the Collector actually is The OpenTelemetry Collector is an ordinary process (a Go binary, usually shipped as the `otelcol` core distribution, the larger `otelcol-contrib`, or a custom build produced by the OpenTelemetry Collector Builder). It receives telemetry over the network or from files, runs it through processors, and exports it onward. It is **not** part of the protocol: an SDK that speaks OTLP can talk directly to any backend that accepts OTLP. So the question is always "what does this extra hop earn?" ## Job 1 — decoupling the app from the backend Without a Collector, every service embeds the destination endpoint, the vendor API key, the retry policy and any redaction logic. Changing vendors, adding a second destination during a migration, or lowering export volume means rebuilding and redeploying every service. With a Collector, applications export to one local, unauthenticated endpoint (typically OTLP/gRPC on `4317` or OTLP/HTTP on `4318` on localhost or a cluster service), and everything downstream is Collector configuration. Dual-shipping to old and new backends for a migration window is two exporters in one pipeline. It also contains credentials: the backend token lives in one deployment's environment, not in every service image. ## Job 2 — buffering and retry outside the application process SDK exporters have small in-process queues by design; they must never consume the heap that serves user traffic, so under a backend outage they drop. A Collector can hold a much larger bounded queue, retry with exponential backoff, and — with the `file_storage` extension enabling a persistent sending queue — survive its own restart with data on disk. The failure mode moves out of the application, which is the point: telemetry loss should never become an application incident. ## Job 3 — enrichment the application cannot do A process does not know its Kubernetes pod name, namespace, node, deployment, or its cloud region and instance id — and usually should not hold the API credentials required to look them up. A node-local Collector can, and processors such as `k8sattributes` and `resourcedetection` attach those as resource attributes to every record passing through. This is why the agent tier exists even in shops that also run a gateway. ## Job 4 — translation and whole-fleet functions Receivers exist for Prometheus scrape, Jaeger, Zipkin, statsd, fluentforward, host metrics, container logs and Kafka. That lets a Collector be the bridge from a legacy telemetry estate into OTLP without touching legacy apps. And some processing is *only* possible where many streams converge: tail-based sampling needs every span of a trace, span-derived RED metrics need aggregation across requests, cardinality limiting needs a view of the whole series population. No single application process can do these. ## The costs, stated honestly It is one more thing to deploy, version, monitor, capacity-plan and debug. It has its own failure modes — out-of-memory under a traffic spike, a full sending queue silently dropping, a misconfigured pipeline that starts but exports nowhere. It adds latency to *telemetry delivery*, though not to user requests, since export is asynchronous. And it cannot fix what was never captured: missing instrumentation, a wrong `service.name` at the source, or a head-sampling decision already made in the SDK that discarded the span before it ever left the process. ## When direct export is fine A prototype, a local dev loop, a single service, or a small system with one backend and no compliance requirement on egress. Short-lived serverless functions are a special case: there may be no node-local agent to send to and no time to drain a queue, so many teams point them at a gateway Collector rather than an agent, or export directly and accept the coupling. ## Availability Because the Collector sits in the path of telemetry only, its outage must degrade observability, never the application. SDK exporters are expected to fail quietly and keep serving. Reduce blast radius by running the agent tier per node (one node's telemetry at risk) and replicating the gateway tier behind a load balancer, and watch the Collector's own internal metrics rather than assuming silence means health.

  • If the Collector goes down, what happens to the application?
    Nothing user-visible should happen. SDK exporters are required to fail without propagating errors into application code; they log, retry within a bounded queue, and drop when it fills. You lose telemetry for the outage window, which is why the agent tier is deployed per node to limit blast radius and the gateway tier is replicated. A disk-backed persistent queue can preserve data across a Collector restart, but nothing preserves data the SDK already dropped.
  • Name something a Collector cannot fix that you must get right in the application.
    Anything decided before export: instrumentation that was never added, a head-sampling decision that discarded a trace at creation time, span names and attributes with the wrong semantics, and the identity of the service itself. You can rewrite `service.name` in a transform processor, but then two teams' config disagree about who owns the data — better to set it correctly at the source.

A mail room. Every desk could stamp and post its own parcels, but you put a mail room in the middle so postage accounts, address changes, redaction of sensitive contents and courier switches happen once, not at every desk.

saying these in an interview costs you the question

  • Believing OTLP requires a Collector, or that the Collector is part of the SDK
  • Claiming the Collector instruments the application — it never sees code, only exported records
  • Assuming a Collector guarantees no data loss; its queue is bounded and drops when full unless disk-backed
  • Treating the Collector as a storage backend rather than a pipeline
  • Saying it removes all latency concerns — export is async either way; the win is operational, not latency

context

open as a page

What does the W3C `traceparent` HTTP header carry, field by field, and what should a receiving service do with its sampled flag and with the accompanying `tracestate` header?

level: juniorimportance: must knowfreq 68%

basics

~20 s

traceparent is four hyphen-separated fields: version, 32-hex trace id, 16-hex span id of the caller, and 2-hex flags whose lowest bit means sampled. The receiver continues the same trace id, parents its span on that span id, and forwards tracestate unchanged apart from its own entry.

open as a page

What fields make up an OpenTelemetry span, and what do its `kind` and `status` fields actually mean to a backend that receives it?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A span carries identity (trace id, span id, parent span id), a name, a kind, start and end timestamps, attributes, events, links and a status. Kind tells the backend the span's role in a call (server, client, producer, consumer, internal); status is Unset, Ok or Error.

open as a page

OpenTelemetry offers a zero-code (agent-based) way to instrument an application and a code-based way using its API. What does each actually capture, what can neither capture, and how do the two combine inside one process?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Zero-code agents patch libraries at load time and emit spans for inbound requests, outbound calls, database and messaging clients with no source changes. Code-based instrumentation adds spans and attributes only your code knows. Both feed one SDK and nest.

open as a page

Your services are instrumented with the OpenTelemetry SDKs and the telemetry has to reach an observability backend. Make the case for exporting in OTLP (the OpenTelemetry Protocol) rather than using a backend-specific exporter inside the process, say when a native exporter is still the right call, and explain how you would choose between OTLP's gRPC and HTTP transports.

level: middleimportance: must knowfreq 55%

basics

~20 s

OTLP is the one exporter every OpenTelemetry SDK ships and it carries traces, metrics and logs, so the destination stays a config change. A native exporter compiles a vendor's protocol into the process, making a backend switch a redeploy. Use gRPC in-cluster, HTTP where proxies or browsers break it.

open as a page

Walk through how an OpenTelemetry Collector configuration is structured — receivers, processors, exporters, connectors, extensions and the `service` section — and explain what determines the order data flows through it.

level: middleimportance: must knowfreq 56%

basics

~20 s

Top-level blocks declare components by id; the service.pipelines block wires them into per-signal pipelines. Data enters a receiver, passes processors in the declared order, then fans out to all exporters in parallel. Declared but unreferenced components are never started.

open as a page

What is Baggage in OpenTelemetry, how does it differ from span attributes, and what risks come with adding fields to it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Baggage is a set of key/value pairs carried in the Context and propagated to every downstream service in the baggage header. Unlike span attributes, which stay on one local span, baggage travels everywhere — so it costs bytes on every hop, is untrusted when inbound, and is not automatically recorded on spans.

open as a page

OpenTelemetry metrics offer counters, up-down counters, gauges and histograms, each in a synchronous and an asynchronous (observable) form. How do you choose between them, and what does "asynchronous" actually change?

level: middleimportance: must knowfreq 46%

basics

~20 s

Counter for monotonic totals, up-down counter for values that rise and fall, histogram for distributions you want percentiles from, gauge for a current value. Synchronous instruments are called where the event happens; asynchronous ones are read by a callback at collection time.

open as a page

You are adding hand-written OpenTelemetry spans to a service. What must be true at span creation, at activation and at end for the span to be sampled, parented and exported correctly — and what makes a hand-written span useless even when it is technically valid?

level: middleimportance: must knowfreq 52%

basics

~20 s

At creation the sampler sees only parent context, trace id, name, kind, links and the attributes passed there. Activation is separate from starting — an unactivated span orphans nested work. End on every path or nothing exports. Useless spans duplicate auto-instrumentation or carry no decision-changing detail.

open as a page

In an OpenTelemetry SDK, what is a Resource, why does telemetry show up under a name like 'unknown_service' when it is not configured, and how do OTEL_SERVICE_NAME, OTEL_RESOURCE_ATTRIBUTES and resource detectors interact?

level: middleimportance: must knowfreq 55%

basics

~20 s

A Resource is the immutable set of attributes identifying the entity producing telemetry — service name, version, instance, host, container. It is set once at SDK startup, attached to every signal. With no service name configured the SDK falls back to a placeholder like unknown_service.

open as a page

Explain the built-in trace samplers in the OpenTelemetry SDK — AlwaysOn, AlwaysOff, TraceIdRatioBased and ParentBased. Which is the default, what information does a sampler see when it runs, and what do the OTEL_TRACES_SAMPLER and OTEL_TRACES_SAMPLER_ARG environment variables set?

level: middleimportance: must knowfreq 52%

basics

~20 s

AlwaysOn and AlwaysOff record all or nothing. TraceIdRatioBased keeps a fraction derived from the trace id, so every service decides alike. ParentBased follows the incoming decision and delegates only at the root; the default is ParentBased wrapping AlwaysOn. OTEL_TRACES_SAMPLER picks one, OTEL_TRACES_SAMPLER_ARG parameterises it.

open as a page

You need OpenTelemetry metrics to land in a Prometheus-compatible metrics store. Walk through the push-versus-pull options, and the two data-model mismatches — aggregation temporality and metric naming — that you have to resolve on the way.

level: seniorimportance: must knowfreq 45%

basics

~20 s

Either expose a scrape endpoint (a Prometheus exporter that Prometheus pulls) or push — OTLP into Prometheus's OTLP receive endpoint, or remote-write from a collector. Prometheus needs cumulative counters, so delta temporality must be converted; and OTel metric names must be normalised (dots to underscores, unit suffixes, _total on monotonic sums), while resource attributes land on a separate target_info series.

open as a page

Compare deploying the OpenTelemetry Collector as a per-host or sidecar agent versus a central gateway cluster. What work belongs in each tier, and how do you scale the gateway tier without breaking stateful processing?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Agents run next to workloads for local-only work: a cheap local endpoint, host and pod enrichment, log/host-metric collection. Gateways are a shared cluster for egress control, credentials, routing and whole-trace work. Scale gateways horizontally, but stateful processors need trace-affinity routing.

open as a page

Explain how tail-based sampling works in the OpenTelemetry Collector — the decision buffer, policies, and late-arriving spans — and what it costs compared with deciding at span creation time.

level: seniorimportance: must knowfreq 44%

basics

~20 s

The tail_sampling processor buffers spans by trace id for a wait window, then evaluates policies (latency, error status, attribute match, probabilistic) against the assembled trace and keeps or drops the whole trace. It costs memory and requires trace affinity.

open as a page

OpenTelemetry keeps a "current context" in an implicit, scope-based store within a process. What exactly breaks when work is handed to a thread pool, a callback, or a reactive/coroutine pipeline, and how is it fixed?

level: seniorimportance: must knowfreq 58%

basics

~20 s

The implicit store is per-thread (or per-async-task), so work that moves to another thread sees an empty or stale context. Spans created there become new roots or attach to the wrong parent. Fix by capturing the Context at hand-off and re-attaching it inside the task, or by passing context explicitly.

open as a page

Explain delta versus cumulative aggregation temporality in OpenTelemetry metrics: what each data point carries, how a consumer detects a counter reset, and how the choice interacts with a scrape-based versus a push-based backend.

level: seniorimportance: must knowfreq 36%

basics

~20 s

Cumulative points report the running total since a start time; delta points report only what happened in the interval since the previous point. Consumers detect resets from a decreased value or a changed start timestamp. Scrape-based backends want cumulative; many push backends prefer delta.

open as a page

An application is running with the OpenTelemetry SDK enabled, but no traces reach the backend and nothing obvious is failing. How do you bisect the SDK's configuration to find where the data is being lost?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Walk the path in order: is the SDK actually installed and enabled, is anything creating spans, did the sampler keep them, did the processor queue and not drop them, did the exporter succeed, and is the data filed under a service name and destination you are searching. Turn on SDK self-diagnostics first.

open as a page

Compare the simple and batching span processors in the OpenTelemetry SDK. What do the queue size, schedule delay, maximum export batch size and export timeout settings control, what happens when the queue fills, and what must a short-lived process do before exiting?

level: seniorimportance: must knowfreq 45%

basics

~20 s

The simple processor exports each span synchronously as it ends — correct but slow, for debugging. The batching processor enqueues ended spans and a background thread exports batches on a timer or when full; a full queue drops spans silently. Short-lived processes must force-flush and shut down or lose buffered spans.

open as a page

Describe the inject/extract carrier model OpenTelemetry uses to move context across a process boundary, and how you would apply it to a transport that is not HTTP, such as a message broker's record headers.

level: middleimportance: should knowfreq 48%

basics

~20 s

A text-map propagator exposes inject(context, carrier, setter) and extract(context, carrier, getter). The setter and getter abstract the carrier, so the same propagator writes into HTTP headers, gRPC metadata or broker record headers. For messaging you inject when producing and extract when consuming.

open as a page

OpenTelemetry ships a tracing API artifact separately from its SDK. Why is the split there, what should a reusable library depend on, what happens at runtime when no SDK is installed, and what is the tracer's name and version used for?

level: middleimportance: should knowfreq 45%

basics

~20 s

The API is the calling surface; the SDK is the implementation that samples, processes and exports. Libraries depend on the API only, so with no SDK present every call is a cheap no-op. The tracer's name and version identify the instrumentation scope on each span.

open as a page

You are sending OpenTelemetry log records to a log backend such as Loki. Explain how the OTel log record maps onto that backend's storage model, what makes logs correlate with traces, and where the cardinality trap is.

level: seniorimportance: should knowfreq 26%

basics

~20 s

An OTel log record has a body, a severity, resource and record attributes, and optional trace and span ids. A label-indexed store like Loki can only afford a handful of low-cardinality labels, so a small subset of resource attributes becomes stream labels and everything else must ride as non-indexed structured metadata or in the body. Correlation works because the trace and span ids travel on the record itself.

open as a page

What does the OpenTelemetry Operator for Kubernetes actually do? Name the custom resources it introduces and explain what each one buys you compared with deploying the same thing by hand.

level: seniorimportance: should knowfreq 30%

basics

~20 s

It manages telemetry infrastructure declaratively. An OpenTelemetryCollector resource deploys and configures collectors in one of four modes — deployment, daemonset, statefulset or sidecar. An Instrumentation resource describes language agents and endpoints, and an admission webhook injects them into annotated pods via init containers, so applications get auto-instrumentation without changing their images.

open as a page

An OpenTelemetry Collector is dropping telemetry under load and occasionally restarting after running out of memory. Explain the roles of the memory_limiter processor, the batch processor and the exporter's sending queue and retry, and how you would locate where data is being lost.

level: seniorimportance: should knowfreq 40%

basics

~20 s

memory_limiter refuses new data when heap crosses a soft limit so the process survives; batch amortises export cost; the exporter's sending_queue plus retry absorb backend outages and drop when full. Locate loss with the Collector's own receiver/processor/exporter counters.

open as a page

Downstream services are producing spans, but instead of continuing the caller's trace they start new traces of their own. Walk through how you would diagnose a cross-service context-propagation failure.

level: seniorimportance: should knowfreq 54%

basics

~20 s

Prove first whether the context header arrives at the callee. If it does not, something between them stripped it or the caller never injected. If it does, the callee's propagator format or extraction is wrong. Distinguish new trace ids (propagation failure) from missing spans (sampling).

open as a page

Describe the OpenTelemetry log record as a data structure: what fields it carries, why it has two timestamps, and how its 1-24 severity numbering is meant to be used.

level: seniorimportance: should knowfreq 34%

basics

~20 s

It carries an event timestamp (which may be absent), an always-present observed timestamp, a numeric severity 1-24 in six bands of four plus the original severity text, an any-typed body, attributes, an event name, and trace id, span id and trace flags fields.

open as a page

OpenTelemetry calls its Logs API a bridge API rather than something application developers should call directly. What does that mean in practice, how does a log record acquire a trace id, and what are the ways of getting an existing application's logs into OTLP?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Applications keep using their existing logging library; an OpenTelemetry appender bridges records into a LoggerProvider, which processes and exports them as OTLP. Records emitted while a span is active pick up its trace and span ids automatically, correlating logs with traces.

open as a page

What are OpenTelemetry's semantic conventions, why are they versioned and stabilised separately from the SDKs, and how would you migrate a fleet across a breaking attribute rename such as http.method becoming http.request.method?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Semantic conventions are the agreed names and value shapes for attributes — the reason telemetry from unrelated libraries is queryable together. They version independently, stabilising area by area. Breaking renames are migrated with a dual-emit opt-in: emit old and new, cut dashboards over, then drop the old.

open as a page

How do you configure an OTLP exporter in an OpenTelemetry SDK — gRPC versus HTTP transport and their conventional ports, the endpoint environment variables and their differing path rules, headers, compression and timeout — and what does a 'partial success' response from the receiver mean?

level: seniorimportance: should knowfreq 42%

basics

~20 s

OTLP runs over gRPC (conventionally port 4317) or HTTP with protobuf or JSON payloads (port 4318). A generic endpoint variable has the per-signal path appended; a per-signal endpoint variable is used exactly as given. Headers carry auth, gzip compresses, and partial success means some items were rejected while the request itself succeeded.

open as a page

Your organisation runs a commercial APM vendor's proprietary agent across its services and is considering moving instrumentation to OpenTelemetry. Argue both sides honestly, and describe how you would actually run the migration.

level: principalimportance: should knowfreq 34%

basics

~20 s

OpenTelemetry buys instrumentation you own and can point anywhere, with a shared semantic vocabulary; a proprietary agent buys depth, defaults and support that the vendor tunes end to end. Migrate incrementally per service through a collector that can fan out to both backends, with an agreed dual-run window, a comparison method, and a date to stop paying twice.

open as a page

showing 1–30 of 34