skip to content

OpenTelemetry

The vendor-neutral standard for producing telemetry: one set of SDKs and one wire protocol for traces, metrics and logs, plus a collector that reshapes and routes them. Interviewers care because OTel is how you avoid rewriting instrumentation every time the observability vendor changes.

on this pageshow

questions

page 2 of 2

Telemetry spend is doubling every quarter and legal wants user email addresses removed from span attributes. How would you use OpenTelemetry Collector processors — filter, transform (OTTL), attributes — and connectors to shape volume and scrub data, and what breaks if you get it wrong?

level: principalimportance: should knowfreq 29%

basics

~20 s

Scrub PII at the agent with transform/OTTL or redaction so raw values never leave the node; shape volume centrally with filter, attribute pruning and aggregation connectors. Measure before dropping, and never break trace continuity or metric series continuity.

open as a page

You need latency percentiles across dozens of services. Compare explicit-bucket histograms, exponential (base-2) histograms and the legacy summary data point in OpenTelemetry, and explain why attribute cardinality is a far bigger problem on a metric than on a span.

level: principalimportance: should knowfreq 27%

basics

~20 s

Explicit buckets need boundaries chosen up front and only merge across services if identical. Exponential histograms derive buckets from a scale, cover any range and merge by downscaling. Summaries carry pre-computed quantiles that cannot be re-aggregated. Metric attributes multiply stored series; span attributes do not.

open as a page

You own observability across a few hundred services and must decide how deeply each is instrumented with OpenTelemetry. How do you split effort between running an auto-instrumentation agent everywhere and hand-written spans, and how do you keep the result consistent and affordable?

level: principalimportance: should knowfreq 30%

basics

~20 s

Make the agent the default everywhere: it buys uniform boundary coverage and context propagation with no per-team work. Spend scarce manual instrumentation only where domain decisions live. Govern a shared attribute vocabulary, and control cost through span granularity rather than by removing instrumentation.

open as a page

You must choose a trace sampling strategy for a fleet of services that call one another. How do you keep the decision consistent across a whole trace, how do you decide between deciding at the entry point and deciding later in a collector, and how do you set the rate?

level: principalimportance: should knowfreq 35%

basics

~20 s

Decide once at the entry point and have everyone else inherit, using parent-based sampling everywhere. Head sampling is cheap and blind to outcomes; deciding after the trace completes, in a buffering collector, can keep errors and slow traces but costs buffering and complexity. Rate follows an ingestion budget, not a round number.

open as a page

showing 31–34 of 34