skip to content

An application is running with the OpenTelemetry SDK enabled, but no traces reach the backend and nothing obvious is failing. How do you bisect the SDK's configuration to find where the data is being lost?

level: seniorimportance: must knowfreq 48%

answer

  1. console exporter = fastest bisection
  2. no SDK ⇒ non-recording spans, no error
  3. inherited not-sampled looks like breakage
  4. 404 = per-signal endpoint missing /v1/traces
  5. check unknown_service and the time window

basics

~20 s

Walk the path in order: is the SDK actually installed and enabled, is anything creating spans, did the sampler keep them, did the processor queue and not drop them, did the exporter succeed, and is the data filed under a service name and destination you are searching. Turn on SDK self-diagnostics first.

solid answer

~50 s

Bisect the pipeline rather than guessing. **Creation**: is a real SDK installed, or is the API returning no-op spans because nothing registered a provider — and is the global disable flag set? **Sampling**: an always-off sampler, a tiny ratio, or an inherited not-sampled decision from upstream produces perfectly healthy code that exports nothing; temporarily force always-on. **Processing**: check the SDK's dropped-span counter — a full queue discards silently. **Export**: enable SDK internal logging and look for connection refused, TLS failures, 401/403, or the classic 404 from a per-signal endpoint missing its path; also check for partial-success rejections, which are loss that looks like success. **Attribution**: data may be arriving under `unknown_service`, a different environment, or a tenant you are not querying. The fastest discriminator is swapping in a console exporter: if spans print, the problem is downstream of the SDK; if not, it is creation or sampling.

go deeper

for a junior

Show a checklist mindset: is the SDK really installed, is sampling on, is the endpoint right, and does the service name match what you are searching.

for a middle

Order the stages and name the distinctive symptom of each, especially the no-op provider and the 404 endpoint path bug.

for a senior

Lead with SDK self-diagnostics and the console-exporter bisection, distinguish absent traces from holed traces, and extend the same method to the collector hop.

for a principal

Turn the post-mortem into monitoring: export failure rate, dropped spans, rejected items and a per-service reporting signal, so the next occurrence is detected rather than reported.

## Bisect, do not guess The path is: SDK installed → spans created → sampler keeps them → processor queues them → exporter ships them → backend stores them → your query finds them. Each stage has a distinct symptom. Test the middle first, then halve again. ## Stage 0 — turn the lights on Before anything else, enable the SDK's **internal diagnostics** (each SDK has its own switch for internal logging level) and, if available, its self-metrics. The SDK is usually already telling you the answer — export failures, dropped span counts, rejected items — and is being ignored because the output is off by default. The single fastest bisection is to swap in a **console/logging exporter** (or add one). Spans printed to stdout prove creation, sampling and processing are all fine, and move the whole investigation to the network and the backend. ## Stage 1 — is anything being created? Two classic causes of a silent nothing: - **No SDK is registered.** The API's default provider returns non-recording spans: code runs, no error, no data. Common when the agent was not actually attached, the dependency shipped only the API, or instrumentation grabbed a tracer before registration. - **The SDK is globally disabled** (`OTEL_SDK_DISABLED=true`) or the traces exporter is set to `none` (`OTEL_TRACES_EXPORTER=none`) — both legitimate settings that someone left in a base image or a shared config map. Also check the obvious: is the code path you are exercising instrumented at all, and did a framework upgrade silently move past the agent's supported version range? ## Stage 2 — did the sampler keep them? Sampling produces *complete absence*, not partial data, in two ways: an always-off sampler or a ratio so low that your handful of test requests all lost; or a parent-based sampler that inherited a **not-sampled** decision from an upstream service — so a service is behaving exactly as configured while looking broken. Temporarily forcing always-on and re-testing settles it in one deploy. If traces appear with *holes* rather than not at all, sampling is not your problem — look at processing, export, or an uninstrumented hop. ## Stage 3 — did the processor drop them? The batching processor discards spans when its queue is full, with no application-visible error. Its dropped-span metric is the evidence. A full queue usually means the exporter is slow or failing, so a drop count is often a symptom of stage 4 rather than a cause. Special case: a short-lived process that exits before its queue is flushed loses everything, intermittently and confusingly. ## Stage 4 — did the export succeed? With internal logging on, the errors are usually explicit: - **Connection refused / DNS failure** — wrong endpoint host, or the collector is not where you think. - **404** — a per-signal endpoint variable set without its `/v1/traces` path, or a path doubled onto the generic variable. - **Protocol mismatch** — gRPC configuration pointed at the HTTP port or vice versa. - **TLS errors** — untrusted certificate, missing client certificate for mTLS. - **401/403** — missing or wrong credentials in the headers variable. - **Timeouts** — reachable but too slow, which also drives stage-3 drops. - **Partial success with a rejected count** — accepted request, rejected data, never retried, and invisible unless you read the log. ## Stage 5 — is it there but not where you are looking? The data may be arriving and your query missing it: - Filed under a placeholder such as `unknown_service` because no service name was configured — so every service-name filter finds nothing. - A different environment or namespace attribute than the one your dashboard filters on. - Sent to a different tenant, project or endpoint than the one you are viewing. - Clock skew placing spans outside the time window you are searching. If a collector sits in between, repeat the same bisection one hop later: the collector has its own receivers, processors, exporters and its own self-telemetry, and it can drop or filter data that the SDK exported perfectly. ## Closing the loop Every stage above that failed silently should end up with monitoring: export failure rate, dropped spans, rejected items, and a coverage signal for services that stop reporting. The reason this investigation is hard the first time is that each stage was designed to fail quietly so telemetry never harms the application — which is correct, and is exactly why the SDK's own telemetry has to be watched.

  • Traces appear but each one is missing whole services. Does that point at sampling?
    Rarely. A parent-based sampling decision is inherited, so a sampled trace is normally sampled everywhere — absence would be all-or-nothing. Holes point instead at an uninstrumented or non-exporting hop, dropped context between services, or that hop's exporter failing. Check whether the missing service produces spans at all before touching sampling.
  • You add a collector between the applications and the backend and data stops arriving. How does that change the investigation?
    You now have two pipelines and repeat the same bisection at the second hop. Confirm the SDK exports successfully (its own logs), then read the collector's self-telemetry: are its receivers accepting, are processors filtering or dropping, are its exporters failing or being throttled. A collector that accepts data and fails to forward it looks identical from the application side to a healthy pipeline.

saying these in an interview costs you the question

  • Changing several settings at once instead of bisecting the pipeline
  • Never enabling SDK internal diagnostics, then calling the failure silent
  • Forgetting that an inherited not-sampled decision makes a correctly configured service emit nothing
  • Assuming a successful export means the data was stored — partial success rejects items without retry
  • Not checking whether the data landed under unknown_service or outside the queried time window

context