skip to content

Observation API & Distributed Tracing

The Observation API and distributed tracing: one instrumentation producing metrics, traces and logs, handlers and conventions, the OpenTelemetry and Brave bridges, context propagation and baggage, and sampling. Interviewers ask how you debug a slow request that crosses five services.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

25

What is Micrometer Tracing, and what does a 'tracing bridge' (bridge-otel / bridge-brave) actually do?

level: juniorimportance: must knowfreq 62%

answer

  1. facade + one bridge JAR = tracer choice
  2. bridge-otel = OpenTelemetry, bridge-brave = Brave/Zipkin
  3. replaced Spring Cloud Sleuth in Boot 3
  4. bridge != exporter
  5. code imports only Micrometer types

basics

~20 s

Micrometer Tracing is a vendor-neutral facade for distributed tracing. A bridge plugs a real tracer (OpenTelemetry or Brave/Zipkin) behind that facade, so your code uses one API while the bridge does the actual span work.

solid answer

~40 s

Micrometer Tracing is the tracing facade that replaced Spring Cloud Sleuth in Spring Boot 3. Your code depends only on its neutral API — the Tracer, Span, and SpanInScope abstractions. It does no tracing itself; it delegates to a concrete tracer chosen by which bridge JAR is on the classpath. Add micrometer-tracing-bridge-otel and you get OpenTelemetry as the SDK; add micrometer-tracing-bridge-brave and you get Brave (the OpenZipkin library). You pick exactly one bridge. The bridge translates Micrometer's Span/TraceContext calls into that library's objects, and a separate exporter dependency ships the finished spans to a backend such as Zipkin or an OTLP collector. This lets you swap tracer implementations without touching application code.

code

java · 23 lines
java
// build.gradle — facade + ONE bridge + ONE exporter
// implementation 'org.springframework.boot:spring-boot-starter-actuator'
// implementation 'io.micrometer:micrometer-tracing-bridge-otel'   // choose OTel...
// implementation 'io.opentelemetry:opentelemetry-exporter-otlp'   // ...and OTLP export

// Application code depends ONLY on the Micrometer facade:
import io.micrometer.tracing.Tracer;
import io.micrometer.tracing.Span;

@Service
class PricingService {
    private final Tracer tracer; // injected; impl comes from the bridge
    PricingService(Tracer tracer) { this.tracer = tracer; }

    void price() {
        Span span = tracer.nextSpan().name("price");
        try (Tracer.SpanInScope scope = tracer.withSpan(span.start())) {
            span.tag("tier", "gold");
        } finally {
            span.end();
        }
    }
}

go deeper

for a junior

Know it's a facade + a bridge picks the real tracer (OTel or Brave); replaced Sleuth.

for a middle

Distinguish bridge from exporter; know the two bridge artifacts and matching exporters.

for a senior

Explain Observation→span integration and that only sampled spans export while ids always propagate.

for a principal

Reason about migration from Sleuth/Brave to OTel, standardization on OTLP, and ecosystem trade-offs.

## The problem it solves **Distributed tracing** records the path of one request as it hops across services. Each unit of work is a **span** (an operation with a start time, duration, name, tags, and events); spans sharing a **trace id** form one **trace**. To make this work you need a library that creates spans, propagates ids across network calls, samples, and exports to a backend. Historically Spring used **Spring Cloud Sleuth**, which was tightly coupled to Brave. In **Spring Boot 3+** that was replaced by **Micrometer Tracing**. ## What Micrometer Tracing is Micrometer Tracing is a **facade** — a thin, vendor-neutral API. Its key abstractions: - `io.micrometer.tracing.Tracer` — creates/starts spans and exposes the current span. - `io.micrometer.tracing.Span` — one operation; you `tag(...)`, add `event(...)`, and `end()` it. - `Tracer.SpanInScope` — an `AutoCloseable` that marks a span as *current* on the thread. - `TraceContext` / `Baggage` — the propagated ids and custom key/value context. Critically, the facade contains **no tracing logic**. It needs a concrete implementation supplied at runtime. ## What a bridge does A **bridge** is the adapter that connects the Micrometer facade to a real tracer library. There are two official bridges, and you add **exactly one**: - `io.micrometer:micrometer-tracing-bridge-otel` → backs the facade with the **OpenTelemetry** SDK. Micrometer `Span` calls become OpenTelemetry `io.opentelemetry.api.trace.Span` calls. - `io.micrometer:micrometer-tracing-bridge-brave` → backs the facade with **Brave** (the OpenZipkin tracer). Micrometer calls become `brave.Span` calls. The bridge is *the choice of tracer implementation*. Your application code never imports OpenTelemetry or Brave types — only Micrometer's — so switching bridges is a dependency change, not a code change. ## The bridge is not the exporter Creating spans and **exporting** them are separate concerns. The bridge produces spans in the chosen library's format; a separate **exporter/reporter** dependency ships them: - Brave bridge → `io.zipkin.reporter2:zipkin-reporter-brave` sends to **Zipkin**. - OTel bridge → `io.opentelemetry:opentelemetry-exporter-zipkin` (to Zipkin) or `opentelemetry-exporter-otlp` (to an **OTLP** collector). Spring Boot auto-configures the exporter when its endpoint property (`management.zipkin.tracing.endpoint` or `management.otlp.tracing.endpoint`) is present. ## Auto-configuration in Spring Boot With `spring-boot-starter-actuator` plus a bridge plus an exporter, Boot's `TracingAutoConfiguration` and bridge-specific auto-config wire a `Tracer`, propagators, and the exporter. Tracing also integrates with **Micrometer Observation**: every `Observation` (HTTP server/client, `@Observed`, scheduled tasks) automatically opens and closes a span via `DefaultTracingObservationHandler`, so you often get traces without writing span code at all. ## When to use which - **OTel bridge** is the modern default: OpenTelemetry is the CNCF standard, supports OTLP natively, and has the broadest ecosystem. - **Brave bridge** suits shops already standardized on Zipkin/Brave or needing a specific Brave feature. ## Gotchas - Adding **both** bridges is a misconfiguration — pick one. - A bridge with **no exporter** still creates spans and propagates context but sends nothing to a backend. - Only **sampled** spans are exported; ids are still propagated for unsampled traces.

  • What replaced Spring Cloud Sleuth, and why the change?
    Micrometer Tracing replaced Sleuth in Spring Boot 3. Sleuth was Brave-coupled and tied to the Spring Cloud release train; Micrometer Tracing is a neutral facade with swappable bridges (OTel or Brave), aligning tracing with the Micrometer Observation API.
  • If you add a bridge but no exporter dependency, what happens?
    Spans are still created and trace context is still propagated across calls (and shows up in MDC/logs), but nothing is shipped to a backend like Zipkin — you get correlation ids in logs but no trace UI.

saying these in an interview costs you the question

  • Thinking Micrometer Tracing itself creates/exports spans without a bridge
  • Adding both bridge-otel and bridge-brave at once
  • Believing the bridge also exports (confusing bridge with the exporter dependency)
  • Claiming Spring Cloud Sleuth is still the Boot 3 tracing library

context

open as a page

What does the @Observed annotation do in Spring, and what bean must you register for it to work?

level: juniorimportance: must knowfreq 45%

basics

~20 s

@Observed marks a method so Micrometer creates an Observation around each call (producing metrics and/or a trace span). For it to work you must register an ObservedAspect bean, which uses Spring AOP to wrap the annotated method.

open as a page

What is Micrometer's Observation API, and how do you record a piece of work using an ObservationRegistry?

level: juniorimportance: must knowfreq 55%

basics

~20 s

The Observation API lets you wrap a piece of code once so Spring can produce monitoring signals from it. You create an Observation on an ObservationRegistry with a name and call observe(...) around your code.

open as a page

What is trace context propagation, and what does the W3C traceparent header carry across an HTTP call?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Context propagation passes the current trace's IDs to the next service so their spans join one trace. With Micrometer Tracing, Spring auto-adds a W3C traceparent header carrying the trace ID, the caller's span ID, and a sampled flag.

open as a page

What does the management.tracing.sampling.probability property do in Spring Boot, and what is its default value?

level: juniorimportance: must knowfreq 70%

basics

~10 s

It sets the fraction of traces that get recorded and exported, from 0.0 (none) to 1.0 (all). The default in Spring Boot is 0.1, meaning about 10% of traces are sampled.

open as a page

What is an ObservationHandler and what do its onStart, onStop, and supportsContext callbacks do?

level: middleimportance: must knowfreq 40%

basics

~20 s

An ObservationHandler is a listener that reacts to an Observation's lifecycle. onStart runs when the observation starts, onStop when it stops (e.g. to record timing), and supportsContext(context) decides whether this handler should handle a given observation at all.

open as a page

Explain the 'one instrumentation, many signals' model: how does a single Observation produce metrics, traces, and logs together?

level: middleimportance: must knowfreq 60%

basics

~20 s

You instrument code once with an Observation. The registry's handlers each react to its lifecycle events — one handler records a metric, another creates a trace span, another logs — so one instrumentation feeds all three pillars.

open as a page

What is Baggage in Micrometer Tracing, and how do you propagate a business field (e.g. tenantId) and get it into logs?

level: middleimportance: must knowfreq 60%

basics

~20 s

Baggage is arbitrary key/value data attached to the trace context and carried to downstream services alongside the trace IDs. Declare fields with management.tracing.baggage.remote-fields to propagate them, and correlation.fields to copy them into the logging MDC.

open as a page

Explain head-based sampling and how the sampling decision is propagated between services via the sampled flag.

level: middleimportance: must knowfreq 60%

basics

~20 s

The sample-or-not decision is made once at the start (head) of a trace and stored on the context. It's carried to downstream services in the trace-context headers (e.g. the W3C traceparent flags byte), so every service in the trace makes the same keep/drop choice.

open as a page

Walk through the manual span/scope lifecycle with the Micrometer Tracer API. Why must you close the scope AND end the span, and what breaks if you don't?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Create a span (tracer.nextSpan().name(...)), start it, open a scope with tracer.withSpan(span) so it becomes 'current', do work, then close the scope (try-with-resources) and call span.end(). Skipping end() loses timing/export; leaving a scope open leaks context onto later work on that thread.

open as a page

What are KeyValues on an Observation, and what is the difference between low-cardinality and high-cardinality KeyValues?

level: seniorimportance: must knowfreq 58%

basics

~20 s

KeyValues are key/value tags you attach to an Observation. Low-cardinality ones (few distinct values, like status) become metric tags and span tags. High-cardinality ones (many values, like user ID) go only on the span, not on metrics.

open as a page

How do you choose between the OpenTelemetry bridge and the Brave bridge, and what dependencies pair with each for Zipkin vs OTLP export?

level: middleimportance: should knowfreq 48%

basics

~10 s

Pick one bridge. OTel bridge (micrometer-tracing-bridge-otel) is the modern default and exports via OTLP or Zipkin. Brave bridge (micrometer-tracing-bridge-brave) exports to Zipkin via zipkin-reporter-brave. Never add both.

open as a page

How do you stop certain observations (like actuator/health traffic) from being recorded using ObservationPredicate?

level: middleimportance: should knowfreq 30%

basics

~20 s

Register an ObservationPredicate on the ObservationRegistry. It's a function of (name, context) returning true/false; if it returns false the observation becomes a no-op and nothing (no metric, no span) is recorded. Use it to filter out noise like /actuator paths.

open as a page

How do W3C and B3 propagation differ, and how do you configure which one Spring uses?

level: middleimportance: should knowfreq 55%

basics

~20 s

Both encode the same trace/span IDs but in different headers. W3C uses one traceparent header (the Boot 3 default); B3 (from Zipkin/Brave) uses either a single b3 header or multiple X-B3-* headers. Set management.tracing.propagation.type to W3C or B3.

open as a page

How does Micrometer Tracing hook into the Observation API so Spring MVC/WebClient calls get spans automatically, and how does context propagate to a downstream service?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Spring instruments HTTP and other work as Observations. Tracing registers ObservationHandlers that turn each Observation into a span. For outbound calls a propagating handler injects trace headers; on the server side another handler extracts them, so the downstream span joins the same trace.

open as a page

What is an ObservationConvention, and what is the difference between low-cardinality and high-cardinality KeyValues?

level: seniorimportance: should knowfreq 35%

basics

~20 s

An ObservationConvention centralizes how an observation is named and tagged. It produces low-cardinality KeyValues (few distinct values — safe as metric tags/dimensions) and high-cardinality KeyValues (many distinct values like an id — attached to spans only, never to metric dimensions).

open as a page

What is an Observation.Context, and how do ObservationConvention and a custom Context work together in reusable instrumentation?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Observation.Context is a mutable data bag that travels through an Observation's lifecycle, carrying its name, KeyValues, error, and any custom fields. An ObservationConvention reads that Context to decide the Observation's name and KeyValues in one reusable place.

open as a page

How does trace context propagate across non-HTTP boundaries like Kafka or RabbitMQ, and what do you do when auto-instrumentation isn't available?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Context travels as message headers instead of HTTP headers. Spring's Kafka/RabbitMQ instrumentation injects traceparent/baggage into producer record headers and extracts them on the consumer. Where no auto-instrumentation exists, use the Propagator API to inject on send and extract on receive manually.

open as a page

Why not just set sampling probability to 1.0 everywhere? Discuss the cost-vs-coverage trade-off.

level: seniorimportance: should knowfreq 50%

basics

~20 s

Sampling 100% is accurate but expensive: more CPU/memory per request, more network to the collector, and much higher storage/backend cost. Lower probabilities cut cost but risk missing the rare errored or slow request. You balance budget against how much visibility you need.

open as a page

How would you configure a custom sampling strategy in Spring Boot — for example, always sampling a critical endpoint but rate-limiting the rest?

level: seniorimportance: should knowfreq 35%

basics

~10 s

For simple cases set management.tracing.sampling.probability. For custom logic, define your own Sampler bean (Brave's brave.sampler.Sampler or OpenTelemetry's io.opentelemetry.sdk.trace.samplers.Sampler), which overrides the property. Brave offers RateLimitingSampler and per-path samplers you can compose.

open as a page

In production, sampling probability is 0.1 and users report 'traces are missing.' Explain how sampling, context propagation, and the exporter interact, and what you'd verify.

level: principalimportance: should knowfreq 33%

basics

~20 s

At 0.1 only ~10% of traces are exported by design — that's expected, not a bug. The sampling decision is made once at the trace root and propagated downstream, so all services agree. Verify the endpoint, that ids still appear in logs (proving propagation works), and consider raising probability or using tail sampling in the collector.

open as a page

Walk through the full Observation lifecycle and explain how handlers, scopes, predicates, and conventions interact when you build custom instrumentation.

level: principalimportance: should knowfreq 20%

basics

~20 s

Create an Observation with a convention and context, start it (onStart fires), open a scope so nested work sees trace context (onScopeOpened), run the work, record errors (onError), then stop it (onStop records timing/finishes the span). Predicates decide up front whether it runs at all; conventions decide naming/tags; handlers do the actual recording.

open as a page

Walk through the full lifecycle of a manually managed Observation (createNotStarted → start → openScope → error → stop). What breaks context propagation and how do you avoid it?

level: principalimportance: should knowfreq 30%

basics

~20 s

Create with createNotStarted, start() it, openScope() to bind it to the thread (so child spans/logs correlate), run work, call error(e) on failure, close the scope, then stop(). Forgetting the scope breaks correlation; forgetting stop() leaks. observe() does all this safely; async work needs context propagation.

open as a page

As a platform owner, what governance and failure-mode concerns would you set for Baggage propagation across a large service fleet?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Restrict baggage to a small, approved allow-list of small non-sensitive fields; forbid secrets/PII since baggage is forwarded everywhere and logged. Standardize field names and propagation format fleet-wide, strip baggage at public trust boundaries, and watch header-size/overhead on high-volume paths.

open as a page

What are the fundamental limitations of probabilistic head-based sampling, and how would you architect around them at scale?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Head-based sampling decides before the outcome is known, so it randomly drops errors and slow requests and can't guarantee capturing rare problems. At scale you add tail-based sampling in a collector, rate limiting, per-route rates, force-sample-on-debug, and lean on metrics/exemplars so tracing isn't your only signal.

open as a page