skip to content

You own the observability platform and want engineers to be able to move from a metric spike, to one slow request, to that request's logs, for any service in the fleet. What conventions and policies have to be true across teams for that navigation to work reliably, and where do you accept that it will not?

level: principalimportance: nice to knowfreq 22%

answer

  1. Navigation is a pipeline contract, not a UI feature
  2. One service identity, mapped into every store's label vocabulary
  3. Trace id: structured field, not regex, not an index label
  4. Tail sampling on latency/errors so exemplar links resolve
  5. Retention horizon published; UIDs pinned; config as code

basics

~20 s

Fleet-wide you need one service-identity convention mapped consistently into every store, the trace id present and exactly matchable in logs, exemplars on the key latency metrics, a sampling policy that keeps the traces those exemplars name, and retention windows aligned. Accept broken navigation outside trace retention and for unsampled, unremarkable requests.

solid answer

~60 s

Cross-signal navigation is a **contract between pipelines**, not a Grafana feature. Five things must hold fleet-wide: 1. **One service identity**, mapped deterministically into each store's label vocabulary, so a jump built from a span attribute finds the matching log stream every time instead of per-service. 2. **Trace id in logs, exactly matchable** — as a structured field or structured metadata, never only as free text needing a regex, and never as an index label (cardinality). 3. **Exemplars on the RED-style latency metrics**, with the exemplar path enabled all the way through the metrics store, including per-tenant limits. 4. **Sampling aligned with the workflow**: tail-based on latency and errors, so the traces exemplars and log links point at are exactly the ones kept. Head sampling at a low rate guarantees mostly dead links. 5. **Retention alignment**, or at least a stated window inside which links are expected to work. Delivery matters too: pin data source **UIDs** and provision derived fields, trace-to-logs and correlations as code, so navigation is identical in every environment and reviewable. What I accept: no navigation beyond trace retention, none for unsampled routine requests, and higher cost only where the workflow earns it.

go deeper

for a junior

Not expected. The takeaway is that links only exist because the application emitted a shared identifier.

for a middle

Name the prerequisites — consistent service labels, trace id in logs, exemplars enabled — without needing to own the policy trade-offs.

for a senior

Argue the sampling and retention coupling concretely and insist the navigation configuration be provisioned as code with pinned UIDs.

for a principal

Present it as a contract with published horizons, a conformance check at onboarding, tiered investment by service criticality, and an explicit ingest-versus-query cost position.

## The framing Engineers experience cross-signal navigation as a UI feature. It is not. Grafana renders links over identifiers the telemetry already contains; if the identifiers are absent, inconsistent or point at data that no longer exists, the UI has nothing to render. So the platform question is: *what invariants must every service's telemetry satisfy, and who enforces them?* ## The invariants **Service identity.** Each store has its own vocabulary — a resource attribute on spans, a stream label on logs, a metric label. Pick one canonical identity and one deterministic mapping into each store's naming, applied at the collection layer, so navigation configuration is written once rather than per service. When teams configure their own agents, this is the invariant that erodes first: two services in the same cluster end up with different label names and one of them silently loses its log jump. **Request identity.** The trace id must reach the log store in a form that can be filtered exactly — a structured field or structured metadata. Not a regex over free text (a heuristic that breaks silently on format change), and not an index label (unbounded cardinality). This is a logging-library convention plus a pipeline convention, and it is cheap to enforce with a shared logging configuration. **Aggregate-to-individual bridge.** Exemplars on the latency histograms of the metrics that people actually alert on. The chain is only as strong as its weakest link: instrumentation records them, exposition carries them, the metrics store admits them (per-tenant exemplar limits default to zero in some deployments), the data source links them. **Sampling policy.** This is where most platforms quietly fail. With head sampling at a low rate, a metric exemplar still records the trace id of the request it sampled, but that trace was never exported — so the link is a promise the platform cannot keep. Tail-based sampling keyed on latency, errors and a small always-on baseline makes the kept set nearly identical to the interesting set, which is precisely what exemplars and log links point at. Frame it explicitly as *coupling sampling policy to the navigation product*, not as a cost knob chosen in isolation. **Retention alignment.** Metrics live for months, logs for weeks, traces for days, because that is what they cost. Every link crossing from a longer-retained signal into a shorter-retained one has a horizon beyond which it dead-ends by construction. Two honest responses: publish the horizon ("trace links work for 72 hours"), and consider longer retention only for the small set of traces that tail sampling already marked as interesting. ## Where the cost lands There is an ingest-versus-query trade running through all of it. Enriching at ingest — adding consistent labels, promoting the trace id into structured metadata, keeping tail-sampled traces — costs storage and pipeline complexity but makes navigation instant and reliable. Deferring to query time — regex extraction, wide time windows, cross-store guessing — costs nothing at ingest and gives a slow, flaky experience precisely when someone is debugging an incident. Spend at ingest where the navigation is used, and nowhere else. Cardinality is the hard constraint on the enrich-more instinct. Every label you add for correlation multiplies series or streams. The discipline: identity dimensions become labels; request-scoped identifiers stay out of labels and live in the payload or in a structured-metadata-style side channel designed for high-cardinality attachment. ## Delivery and governance All of the navigation configuration — derived fields, trace-to-logs mappings, exemplar links, correlations — resolves targets by **data source UID**. If UIDs are generated per instance, links work in the environment where someone clicked them into place and nowhere else. Pin UIDs in provisioning, ship the navigation configuration as code, and review it like any other platform contract. Then make conformance visible rather than assumed: a small check that, per service, answers whether logs carry a matchable trace id, whether the latency histogram emits exemplars, and whether the identity labels match the convention. Onboarding a service means passing that check. ## What I explicitly do not promise - Navigation outside the trace retention horizon. - A trace for every request; sampling means most routine requests have none, and that is the correct economic answer. - Uniform depth across the fleet — a tier-1 payment path can justify a much higher sampling floor and longer trace retention than an internal batch job, and pretending otherwise either bankrupts the platform or starves the paths that matter.

  • A team argues for head sampling at 1% to control trace cost. How do you respond given the navigation goals?
    Head sampling at 1% means roughly 99 of every 100 exemplar and log links point at traces that were never stored, so the navigation feature effectively does not exist while still costing configuration and user trust. The counter-proposal is tail-based sampling keyed on latency and errors plus a small always-on baseline: similar or lower storage, but the kept set overlaps almost exactly with the set anyone would ever click through to. If tail sampling is not available, raise the head rate only for the services where the workflow is actually used.
  • How do you keep the correlation labels from wrecking cardinality?
    Separate identity from request scope. Identity dimensions — service, environment, namespace, maybe route class — are bounded and may become labels. Request-scoped identifiers such as trace, span, user or order ids are unbounded and must never become index labels; they belong in the payload or in a high-cardinality side channel such as structured metadata, which exists precisely so they can be filtered without joining the index. Enforce it with a limit and an alert on stream/series growth, not with good intentions.

saying these in an interview costs you the question

  • Treating cross-signal navigation as a Grafana configuration exercise rather than a telemetry contract
  • Promising links that cannot resolve because sampling or retention was chosen without reference to the workflow
  • Adding trace or request ids as index labels to make joins easy
  • Letting each team configure its own agent naming, so navigation works service-by-service at random
  • Configuring navigation in the UI of one instance and assuming it transfers, when links resolve by data source UID

context