skip to content

You're setting the observability standard for an organization with 200 microservices owned by 40 teams, each currently picking its own logging format and tracing library ad hoc. What decisions do you need to make about traces, metrics, logs, and sampling to make cross-service debugging tractable without bankrupting the observability budget?

level: principalimportance: should knowfreq 40%

answer

  1. standardize instrumentation layer (OpenTelemetry) org-wide
  2. shared log schema + mandatory propagation contract
  3. hybrid sampling (head + tail) for cost control
  4. cardinality limits enforced at ingestion
  5. paved-road libraries + enforcement, not just a mandate

basics

~20 s

You pick one shared standard, like OpenTelemetry, for how every team generates traces, metrics, and logs, so they all fit together and can be searched or joined across teams. You also set shared rules for what gets sampled and how detailed logs can be, so cost stays under control as the company grows.

solid answer

~1 min

The core decision is standardization: mandate a single instrumentation layer, in practice OpenTelemetry, as the shared API and wire format for traces, metrics, and logs across every team, so a trace started in one service can be correctly continued by any other service regardless of language or team, and so every team's telemetry lands in the same backend with the same shape. On top of that you set concrete, enforced conventions: a required correlation and trace-context propagation contract across all synchronous and asynchronous boundaries, including third-party integrations, a shared log schema with mandatory fields like service, trace_id, and level, a small set of standardized RED and USE dashboard templates so on-call engineers can read any team's dashboard without ramping up, and explicit cardinality limits on metric labels enforced at the ingestion layer. For cost, you set an explicit sampling policy, typically a hybrid of head-based sampling to cap raw ingestion plus tail-based rules to guarantee errors and outliers are kept, with retention tiers, short cheap retention for high-volume raw data, longer retention only for aggregated or sampled data. None of this works without also owning enforcement: linting and CI checks for missing instrumentation, a paved-road library every team consumes by default, and cost and usage dashboards attributing observability spend back to each team so incentives to avoid cardinality abuse and over-logging are aligned.

go deeper

for a junior

Not typically expected to design this; if asked, should at least recognize that having every team pick its own tools makes cross-team debugging hard.

for a middle

Should recognize the value of a shared instrumentation standard and structured log schema, even without designing the full rollout.

for a senior

Should be able to design the standard itself, OpenTelemetry, shared schema, RED and USE dashboards, sampling policy, for a single team or a handful of services.

for a principal

Should design the org-wide rollout: enforcement mechanisms, cost and retention tiering, cardinality guardrails, and the autonomy-versus-control trade-off across 40 independently-shipping teams.

## The problem at this scale At the scale of 200 services and 40 teams, the central problem stops being how do I instrument my service and becomes how do I make 40 independently evolving teams' telemetry composable, comparable, and affordable, without dictating every implementation detail. Left alone, teams converge on incompatible choices: - one team's JSON logs use timestamp, another's use ts; - one team traces with a vendor SDK that doesn't propagate context the same way another team's older library does; - and the result is that a single user-facing incident touching five teams' services can't be debugged as one story, because the signals don't line up. The job of an observability strategy at this scale is fundamentally standardization plus enforcement, more than it is choosing clever individual techniques. ## Decision one: the instrumentation layer The first decision is the instrumentation layer itself: standardize on **OpenTelemetry**, or in older organizations retrofit toward it, as the single API surface every service uses to emit traces, metrics, and logs, and as the wire protocol, OTLP and W3C Trace Context, every service uses to talk to a shared collector tier. This matters because OpenTelemetry decouples what a team instruments from where that data ultimately lands, letting the org swap or run multiple backends, Jaeger, a commercial APM, an in-house store, behind the same instrumentation without re-instrumenting 200 services, and it guarantees that a trace context generated in a Go service and continued in a Java service actually stitches together correctly, since both speak the same propagation format. ## Decision two: the mandatory conventions The second decision is a small set of mandatory conventions layered on top of that shared instrumentation, since a shared library alone doesn't stop teams from using it inconsistently. - This includes **a required propagation contract**: every synchronous call, every message queue publish and consume, every scheduled job must propagate trace context, verified by paved-road client libraries that do it automatically and, ideally, by CI or runtime checks that flag services emitting orphaned traces. - It includes **a shared structured-log schema**: a small mandatory field set, timestamp, service name, environment, `trace_id`, `span_id`, severity, so that logs from any service can be correlated with traces from any service, and so a central log query can be written once and work everywhere. - It includes **standardized dashboard templates** built on RED for services and USE for infrastructure, so an on-call engineer paged into an unfamiliar team's service can read its dashboards using muscle memory from their own team's dashboards, rather than learning a new layout under pressure. ## Decision three: explicit cost control The third decision is explicit cost control, since at this scale, uncontrolled telemetry volume is a real budget line item, not an abstraction. - This means **an explicit sampling policy**, most realistically a hybrid: head-based sampling at the edge to bound raw ingestion volume into the collector tier, since ingestion cost scales with total spans generated, not just what's ultimately stored, combined with tail-based rules at the collector to guarantee that errors and high-latency outliers are retained regardless of the head-based rate, so rare but important incidents aren't lost to chance. - It also means **retention tiering**: high-fidelity raw traces and debug-level logs kept cheaply for a short window, hours to a few days, long enough for active incident response, with only aggregated or heavily-sampled data retained longer for trend analysis, since indefinite full-fidelity retention at 200-service scale is rarely justifiable. - And it means **hard, enforced cardinality limits** on metric labels at the ingestion layer, since a single team's cardinality mistake, an unbounded label like a raw user ID, can otherwise degrade or crash a metrics backend shared by all 40 teams, turning one team's bug into an organization-wide outage of the observability system itself. ## Control against autonomy The trade-off running through every one of these decisions is central control versus team autonomy. Mandating a single instrumentation standard and schema is real friction for individual teams: migration cost for existing services, a slower on-ramp for a new team that would rather just add a print statement, and a central function, a platform or observability team, that must build and maintain the paved-road libraries, collector infrastructure, and enforcement tooling that makes the standard the path of least resistance rather than a rule nobody follows. Get this balance wrong in the direction of too little enforcement and you're back to 40 incompatible dialects within a year, no matter what standard was announced; get it wrong in the direction of too much centralization and teams route around the mandate with shadow logging or their own separate APM subscriptions, fragmenting things anyway while also duplicating cost. ## When the tooling becomes the incident The clearest failure mode at this scale isn't any single team's mistake, it's an aggregate one: the observability system itself becomes the incident. - A shared metrics backend running out of memory from one team's cardinality mistake, - a shared log cluster's ingestion pipeline backing up because one service started logging at debug level in production, - or a collector tier falling over under 100% head-based sampling turned on just for today during an incident and never turned back off, all take down visibility for every other team at exactly the moment it's needed most. A concrete real-world instance of this pattern: large organizations like Uber, which built Jaeger for this exact reason, and Google, with its internal Dapper system, built or adopted org-wide tracing standards specifically because ad hoc, per-team instrumentation stopped scaling once request chains routinely crossed dozens of services owned by dozens of teams.

  • How would you actually get 40 existing teams, each with their own logging setup, to migrate to a shared standard without a top-down mandate causing revolt?
    Make the shared standard the path of least resistance rather than a bureaucratic requirement: ship a paved-road library that wraps OpenTelemetry with your org's schema pre-configured, so adopting it is less work than what a team currently has, bundle migration into a broader platform initiative teams already need, and use CI checks or gentle nudges, like dashboards showing a service is missing trace propagation, rather than blocking deploys outright, at least initially.
  • What's the risk of letting each team choose its own sampling rate independently?
    Inconsistent sampling rates across services in the same call chain break trace completeness: if one service samples a trace in but another service, further downstream in the same request, independently decides not to sample, you get a partial trace missing exactly the hop you needed. Sampling decisions need to be made once, at the root of a trace, and honored by every downstream service via the propagated sampled flag, not re-decided independently at each hop.
  • How do you prevent one team's cardinality mistake in metrics from taking down the shared backend used by all 40 teams?
    Enforce label-cardinality limits at the ingestion layer itself, for example rejecting or dropping series generation beyond a per-team quota, rather than relying on every team to self-police; this turns a shared-infrastructure incident into an isolated, contained problem for the offending team's own metrics rather than an org-wide outage.

It's like standardizing electrical outlets and voltage across every building in a city instead of letting each building pick its own plug shape and voltage: any one building's wiring might work fine on its own, but nothing plugs into anything else, and the city can't run shared infrastructure, like a power grid, or here a shared observability backend, across all of them without that common standard.

saying these in an interview costs you the question

  • Proposes letting every team keep choosing its own tools and schema as long as they try to be consistent
  • No mention of enforcement mechanisms such as CI checks or ingestion-layer limits, only a written standard
  • Ignores cost and retention as a first-class design constraint
  • Doesn't address how sampling decisions stay consistent across teams and services
  • Treats this as a purely technical problem with no mention of adoption or migration strategy

context