skip to content

You own observability across a few hundred services and must decide how deeply each is instrumented with OpenTelemetry. How do you split effort between running an auto-instrumentation agent everywhere and hand-written spans, and how do you keep the result consistent and affordable?

level: principalimportance: should knowfreq 30%

answer

  1. agent = floor, mandated by platform
  2. propagation gap breaks traces, not just blanks them
  3. manual spans only where decisions live
  4. one bytecode agent per process, ever
  5. govern namespace + cardinality in review

basics

~20 s

Make the agent the default everywhere: it buys uniform boundary coverage and context propagation with no per-team work. Spend scarce manual instrumentation only where domain decisions live. Govern a shared attribute vocabulary, and control cost through span granularity rather than by removing instrumentation.

solid answer

~60 s

Treat coverage as a **floor plus targeted depth**. The floor is the agent, mandated in the base image or platform template: every service gets inbound and outbound spans and, crucially, context propagation — which is what makes cross-service traces exist at all. It requires no team effort, so adoption does not depend on goodwill. Depth is bought deliberately. Hand-written spans belong where a request makes a decision the agent cannot see: routing to a fallback, tenant class, cache miss path, batch stage boundaries, and any in-house transport the agent does not recognise. Everything else stays uninstrumented on purpose. Consistency comes from a governed vocabulary: standard conventions plus a company namespace with owners, enforced in review rather than by hope. Cost is driven by span *count*, so the levers are granularity (never a span per loop iteration or per retry attempt) and, separately, the sampling and export configuration. Rollout risks: never two bytecode agents in one process, canary the agent for startup and latency regressions, and keep per-instrumentation and global kill switches.

go deeper

for a junior

Not expected to own this; be able to say the agent gives broad coverage cheaply and manual spans add business meaning.

for a middle

Argue the floor-plus-depth split and give concrete examples of what deserves a hand-written span.

for a senior

Add rollout mechanics: canaries, kill switches, agent conflicts, version lag, and granularity as the cost lever.

for a principal

Own it as policy — platform-mandated agent, governed vocabulary with owners, coverage measured by trace completeness and incident answerability, migrations as cutovers.

## Framing the decision At fleet scale the constraint is not technical capability but the number of engineer-hours you can spend per service. So decide once, centrally, what every service gets for free, and reserve human effort for what only humans can supply. ## The floor: agent everywhere Auto-instrumentation should be part of the platform, not a per-team project — baked into base images, launch templates or an injecting operator, on by default. Two reasons it is worth mandating rather than recommending: 1. **Propagation**, not spans, is the real deliverable. A service without instrumentation is not merely a blank in the trace: it *breaks* the trace, because incoming context is never extracted or re-injected downstream. One unpatched hop can orphan everything behind it. Uniformity is a correctness property. 2. **Marginal cost per team is zero.** Anything requiring per-team work will have long-tail non-adoption, and the long tail is exactly where incidents get hard. ## Where to buy depth Manual instrumentation is worth its cost only where it answers a question the boundary map cannot: - Branches that change behaviour: fallback vs primary, cache hit vs miss, degraded mode. - Business identity as attributes: tenant class, plan, region, request kind — the dimensions your incidents are actually sliced by. - Stage boundaries in long-running or batch work, where a single overall span hides everything. - In-house or exotic transports the agent does not recognise; these need a wrapper that both creates spans and injects context. And where not to: leaf helpers, per-iteration loop bodies, individual retry attempts, and pure functions. If a span would never change a decision during an incident, it is data you pay to store and scroll past. ## Consistency: govern the vocabulary At this scale the failure mode is not missing telemetry, it is telemetry that cannot be joined — three names for tenant, four spellings of environment. Publish a small attribute standard: use standard semantic conventions wherever one exists, put everything else in a company namespace with a named owner per key, and cap cardinality on anything that may become a metric dimension. Enforce it in code review or a lint rule; a wiki page alone does not survive contact with a hundred teams. Instrumentation scope naming deserves the same treatment, so you can tell whose instrumentation emitted a span. ## Cost Spend scales with exported spans, and the biggest multipliers are granularity choices made in code, not backend pricing. In order: eliminate per-iteration and per-attempt spans; keep attribute counts and value lengths sane; then apply retention policy through sampling and export configuration, which is a separate lever with its own trade-offs. Do not manage cost by disabling instrumentation, because that also disables propagation and blinds the traces you kept. ## Rollout risk - **Never run two bytecode-instrumenting agents in one process.** A vendor agent plus a general one produces duplicated spans at best and class-transformation conflicts or crashes at worst. Migrations are cutovers, service by service. - **Canary.** Agents add startup time (class transformation) and a small steady-state overhead. Measure both on a representative service before fleet-wide rollout, especially where cold-start latency is user-visible. - **Keep switches.** A global disable plus per-instrumentation toggles let you neutralise a bad interaction in one deploy without ripping the agent out. - **Version lag is real.** Agents support library versions with a delay; a framework upgrade can silently lose instrumentation. Track coverage as a metric — services reporting spans, and traces completing end to end — so silent loss is visible. ## How you know it worked Define success as answerability, not adoption: can an on-call engineer take a user-visible symptom and reach the responsible component from telemetry alone? Sample real incidents and check. That measure will tell you where to buy the next increment of depth, far better than counting instrumented services.

  • A team wants to keep their existing vendor agent and add the OpenTelemetry agent alongside it. What do you tell them?
    Not in the same process. Two agents transforming the same classes duplicate spans, disagree on context, and can fail class loading outright. Plan a cutover instead: run the vendor agent until the OpenTelemetry pipeline is verified for that service, then switch in one deploy, with a rollback. If both backends are needed for a while, fan out from the collector rather than from the process.
  • How do you measure whether instrumentation coverage is actually adequate rather than merely widespread?
    Count answerability, not adoption. Track the share of services emitting spans and, more importantly, the share of traces that complete across every hop without an orphan — a propagation gap is invisible in adoption counts. Then review real incidents: was the responsible component reachable from telemetry, or did someone have to read code and logs? Those gaps tell you where depth is worth buying.

saying these in an interview costs you the question

  • Leaving agent adoption to individual teams and treating a missing service as a mere blank rather than a broken trace
  • Running a vendor bytecode agent and the OpenTelemetry agent in the same process
  • Controlling cost by disabling instrumentation, which also disables context propagation
  • Letting each team invent its own attribute names with no owned namespace or cardinality rules
  • Rolling an agent fleet-wide without canarying startup and steady-state overhead

context