skip to content

Agent vs Library Instrumentation

The two ways telemetry gets into your code: attach an agent that instruments automatically, or call an SDK explicitly. Interviewers probe the coverage, overhead, and lock-in trade-offs behind that choice.

on this pageshow

questions

4

What can an automatic instrumentation agent see inside a running service, and what is it structurally blind to?

level: middleimportance: must knowfreq 62%

answer

  1. Coverage comes from a catalogue, not magic
  2. It hooks libraries, not your logic
  3. Spans appear where I/O crosses a boundary
  4. Compare self time against the children
  5. A version mismatch fails silently

basics

~20 s

An automatic instrumentation agent hooks call sites it has a plugin for - web handlers, database drivers, message clients - so it produces spans at library boundaries only. Domain logic between those boundaries stays one unattributed gap in the trace.

solid answer

~50 s

An automatic instrumentation agent attaches to the process at startup and wraps or rewrites code at call sites it recognises: inbound request handlers, outbound HTTP and RPC clients, database and cache drivers, message producers and consumers, plus the thread-pool and callback machinery it needs to keep context attached across hand-offs. That buys real coverage with no code change and an identical vocabulary across every service. What it cannot see is anything with no plugin - your domain logic, an in-house transport, a library version whose internals moved - and any *meaning* above the call: which branch ran, which tenant, why a fallback fired. The result is a trace shaped like your I/O rather than like your program: correct at the boundaries, opaque in between, with the service's own **self time** appearing as one flat interval nobody can attribute.

code

text · 5 lines
text
POST /orders/school-meal-week          847 ms   [agent: server span]
  SELECT menu_items ...                 14 ms   [agent: db client]
  SELECT pupil_allowance ...            11 ms   [agent: db client]
  POST payments.internal/charge         63 ms   [agent: http client]
  (759 ms unaccounted: eligibility rules, price banding, receipt build)

go deeper

for a junior

Be ready to say what you get for free when an automatic instrumentation agent is attached: spans for inbound requests and for calls to common libraries, with no code change. Knowing that its coverage comes from a fixed catalogue of supported libraries is enough at this level.

for a middle

Explain the mechanism: the agent wraps recognised call sites at load time, so spans appear at I/O boundaries. Be able to describe what falls outside that - your own logic, unrecognised libraries, and version drift that fails the match silently.

for a senior

Show you can read a trace with a large unattributed self time and diagnose it rather than declare the agent broken: check what actually installed, check library versions, confirm the work is in-process, then instrument the gap deliberately.

for a principal

Own the rollout tradeoff. Agent coverage is the only affordable baseline across a large fleet, but it is third-party code in every process, it lengthens start-up, and its span volume scales with traffic - so decide where the baseline sits and who pays for depth above it.

## What "automatic" means mechanically An automatic instrumentation agent is code loaded into the application process before or alongside your own code: a runtime agent that rewrites classes as they load, an import hook that wraps functions, a preloaded shared library, or a runtime switch that enables hooks already compiled in. The mechanism differs per language; the shape does not. **The agent carries a catalogue of instrumentations, each matching a named library within a version range.** When a matching type or function is loaded, the agent wraps it — entering starts a span, leaving ends it, and the error path records the failure on it. Everything else follows from that catalogue. Coverage is exactly the set of libraries somebody wrote a plugin for, running at versions those plugins still match. Nothing in the word *automatic* implies the agent understands your program. ## What it reliably gives you - **A server span per inbound request** — protocol, route template, status, duration. - **A client span per outbound call it recognises** — HTTP and RPC clients, database drivers, cache clients, message producers and consumers. - **Context plumbing** — writing trace-context headers onto outbound requests, reading them off inbound ones, and following the runtime's thread pools, executors and callback machinery so the active trace survives a hand-off. - **Errors at those boundaries** — the exception that escaped a hooked call, attached to the span that was open when it escaped. - **Process identity** — host, runtime, version, plus whatever service name you configure. That is not a small gift. It is uniform across services, it costs no engineering time per service, and it is the only realistic way to get a baseline on a fleet nobody has time to instrument by hand. ## What it is structurally blind to | Blind spot | Why the agent cannot see it | | --- | --- | | Your domain logic | No library call is crossed, so no hook fires | | An in-house transport or framework | No plugin was ever written for it | | Which branch ran, and why | The hook sees the call, not the decision behind it | | Business identity (tenant, plan, cohort) | Those values live in your objects, not in the call signature | | A library version outside a plugin's match range | The match fails silently — no error, simply no spans | | Time below the hook point, such as lock waiting | The hook measures only the call it wrapped | The last two bite hardest. A silent match failure is indistinguishable from "that code never ran", and any time not spent inside a hooked call simply disappears into the parent's duration. ## What that does to the shape of the trace Take a school-meal ordering service at its 5,400-request-per-second lunchtime peak. Agent-only, one order request renders as a server span of 847 ms containing three children: two database calls of 14 ms and 11 ms, and one outbound call to an internal payments service of 63 ms. Eighty-eight milliseconds are accounted for. The other 759 ms are not. The agent is telling the truth. The request did take 847 ms and 88 ms of it was I/O the agent recognises. What it cannot tell you is that 612 ms of the remainder went into evaluating eligibility rules, because no call site in its catalogue was crossed while that ran. The trace is **correct at the boundaries and opaque in between** — shaped like your I/O rather than like your program. Engineers misread that gap in two opposite directions. Some conclude the trace is broken; it is not, it is complete for what it covers. Others conclude the service was waiting on something invisible; it was not, it was running. The right reading of a large difference between a span's duration and the sum of its children — its **self time** — is "the answer lives in code this agent does not instrument". ## Diagnosing before you blame the agent 1. Compare each span's duration with the sum of its children. A large self time is a pointer, not a fault. 2. Check what the agent actually installed. Agents report which instrumentations loaded; a plugin that did not match will be absent from that report. 3. Check library versions against the agent's supported range, especially straight after a dependency bump. 4. Confirm the work is genuinely in-process. An outbound call made over a raw socket, with no recognised client library, also surfaces as self time. 5. Only then add instrumentation, and add it exactly where the gap is rather than everywhere. ## The costs you are accepting Class rewriting or import patching happens at startup, so start-up time grows measurably and start-up failures become a new class of incident. Every span costs allocation, serialisation and network bytes: a service emitting six spans per request at 5,400 requests per second is producing 32,400 spans per second before any sampling decision is taken. And an agent is third-party code running inside your process — a plugin interacting badly with a library version can change exception behaviour or timing. Roll an agent out the way you roll out a deploy: canary first, with a switch that turns it off.

  • A service upgrades its HTTP client library and its outbound spans disappear, with nothing in the error log. What happened?
    The agent's plugin matches specific types and version ranges of that library. When the internals move outside the range, the match simply does not fire - there is no exception, so the only symptom is missing spans. Check the agent's report of which instrumentations installed, then align the agent version with the library version.
  • Why can an agent-produced trace look complete and still fail to explain why a request was slow?
    Because it times boundaries, not decisions. You learn which hop consumed time, but the service's own self time is one undifferentiated interval, and no span carries the domain identity - tenant, route class, which branch ran - that would let you find comparable slow requests and see what they have in common.
  • Does attaching an automatic agent change how the application behaves?
    Yes, enough to treat it as a real deploy. Rewriting or patching code lengthens start-up, spans cost allocation and bytes at request rate, and a plugin interacting badly with a library version can alter exception paths or timing. Canary it, watch start-up and error rates, and keep a switch that turns it off.

An automatic agent is like cameras mounted in a building's doorways: you know exactly when someone entered and left each room, and nothing at all about what they did while inside.

saying these in an interview costs you the question

  • Claims an agent instruments all application code automatically
  • Assumes missing spans always mean the agent has crashed
  • Believes agent attachment costs nothing at start-up or per span
  • Reads a flat span with no children as the service being idle
  • Cannot name anything an agent is structurally unable to see
  • Treats agent coverage as a substitute for domain attributes
open as a page

When an automatic instrumentation agent is already running, which code still earns a hand-written span?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Code the agent leaves as an unattributed gap you need explained, plus attributes carrying domain identity and outcome. Instrument decisions and meaningful units of work, not every function - and prefer an attribute over a new span.

open as a page

What creates instrumentation lock-in across a service fleet, and what would you standardise to reduce it?

level: principalimportance: should knowfreq 40%

basics

~20 s

Lock-in lives in what was built on top of the telemetry, not in the code that emits it. Changing the emitting mechanism is a rollout; re-authoring every dashboard and alert rule that hardcodes an attribute name is the real cost.

open as a page

How do an in-process instrumentation agent, kernel-level observation and a sidecar proxy differ in what telemetry each can capture?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Three vantage points, three blind spots. An in-process agent sees application objects and can carry trace context; kernel-level observation sees sockets and syscalls in any language but no application meaning; a sidecar proxy sees only traffic that crosses the network.

open as a page