skip to content

When an automatic instrumentation agent is already running, which code still earns a hand-written span?

level: seniorimportance: should knowfreq 48%

answer

  1. The agent drew only the skeleton
  2. Instrument decisions, not every function
  3. Attributes are cheaper than extra spans
  4. Name the incident question it answers
  5. Ask what decision changes with the answer

basics

~20 s

Code the agent leaves as an unattributed gap you need explained, plus attributes carrying domain identity and outcome. Instrument decisions and meaningful units of work, not every function - and prefer an attribute over a new span.

solid answer

~40 s

The agent draws a skeleton at I/O boundaries, which answers *which hop* was slow. Explicit instrumentation adds the three things only your code knows: the **unit of work that means something** to the business rather than to the transport, **identity** that makes a trace findable (tenant, route class, variant, batch size), and the **outcome or decision** - a fallback taken, a cache miss, a rejection, a retry. The selection rule is a pair of questions: which incident question does this span answer, and what decision changes with the answer? If both are unanswerable, it is decoration. And prefer an attribute on a span you already emit over a new span: attributes ride on records already in flight, whereas an extra span at a 5,400-request-per-second peak is 5,400 more records per second, forever.

code

pseudocode · 12 lines
pseudocode
span = start_span("eligibility.evaluate")
span.set_attribute("pupil.year_group", year_group)
span.set_attribute("rules.count", 37)
try:
    band = rules.apply(pupil, menu)
    span.set_attribute("subsidy.band", band)
except RuleTimeout as err:
    span.record_error(err)
    span.set_attribute("fallback.used", "cached_band")
    band = cache.last_band(pupil)
finally:
    span.end()

go deeper

for a junior

Recall that you can add your own spans and attributes on top of what an agent produces, and that the point of an attribute is to make a trace searchable later - for example by tenant, route or outcome.

for a middle

Explain the division of labour: the agent covers the I/O boundaries, and your code adds the meaningful unit of work, the identity and the outcome. Be able to say why an attribute usually beats a new span.

for a senior

Show a selection rule rather than a wish list. Tie each proposed span to an incident question it answers and a decision it changes, and be able to describe the volume cost of one extra span per request at real traffic.

for a principal

Own the discipline across teams: instrumentation is an investment measured in questions answerable under pressure, span names and attribute keys are an interface others bind to, and removing dead instrumentation is legitimate, funded work.

## What the agent already drew An automatic instrumentation agent gives you a skeleton: a span where a request entered, a span for each recognised outbound call, durations and error status on both. That skeleton answers *which hop* was slow. It cannot answer *which decision* was slow, or *for whom*, because those facts live in your objects and never appear in the signature of a library call. So the question is not whether to add explicit instrumentation on top; it is what to add, and how to stop. Both failure modes are common and both are expensive: services with an agent and nothing else, whose traces bottom out in an unattributable gap, and services with three hundred hand-written spans that mirror the call graph and answer nothing. ## Three things only your own code knows - **The unit of work that means something.** The agent names spans after transport operations. Your code knows that the interesting unit is "evaluate eligibility", "reprice the week's menu", "generate the receipt" — the things whose duration you would actually act on. - **Identity.** Tenant, school, plan tier, route class, feature-flag variant, batch size, whether the request came from a retry. These make a trace *findable*. A trace you cannot find when you need it is a trace you did not collect. - **Outcome and decision.** A fallback was taken. A cache missed and the slow path ran. A validation rejected the order. Three attempts happened before one succeeded. The agent sees a successful HTTP 200; only your code knows which of four routes through the function produced it. ## The rule for what earns a hand-written span A new span is worth its cost when you can answer both of these: 1. **What incident question does it answer?** Name the question you have actually asked during an incident and could not answer. "Was the slowness in rules evaluation or in pricing?" is a question. "More visibility" is not. 2. **What decision changes with the answer?** If both possible answers lead to the same action, the span is decoration. Two supporting tests catch most of the rest. If the work is shorter than the noise floor of the surrounding measurements, the span adds bytes and no signal. And if the span would exist once per request per service, ask whether the same fact could be an attribute on a span you already emit. ## Span or attribute? | You want to know | Reach for | | --- | --- | | How long a distinct piece of work took | A span | | Whether that piece of work failed on its own | A span, so it can carry its own status | | Which tenant, route, variant or cohort a request was | An attribute on an existing span | | Which branch or fallback ran | An attribute on an existing span | | How many items a batch held | An attribute on an existing span | Attributes are dramatically cheaper. Adding one to a span you already emit costs bytes on a record already in flight, and unlike a dimension on an aggregated metric it does not multiply anything stored. Adding a span costs a whole new record, per request, forever. In a school-meal ordering service peaking at 5,400 requests per second, one extra span per request is 5,400 additional spans per second; one extra attribute is a few dozen bytes on records you are already paying for. ## What over-instrumentation looks like - Span names that mirror function names, so the trace reads like a stack rather than a story. - Traces with hundreds of spans that nobody expands, because the interesting one cannot be found. - Spans around getters, mappers and constructors added "for completeness". - Hand-written spans duplicating a boundary the agent already covers, so every database call appears twice. - Cost growing quarter over quarter with no new incident question becoming answerable. That last one is the honest measure. Instrumentation is an investment, and the return is measured in questions you can now answer under pressure. ## Keeping it from rotting Span names and attribute keys are an interface, not private detail — dashboards, saved queries and alert rules bind to them, and renaming one silently breaks things nobody will notice until an incident. Keep the names a service emits in one small module rather than as string literals scattered across the codebase, so the whole vocabulary of a service can be read on one screen. Give the handful that other teams depend on a test that fails if the key changes, so the rename is caught before it ships rather than during an outage. And put deletions on the same footing as additions: instrumentation that stopped answering a question is cost with no return, and removing it is legitimate work.

  • When is an attribute on an existing span the better choice than a new child span?
    Whenever you want to slice or find rather than to time. A new span is justified when the work needs its own duration or its own error status; identity, branch taken, batch size and variant do neither. They ride on a record already being emitted, which costs bytes rather than a whole extra span per request.
  • What are the signs a service has been over-instrumented by hand?
    Span names that mirror function names rather than operations, traces with hundreds of spans nobody ever expands, hand-written spans duplicating a boundary the agent already covers, and telemetry cost rising quarter over quarter without any new question becoming answerable during incidents.
  • How do you keep hand-written instrumentation from silently rotting?
    Treat span names and attribute keys as an interface: dashboards, saved queries and alert rules bind to them. Keep a service's names in one small module rather than scattered literals, pin the handful other teams depend on with a test that fails on rename, and delete instrumentation that has stopped answering anything.

saying these in an interview costs you the question

  • Wants a span around every method for completeness
  • Adds spans duplicating boundaries the agent already covers
  • Records durations but no domain identity at all
  • Names spans after functions rather than operations
  • Assumes more spans always mean better debugging
  • Never removes instrumentation once it stops being used