skip to content

A team emits metrics, logs and traces at full fidelity. How would you set each signal's resolution rather than dropping one?

level: principalimportance: should knowfreq 46%

answer

  1. Resolution, not presence
  2. Each signal has its own dial
  3. Full fidelity answers unasked questions
  4. Bias what survives toward failures
  5. Replay an incident against what remains

basics

~20 s

Full fidelity on all three costs more than it answers, but dropping a signal leaves a class of question unanswerable. Set each signal's resolution separately: dimension breadth, per-event detail, kept fraction, sized by what it is asked on a bad day.

solid answer

~50 s

The wrong framing is "which of the three do we cut". Each shape answers a class of question the other two cannot, so removing one converts a cost problem into a blind spot. The decision is **resolution per signal**, and each has its own dial: metrics have a dimension budget and a collection interval, logs have a level and a field policy, traces have a kept fraction and the depth of what is instrumented. Set each dial from what the signal must answer under pressure, not from what it could capture. Then bias what survives toward the interesting: keep failures and slow requests at higher fidelity than healthy ones, since the aggregate already describes the healthy path. Keep it honest by replaying a real incident against what would have survived. If the on-call could still have reached the cause, the resolution is right; if not, raise that one dial.

code

pseudocode · 9 lines
pseudocode
budget the resolution, one dial per signal:

  metrics: dimensions <= 3 per metric, values bounded, interval 15s
           -> answers "is the 320 ms budget missed, where"
  logs:    full detail on failures; one thin record per success
           -> answers "what did this failing renewal do"
  traces:  keep 1 in 100 overall, keep errors and slow requests
           -> ~1,003 kept renewals/day; sees a 3% pathology,
              blind to a 0.02% one -> raise only if that matters

go deeper

for a junior

Recall that keeping everything is a choice with a price, and that each signal can be turned down independently rather than switched off. Knowing the three dials exist is enough at this level.

for a middle

Explain what each dial actually changes: dimension breadth and interval for aggregates, level and record detail for logs, kept fraction and instrumentation depth for traces. Say why keeping fewer traces does nothing for a metrics bill.

for a senior

Demonstrate that you have sized a dial from arithmetic rather than habit — how rare the event is, how many kept requests that implies — and that you have biased what survives toward failures and slow requests.

for a principal

Own the tradeoff. Argue that resolution, not presence, is the decision; defend a specific budget with the incidents that justify it; and be explicit about which questions you are choosing to make unanswerable and why that is acceptable.

## Why "collect everything" is not a policy "Emit everything at full fidelity and decide later" sounds like rigour and behaves like an unpriced default. Three things go wrong with it. - **The cost is paid on every event forever, and the value is realised on a handful of incidents.** Almost all of the fidelity is written, indexed and never read. The spend is certain; the payoff is occasional. - **Volume degrades the thing it was supposed to help.** Queries over a firehose get slower exactly when someone is trying to answer a question quickly, and the noise floor makes the interesting event harder to see, not easier. - **It removes the forcing function.** A team that never has to decide what a signal is for never discovers that half its emissions answer no question anyone asks. The opposite failure is just as bad: cutting a signal entirely because it is the largest line item. That converts a bill into a blind spot, and blind spots surface at 03:00 on the worst night of the quarter. ## Resolution is a per-signal dial | Signal | The dial | Raising it buys | Lowering it costs | |---|---|---|---| | Metrics | dimension breadth; collection interval | finer breakdowns; shorter spikes visible | slower detection; coarser attribution of a failure | | Logs | level, and how much each record carries | full per-event detail on more paths | the ability to reconstruct what one request did | | Traces | fraction of requests kept; depth instrumented | more examples of rare pathologies | the pathology may have no surviving example | The dials are independent, and that independence is the whole point. A metrics bill driven by dimension breadth is not helped by keeping fewer traces, and a trace bill driven by traffic is not helped by pruning dimensions. ## Setting the dials for a real service Take the municipal parking-permit service: a 320 ms p99 budget, an on-call rotation of three people, roughly 100,320 renewals a day. 1. **Start from the question each signal must answer under pressure.** Metrics answer *is the budget being missed, where, and how fast is it getting worse* — that requires an interval short enough to see a two-minute regression and dimensions for office, outcome and channel, not for permit identity. Logs answer *what did this failing renewal actually do* — that requires full detail on failures and much less on the 97% that succeed. Traces answer *which hop consumed the time* — that requires enough kept requests that a slow path has surviving examples, not that every renewal is kept. 2. **Set the dial to the smallest value that still answers it.** Keeping 1 request in 100 gives roughly 1,003 examples a day, which is plenty for a pathology affecting 3% of traffic and useless for one affecting 0.02%. That is an arithmetic decision, not a taste one: pick the rate from the rarity of what you need to see. 3. **Bias what survives toward the abnormal.** Keeping errors and slow requests preferentially costs a fraction of full fidelity and retains almost all of the diagnostic value, because the healthy path is already fully described by the aggregate. 4. **Cap the axis that can explode.** Give metrics a dimension budget per service and refuse dimensions whose value sets are unbounded. This is the only dial where a single careless change multiplies cost overnight. 5. **Re-derive the dials when the service changes shape.** A tenfold traffic increase changes the trace and log arithmetic and leaves the metrics arithmetic alone; a new failure mode changes what has to survive. ## The test that keeps it honest Take the last two real incidents and replay them against what the proposed resolution would have retained. Could the on-call have reached the cause with those aggregates, those records and that fraction of kept requests? This is the only check that is not a matter of opinion, and it usually finds exactly one dial set too low rather than a general shortage of data. Raise that dial, leave the rest alone, and write down which incident justified it so the next reviewer knows why the number is what it is. With an on-call rotation of only three people, the test has a second edge: fidelity nobody has time to interrogate at 03:00 is not fidelity, it is storage. A small rotation is an argument for signals that answer fast and for a shorter path from an alert to an example, and against a wide firehose that assumes a spare afternoon. ## What not to do - **Do not set one kept fraction for the whole estate.** The right fraction depends on traffic and on how rare the interesting event is; a busy service and a quiet one need different numbers to see the same thing. - **Do not cut resolution during the period when incidents are frequent.** That is when the data is being used, and it is the worst possible moment to discover a dial was too low. - **Do not equate more data with better observability.** Data you cannot query in the two minutes you have is not an answer. - **Do not treat the decision as permanent.** These are reversible numbers with a known cost; revisiting them after an incident should be routine and unremarkable.

  • What would make you raise a signal's resolution back up after lowering it?
    An incident that could not be explained with what survived — that is the evidence, and it names the specific dial. Also a genuine change of shape: a stricter latency budget, a new failure mode that is rarer than the current kept fraction can catch, or a traffic increase that moves the arithmetic. Raise the one dial, record the incident that justified it, and treat the change as reversible.
  • How do you keep these dials from becoming a permanent budget argument?
    Attach them to questions rather than to gigabytes. Each service owns a telemetry budget and states what each signal must answer; changes are argued as "this question became unanswerable", which is checkable, instead of "we need more data", which is not. Review the dials at incident review, where the evidence lives, rather than at a cost meeting, where only the bill does.
  • Is there a case for keeping full fidelity on one signal for a while?
    Yes, deliberately and with an end date. A newly launched path, a service in the weeks after a serious incident, or a migration where the failure modes are unknown all justify temporary full fidelity, because the value of an unanticipated question is genuinely high there. The discipline is that it is a decision with an owner and an expiry, not a default that nobody revisits.

saying these in an interview costs you the question

  • Proposes dropping traces entirely to reduce the bill
  • Treats storage price as the only cost of full fidelity
  • Applies one kept fraction to every service regardless of traffic
  • Argues more data always means better observability
  • Sets resolution once and never revisits it after incidents