Telemetry spend is doubling every quarter and legal wants user email addresses removed from span attributes. How would you use OpenTelemetry Collector processors — filter, transform (OTTL), attributes — and connectors to shape volume and scrub data, and what breaks if you get it wrong?
answer
- Scrub at the agent (never leaves the node), shape at the gateway
- OTTL contexts: resource / span / spanevent / metric / datapoint / log
- Allow-list of attribute keys beats deny-list (closed vs open world)
- Measure with a count connector before enabling a drop rule
- Prune attributes → drop junk → aggregate → sample, in that order
basics
~20 sScrub PII at the agent with transform/OTTL or redaction so raw values never leave the node; shape volume centrally with filter, attribute pruning and aggregation connectors. Measure before dropping, and never break trace continuity or metric series continuity.
solid answer
~50 sTwo different problems with two different placements. **PII** belongs at the agent tier: an OTTL `transform` statement (`delete_key`, `replace_pattern`, or hashing) or the `redaction` processor removes or masks the value before it crosses the node boundary, and an allow-list of permitted attribute keys ages better than a deny-list of known-bad ones. **Volume** belongs at the gateway, where you can see the whole picture: use a `count` connector to measure what each rule *would* drop before enabling it, then apply `filter` to remove genuinely worthless records (health-check and readiness spans, debug-severity logs), prune fat attributes, and replace raw span retention with aggregation — a `spanmetrics` connector keeps rate/error/duration on 100% of traffic while sampling keeps only exemplar traces. The failure modes are specific: filtering spans mid-trace leaves orphaned children and broken waterfalls; dropping metric data points breaks a cumulative series' continuity so rate calculations produce gaps or false resets; deleting an attribute silently breaks every dashboard, alert and saved query that grouped by it.
code
text · 10 linesAGENT pipeline (per node):
otlp -> transform (delete_key(attributes,"user.email");
replace_pattern(attributes["http.url"], "[0-9]{9,}", "***"))
-> k8sattributes -> batch -> otlp/gateway
GATEWAY pipeline:
otlp -> spanmetrics(connector) # RED metrics on 100% of spans
-> filter (drop spans where name == "GET /healthz")
-> tail_sampling # keep errors, slow, 1% baseline
-> batch -> otlp/vendorgo deeper
Know that the Collector can delete or mask attributes and drop records, and that this is config rather than code changes in the app.
Name the processors (filter, transform/OTTL, attributes, redaction) and the ordering rule that dropping should precede batching.
Place scrubbing at the agent and shaping at the gateway, measure before cutting, and articulate the trace-continuity and cumulative-series failure modes.
Own it as governed policy: allow-list posture, rule ownership and review, staged rollout, drop counters for auditability, and an explicit statement of the spend-versus-answerability trade-off.
## Separate the two goals Cost shaping and data scrubbing get lumped together because both are "processors that remove things", but they have different placements, different owners and different failure modes. Decide them separately. ## Scrubbing: as early as possible The defensible position is that sensitive values never leave the machine that produced them. That argues for the **agent tier**: a node-local Collector applies the rule before anything crosses a network boundary, so a gateway compromise or a vendor misconfiguration cannot expose what was never sent. The tools: the `transform` processor runs OTTL statements scoped to a context (`resource`, `span`, `spanevent`, `metric`, `datapoint`, `log`), so you can `delete_key(attributes, "user.email")`, `replace_pattern(...)` to mask, or hash a value where you still need to correlate on it without reading it. The `attributes` processor offers a simpler delete/update/hash surface, and the `redaction` processor supports an allow-list of permitted keys plus value-pattern blocking. The design point that matters: prefer an **allow-list** of attribute keys to a deny-list. A deny-list is open-world — it stops the leaks you have already thought of, and every new instrumentation library or hand-written attribute is a fresh chance to add a field nobody reviewed. An allow-list is closed and owned by the config, at the cost of friction when teams add legitimate attributes. Whichever you pick, the rule is fleet-wide config, not something each team re-implements. Remember what scrubbing cannot reach: an email embedded in a span *name*, in a URL path captured as `http.route`, in a log message body, or in an exception stack trace. Those need their own statements, and high-cardinality identifiers in span names quietly wreck backend indexing too. ## Shaping volume: measure, then cut The first move is not a filter, it is a measurement. A `count` connector (or a filter running in a parallel pipeline to a `debug`/`nop` exporter) tells you how many records a candidate rule matches, so you can rank rules by bytes saved rather than by intuition. Cost concentrates in surprising places: a chatty health-check endpoint, a client library spanning every retry, a debug log level left on in one service, a metric labelled with a request id. The ordered menu of interventions, cheapest first in terms of information lost: 1. **Prune attributes**, not records. Fat, low-value attributes (full SQL text, whole request headers, base64 payloads) often dominate bytes. You keep every span and lose some detail. 2. **Drop genuinely worthless records.** Readiness/liveness probe spans, `DEBUG`-severity logs in production, metrics nobody queries. `filter` with OTTL conditions does this. 3. **Aggregate instead of storing.** A `spanmetrics` or `servicegraph` connector converts spans into rate/error/duration metrics and service topology — a tiny fraction of the bytes, computed over 100% of traffic. Then keep only a sample of raw traces as exemplars. This is the single biggest lever in most estates. 4. **Sample**, last, because it is the only step that removes evidence you might need. ## What breaks **Trace continuity.** Dropping spans by condition is not the same as dropping traces. Filter out a middleware's spans and every child of those spans becomes an orphan referencing a missing parent; backends render broken waterfalls, and some drop the orphans entirely. Drop whole traces (sampling) or drop leaf spans that have no children — never a span in the middle. **Metric series continuity.** A cumulative series carries a start timestamp and monotonically increasing values; consumers detect resets by watching for a value decrease or start-time change. Intermittently dropping data points from a cumulative series produces gaps that rate functions interpolate over, and dropping and later resuming a series can look like a reset. Filter metrics by *whole series*, deterministically, not by conditions that flip over time. **Downstream queries.** Deleting an attribute breaks every dashboard, alert rule and saved query that grouped or filtered by it — silently, because the query still runs and simply returns fewer or unlabelled series. Attribute removal needs the same change-management as an API break: announce, measure usage where the backend supports it, stage it. **Reproducibility of the past.** Volume rules change what the archive contains, so historical comparisons across a rule change are invalid. Record when each rule shipped where incident reviewers will see it. ## Governance Because this is central config that silently changes what everyone can see, treat it as a reviewed artifact: version-controlled, code-reviewed by the owning teams, rolled out gradually (agent tiers hit every node at once), and each rule carrying an owner and a rationale. Pair every drop rule with a counter of what it dropped, so "why is this missing?" has an answer that does not require reading the config. And revisit annually — rules written for a service that no longer exists quietly cost nothing and confuse everyone. ## The trap to name explicitly Aggressive shaping optimises the bill and degrades exactly the data you need during the incident that justifies the bill. State the trade-off in the terms the business understands: the goal is not minimum telemetry, it is the minimum spend that still answers the questions you will actually ask under pressure.
- Why is dropping individual spans riskier than dropping whole traces?A trace is a tree held together by parent references. Removing a span in the middle leaves its children pointing at a parent that no longer exists, so backends either render a broken waterfall or discard the orphans, and the timing structure becomes unreadable. Sampling removes the whole tree consistently, which keeps every retained trace intact. If you must filter spans, filter leaves with no children.
- An allow-list of attribute keys or a deny-list of sensitive ones — which do you choose and why?An allow-list, because it is closed-world: only keys the configuration names can pass, so a newly added attribute is excluded by default and a library upgrade cannot introduce a leak silently. A deny-list is open-world and only blocks what someone already anticipated. The cost is friction — teams must register new attributes — which is a review burden you can plan for, unlike an unplanned disclosure.
- How do you avoid losing incident-response capability while cutting spend?Preserve aggregate fidelity and cut sampled detail: keep rate, error and duration metrics computed over 100% of traffic via a connector, keep all error and slow traces via tail-sampling policies, and take the reduction from the ordinary successful traffic that nobody queries. Then validate against real incidents — walk a past investigation using only the data the new rules would have kept.
saying these in an interview costs you the question
- Filtering spans in the middle of a trace and expecting the waterfall to survive
- Using a deny-list of sensitive attribute keys and calling the problem solved
- Dropping metric data points conditionally and breaking cumulative series continuity
- Cutting volume without first measuring what each rule would remove
- Scrubbing PII only at the central gateway, after the raw value has already crossed the network
- Assuming attribute deletion is invisible — it silently empties dashboards and alert rules