skip to content

How do you set a content-capture policy for gen_ai.prompt attributes across environments?

level: principalimportance: should knowfreq 28%

answer

  1. Classify the traffic before configuring anything
  2. Default is on, so unset means captured
  3. No per-request control, so split deployments
  4. Give the closed tier a substitute
  5. The store inherits the highest class you allow

basics

~20 s

Classify each service's data, then decide per deployment whether spans may hold message text at all. Because the capture switch is process-wide and defaults to on, enforce it in deployment configuration, and give services that cannot capture a reference-based alternative.

solid answer

~60 s

Start from data classification, not from tooling: for each service, what class of text passes through the model, and is the trace store an approved destination for it? That question has three answers in practice — capture freely (synthetic or internal data), capture redacted, or never capture. Because OpenLLMetry captures message content by default and the toggle is process-wide, the policy must be enforced where deployments are defined rather than left to a developer's environment, and unset must be treated as a violation since it fails open. Services in the "never" class still need debuggability, so give them a substitute: an application-side record id on the span pointing at a store that already has the right retention, access control and deletion path. Then set retention and access on the trace backend to match the most sensitive content you permit, sample full payloads rather than keeping all of them, and rehearse the incident path — capture is not retroactive, so accidental capture is a data-handling incident, not a config typo.

go deeper

for a junior

Understand that whether prompts may be recorded is a policy decision made per service, not a personal preference, and that the safe assumption is that capture is on unless configuration says otherwise.

for a middle

Be able to implement the decision: set the capture variable explicitly in each environment's configuration, verify on a real span, and know that there is no per-request granularity to fall back on.

for a senior

Argue the enforcement design — fail-open default, deploy-time checks, separate deployments per data tier — and provide the substitute for services that cannot capture, such as a record id on the span pointing at a governed payload store.

for a principal

Own the whole frame: data classification per service, retention and access requirements the trace store must meet, hosted versus self-hosted residency, sampling economics, and a rehearsed incident path for accidental capture.

## Why this is a policy question, not a settings question The mechanics are trivial: one environment variable turns message capture on or off. What makes it a leadership problem is that the default is on, the granularity is a whole process, the failure mode is silent, and the consequence lands on a system — the trace backend — that was probably never classified as a store of customer content when it was procured. Someone has to decide, per service and per environment, what is allowed, and then make that decision hard to get wrong. ## Step 1: classify the traffic, not the tool For each service that calls a model, ask what the prompt actually contains: synthetic fixtures, internal employee text, customer-authored text, regulated categories (health, payment, identity), or text belonging to a customer under a contract that restricts sub-processors. The answer places the service into a tier: - **Open**: capture full content. Development, staging with synthetic data, internal tooling. - **Restricted**: capture with redaction or hashing, or capture on a small sample. - **Closed**: never place content on a span. This tiering is the artefact worth writing down, because it survives tool changes. If you replace the tracing vendor next year, the classification still holds. ## Step 2: make the enforcement structural A process-wide switch that defaults to capture means "forgot to set it" equals "captured it". Treat that like any other fail-open default: - Set the variable explicitly in every deployment manifest, including the closed and open tiers, so the intent is visible in review rather than inferred from absence. - Add a check that fails deployment for a closed-tier service if the value is missing or true. - Prefer separating tiers into different deployments over trying to branch within one process, since there is no per-request control. - Verify empirically after rollout by reading a real span, not by trusting the manifest. ## Step 3: give the closed tier something A policy that only takes things away gets circumvented. Engineers turn capture on "temporarily" to debug an incident and it stays on. Provide the substitute up front: - Put an application-generated request or record id on the span, and keep the payload in a store that already has retention limits, access logging and a deletion workflow. - Or redact in a span processor before export — mask detected identifiers, or store a hash so you can tell whether two requests carried the same prompt without holding the text. - Or reproduce in a staging environment with capture on and synthetic inputs. The point is that the trace links to evidence rather than being the evidence, which also moves the compliance burden to a system built for it. ## Step 4: set the store's controls to the highest class you allow The trace backend's retention, access model and residency now apply to whatever content you permit anywhere. That drives concrete asks: short retention on content-bearing traces, role-based access separate from general dashboard access, a supported deletion path for a subject-access or erasure request, and a decision on hosted versus self-hosted for residency. If the backend cannot delete individual traces, the closed tier cannot be relaxed later — worth knowing before, not after. ## Step 5: sample deliberately Even in the open tier, keeping every payload forever is neither cheap nor useful. Payload bytes usually dominate LLM tracing cost, and the debugging value is concentrated in failures and in a representative slice. A defensible default is metadata on everything, full content on a small percentage plus all errors, with a time-boxed override that raises capture for a specific session while an incident is open. Time-boxing matters: the override should expire on its own rather than relying on someone remembering. ## Step 6: rehearse the incident Capture is not retroactive. If content lands somewhere it should not have, changing configuration only stops future spans; the stored data is a separate problem. The runbook needs: how to detect it (a check that queries for content attributes in closed-tier services), who is notified, how affected traces are deleted or aged out, who had access in the meantime, and what the disclosure obligation is. Having this written before it is needed is the difference between a contained incident and a scramble. ## The tradeoffs to name out loud - **Debuggability versus exposure.** Content is the fastest path from "the answer was wrong" to "here is why", and giving it up genuinely slows incident response. Say so rather than pretending the closed tier is free. - **Cost versus completeness.** Full payloads are the majority of the bill; sampling trades rare-case visibility for a large saving. - **Central store versus isolation.** One trace backend for everything gives correlation with the rest of your telemetry; isolating LLM content into a separate, tightly controlled store simplifies compliance and complicates debugging. - **Vendor versus self-host.** Self-hosting removes the sub-processor question and adds an operations burden you must staff. ## What a strong answer sounds like Not "we set the env var to false", but: here is the tiering, here is where it is enforced so it cannot be forgotten, here is what the restricted tier gets instead of raw text, here is what the store's retention and access must be to hold the most sensitive class we allow, and here is the runbook for the day it goes wrong.

  • A team asks to turn content capture on in production for one week to chase a quality bug. How do you answer?
    Ask which tier the service is in. For an open tier, approve with an expiry and a named owner. For restricted or closed, offer alternatives first: reproduce in staging with synthetic inputs, capture a small sampled slice, or capture hashes and ids and pull the payload from the governed store. If you must relax it, time-box it in the manifest so it reverts on its own, notify whoever owns the data classification, and shorten retention for the window.
  • Why is enforcing this in deployment manifests better than documenting it as a convention?
    Because the default is capture. A convention is only checked when someone remembers; an absent variable silently means "record customer text". Putting the value in the manifest for every tier makes intent reviewable in a diff, lets a policy check fail the deploy when a closed-tier service is missing or contradicts it, and leaves an audit trail of when it changed and who approved it.
  • What do you demand from a tracing backend before allowing any content capture?
    Retention you can set short and independently for content-bearing data, access control separate from general dashboard access, a supported path to delete specific traces for an erasure request, residency that matches your obligations, and clarity on sub-processor status. If it cannot delete individual traces, you cannot honour a deletion request, which caps the data classes you may ever allow into it.
  • What does the organisation actually give up by closing content capture on its most sensitive service?
    Speed of diagnosis. Metadata tells you a call was slow, expensive or errored; only the payload tells you the retrieval returned the wrong document or the template lost a variable. Expect longer incident timelines and more reliance on staged reproduction, and budget for the substitute — a governed payload store and the tooling to join it to traces — rather than pretending the loss is free.

saying these in an interview costs you the question

  • Leaves the capture setting unset and assumes it is off
  • Tries to switch capture per user inside one process
  • Classifies the tool instead of the data
  • Assumes disabling capture cleans up traces already stored
  • Sets no retention limit on content-bearing traces

context