How do you set a prompt and completion payload capture policy for LLM traces?
answer
- highest value, highest cost field
- attributes always, payloads sometimes
- sample the decision at the root
- errors captured at 100%
- payloads retained shorter than attributes
basics
~20 sCapture full prompt and completion payloads on a small sampled share of traces, plus always on errors and flagged turns. Keep attribute-only spans everywhere else, cap payload size, redact on the way out, and retain payloads for less time.
solid answer
~50 sPayloads are the highest-value and highest-cost field on an LLM span: they are what turns a wrong answer into a fixable bug, and they are also unbounded text that often contains personal data. So make it a policy rather than a per-service habit. A workable shape: always-on attributes everywhere — model, operation, latency, finish reason, retrieval counts, span structure — which already answer most latency and structure questions; **full payloads on a small percentage of traces**, 3% being a common starting point, so you always hold live examples; **100% capture on errors and on turns your application has already flagged as suspect**; hard truncation with the original length recorded, or store-by-reference with the span holding a pointer to a blob; a redaction hook in the export path so nothing raw is written; and shorter retention for payloads than for attribute-only spans. Note this is orthogonal to whether the trace is sampled at all — it is about what you write on a span you have decided to keep.
go deeper
Know that prompts and completions are optional, expensive extras on a span, and that teams deliberately capture only some of them rather than all.
Be able to explain what attribute-only spans still answer, and why capture is sampled and size-capped rather than unlimited.
Show the operational detail: root-level capture decisions, 100% on errors, truncation shape, store-by-reference, redaction before write, and a forced-capture switch for incidents.
Own the policy as a tradeoff between debuggability, cost and exposure, defensible per domain maturity and regulatory context, with a review cadence rather than a one-time setting.
## Why this needs a policy at all On a conventional service, span attributes are small and bounded. On an LLM span they are not: an assembled prompt can be tens or hundreds of kilobytes, one turn can hold several of them, and the text is exactly the material most likely to contain customer data. Left to each team, the outcome is bimodal — either nobody captures payloads and every investigation dead-ends at 'we cannot see what the model was told', or everybody captures everything and the observability bill and the risk surface both arrive as a surprise. Deciding it once, centrally, with an explicit rationale, is the principal-level move. ## What attribute-only spans already give you Before reaching for payloads, notice how much is answerable without them. Model and operation, latency per span, the finish reason, the number of agent iterations, retrieval candidate counts in and out, tool argument *shapes*, error statuses, span structure. That set answers latency regressions, loop and step-count problems, truncation, and most retrieval-shape bugs. Payloads are needed for one class specifically: 'the structure looks right and the answer is still wrong'. Sizing your capture to that class rather than to everything is the whole trick. ## The levers **Capture rate.** A small sampled percentage of traces gets full payloads. Sample the decision **at the root**, so a chosen trace is captured whole; sampling per span produces traces with payloads on some spans and holes on others, which is worse than none because the missing span is always the interesting one. Stratify by feature and tenant, because a flat percentage of skewed traffic yields nothing at all for the low-volume path that is failing. **Always-on exceptions.** Errors, and turns the application itself marked as suspect, should be captured at 100%. These are rare by definition, so they cost little, and they are the traces someone will actually open. **Size bounds.** Cap payload size, but pick the truncation shape deliberately: a plain prefix cut loses the tail, which is where the answer is, and it silently discards exactly the huge contexts you most wanted to inspect. Record the original length and a stable hash so you can tell a truncated capture from a short prompt, and prefer head-plus-tail truncation, or store-by-reference where the span holds a pointer into object storage and the large body lives outside the trace index. **Redaction on write.** Run payloads through a redaction step in the export path so raw text is never persisted. The detail of what constitutes sensitive content is a separate discipline; the tracing decision is that the hook exists, sits before the write, and fails closed. **Retention split.** Attribute-only spans can live for months cheaply; payloads should not. Different retention for the two is usually the single largest cost lever, and it also shrinks the blast radius of the store. **Tenant controls.** Some contracts forbid capturing content at all. That has to be expressible per tenant, and the default for a tenant with no explicit setting should be the conservative one. ## The tradeoff you should state plainly There is no setting that is right for everyone, and the honest framing is a three-way tension between debuggability, cost and exposure. A team shipping a new agent into an unfamiliar domain should capture aggressively and cut retention hard, because they do not yet know what they will need to look at. A mature, high-volume, regulated product should invert that: rich attributes everywhere, payloads on a thin sample, tight redaction, short retention. Both are defensible. What is not defensible is arriving at either by accident. ## Second-order effects worth anticipating Capture policy quietly determines what investigations are possible months later, so it deserves a review cadence rather than being set once. It also interacts with reproduction: if a failing turn was not among the captured sample, your fallback is to re-run the path with capture forced on, which requires that forcing capture per request is a supported operation rather than a redeploy. Build that switch early — it converts 'we have no data on this incident' into a ten-minute detour. ## Interview shape A strong answer names the levers, gives one concrete default (a low single-digit capture percentage, 100% on errors, payloads retained for days rather than months), explains the root-level sampling requirement, and is explicit that the right point on the curve depends on domain maturity and regulatory exposure. A weak answer is either 'capture everything, storage is cheap' or 'never store prompts' — both skip the actual decision.
- Why not just capture every payload and cut retention to two days?It is a legitimate option and some teams run it. The objections are that payload volume dominates ingest and index cost even briefly, that one store now holds every prompt the product has seen, which is a blast radius rather than a cost line, and that two days is shorter than the feedback loop for slow quality regressions — you notice in week three and the evidence expired. A thin permanent sample plus full capture on errors usually buys more for less.
- What breaks when you truncate captured prompts at a fixed size?The case you most need is the one that gets cut, because the biggest assembled contexts are exactly where context-construction bugs live. A plain prefix cut also loses the tail, which is often where the operative instruction sits. Record the original length and a hash so you can tell truncated from short, prefer head-plus-tail, and for anything large use store-by-reference so the span carries a pointer instead of the body.
- How do you make a small sampled share of captures representative?Decide at the root so a chosen trace is whole, then stratify: sample per feature, per tenant class and per entry point rather than uniformly over traffic. A flat percentage of a skewed mix gives you thousands of examples of the busiest happy path and none of the rare flow that is failing. Add a forced-capture switch per request so an investigation can produce data on demand.
saying these in an interview costs you the question
- Capturing every prompt and completion by default and finding the bill later
- Giving payloads the same retention as attribute-only spans
- Truncating to a fixed prefix and losing exactly the long contexts you needed
- Sampling capture per span, leaving traces with holes
- Writing raw payloads first and planning to scrub them afterwards