skip to content

At high request volume, how do you decide which LangSmith runs to trace at all?

level: principalimportance: should knowfreq 34%

answer

  1. a policy, not a magic percentage
  2. decision is made at the root run
  3. head sampling cannot know the outcome
  4. bias the sample toward risky paths
  5. payload size dominates ingest cost

basics

~20 s

Sample at the root: LANGSMITH_TRACING_SAMPLING_RATE keeps a fraction of traces whole, and tracing_context(enabled=...) lets code decide per request from cheap signals. Trace internal, canary and high-value traffic fully, sample the bulk, and route classes of traffic to separate projects.

solid answer

~50 s

There is no single right rate, so the answer is a policy rather than a number. The blunt instrument is `LANGSMITH_TRACING_SAMPLING_RATE`, a value between 0 and 1 applied at the root run so a kept trace is kept whole — you never get half a tree. The sharper instrument is `tracing_context(enabled=...)`, which lets a request decide from signals available at its start: internal or canary traffic always on, a new prompt version always on, a hot bulk endpoint at a few percent. Route those classes to different projects with `project_name` so investigation traffic is not drowned by volume. The honest limitation is that this is head sampling — the keep/drop decision precedes the outcome, so you cannot say "keep every failure"; pair tracing with metrics and logs that cover 100% of requests, and treat traces as the deep-dive sample rather than the system of record. Payload size, not run count, usually dominates ingest cost, so redacting or truncating bulk traffic is often a better lever than dropping it.

code

python · 18 lines
python
import os
from langsmith import traceable
from langsmith.run_helpers import tracing_context

os.environ["LANGSMITH_TRACING"] = "true"
os.environ["LANGSMITH_TRACING_SAMPLING_RATE"] = "0.05"


@traceable
def handle(request: dict) -> str:
    return "answer"


def serve(request: dict, is_canary: bool) -> str:
    if is_canary:
        with tracing_context(enabled=True, project_name="assistant-canary"):
            return handle(request)
    return handle(request)

go deeper

for a junior

Know that tracing everything is a choice, not a given, and that a sampling rate plus a way to turn tracing on or off per request both exist.

for a middle

Explain that the sampling decision is taken at the root so traces stay whole, and show the per-request switch and project routing that let different traffic classes carry different rates.

for a senior

Argue the policy: full coverage for canary and internal traffic, a low rate for bulk, full-coverage metrics and logs for failure detection, and a runtime override for incidents. Name head sampling's limitation before being asked.

for a principal

Own the budget and the governance — who sets rates, how a team requests more for a rollout, what content classes are traceable at all, and why redacting heavy payloads often beats dropping traces outright.

## Why sampling becomes a decision at all At low volume you trace everything and never think about it. Three things change that as traffic grows. **Cost**: ingest is billed and large-context RAG prompts are heavy, so a fully traced high-volume endpoint can outweigh the model spend it is observing. **Signal**: a project receiving millions of runs an hour is not more useful than one receiving thousands — it is less useful, because the interesting run is harder to find. **Exposure**: every traced request copies more content into a third-party store, and volume multiplies whatever privacy risk one request carries. So the question is not "can we afford it" alone; it is "what do we actually need to be able to reconstruct later". ## The two mechanisms **Global head sampling.** `LANGSMITH_TRACING_SAMPLING_RATE` takes a value from 0 to 1. The decision is made at the root run and applies to the whole tree, which is the property you want: a partially sampled trace, with the retrieval step present and the model call missing, would be worse than no trace. It is set per deployment, so different services or environments can carry different rates without code changes. **Programmatic per-request decisions.** `tracing_context(enabled=...)` wraps a block and switches tracing on or off for everything inside it. Because it is code, the predicate can use anything the request knows at its start: is this internal traffic, is this account on the canary prompt, is this a paying tier, is a debug flag set, did the previous turn in this conversation fail. Placed in middleware alongside the labelling context, it becomes the single place where the policy lives. **Project routing.** `project_name` on `tracing_context` or `@traceable` sends a class of traffic to its own project. This is underrated: separating "bulk production at 2%" from "canary at 100%" from "internal QA at 100%" means each project's monitoring charts describe a coherent population, and a filter over the canary project is not competing with a million background runs. ## The head-sampling limitation, stated honestly All of the above decide before the outcome is known. You cannot express "keep every trace that errored or exceeded three seconds", because at the moment of the decision the request has not run. This is the single most important thing to say out loud in an interview, because the naive answer — "sample 5% but keep all the failures" — is not something head sampling can deliver. What you do instead: - Keep **metrics and logs at full coverage**. Error rate, latency percentiles and token spend should come from your normal telemetry, which is cheap per event, not from the traced sample. Traces are for reconstructing individual executions. - **Bias the sample toward where failures live**: new prompt versions, new models, newly released code paths, accounts that have complained. A 100% rate on a canary that is 1% of traffic costs almost nothing and captures most of what you need. - **Make it reproducible.** If a specific request needs investigating and was not sampled, the fallback is replaying it with tracing forced on — which requires that you retained the inputs somewhere, and that replay is safe. If neither holds, raise the rate for that path instead. - **Allow a targeted override.** A per-request debug flag, or a rate you can raise at runtime without a deploy, converts a multi-hour investigation into a minutes-long one during an incident. ## Cost is mostly payload, not count A trace of a RAG request carrying twenty retrieved chunks is orders of magnitude larger than a trace of a short classification call. Before halving the rate — which halves your visibility — consider truncating or hiding the heavy fields on bulk traffic while keeping the run structure, timings and token usage. You retain latency analysis, error rates and cost attribution for 100% of requests and pay for content only where content matters. That is usually a better trade than dropping whole traces, and it is the move that separates a considered answer from a reflexive one. ## A workable default policy - Development and staging: 100%. - Internal, employee and QA traffic in production: 100%, own project. - Canary or newly deployed prompt/model paths: 100% until the rollout completes, own project. - Bulk production: a low single-digit percentage, content redacted or truncated, plus full-coverage metrics and logs. - An override switch that raises the rate for a named path or account without a deploy. Then revisit it, because the right rate is a function of current traffic, current spend and what your last three incidents actually needed — none of which are stable. ## The organisational half Someone has to own this. If each team sets its own rate against a shared budget, the budget is consumed by whoever traces most eagerly, not by whoever needs it most. Set the policy centrally — defaults per environment, a documented way to request more for a rollout, and a review when a project's volume changes materially — and make the default safe so that a team that thinks about none of this still ships something reasonable.

  • Why can't you configure "sample five percent but keep every failed trace"?
    Because the keep/drop decision is head-based: it happens at the root run, before the request has executed, so the outcome is unknown. Tail sampling would require buffering every trace until completion and deciding afterwards, which is not what this sampling switch does. The practical substitute is full-coverage metrics and logs for failure detection, plus a high rate on the paths where failures are likely.
  • Your ingest bill doubled but request volume was flat. What do you look at first?
    Payload size, not run count. A larger retrieval top-k, a longer system prompt, or a newly instrumented step that dumps whole documents into inputs will double stored bytes with identical traffic. Compare recent runs against older ones for the heavy fields, then truncate or redact those fields on bulk traffic before touching the sampling rate — you keep structure, timings and token usage at full coverage.
  • How do you investigate a specific failed request that was not sampled?
    Either replay it with tracing forced on for that call — which only works if you retained the inputs elsewhere and replay is side-effect free — or raise the rate for that path or account temporarily. Having a runtime override that does not require a deploy is what makes this a minutes-long fix instead of an incident-long one, and it is worth building before you need it.

saying these in an interview costs you the question

  • Claims head sampling can retain all failed traces
  • Sets one global rate and never revisits it
  • Cuts the sampling rate before looking at payload size
  • Mixes sampled bulk and full-fidelity canary in one project
  • Treats sampled traces as the system of record for error rates

context