Prompts traced to LangSmith contain PII — how do you keep content out of runs?
answer
- redaction happens before the run is queued
- two levers: client-wide and per function
- booleans drop, callables transform
- structure and timings survive content removal
- masking is not the same as residency
basics
~20 sStrip payloads before they leave the process: construct langsmith.Client(hide_inputs=True, hide_outputs=True) — or set LANGSMITH_HIDE_INPUTS/LANGSMITH_HIDE_OUTPUTS — to drop them wholesale, or pass process_inputs/process_outputs to @traceable to redact selected fields. Structure, latency, errors and token counts still ship.
solid answer
~50 sBy default these tools record the interesting thing, which is also the regulated thing: full prompts and completions. There are two levers, both in-process. Coarse: build the client as `Client(hide_inputs=True, hide_outputs=True)`, or set `LANGSMITH_HIDE_INPUTS`/`LANGSMITH_HIDE_OUTPUTS` in the environment, so payloads are removed before any run is queued. Both arguments also accept a **callable**, so you can transform rather than delete — hash an email, keep only a domain, keep the field names and drop the values. Fine-grained: `@traceable(process_inputs=..., process_outputs=...)` redacts per function, which is what you want when one step handles identifiers and the rest do not. Either way the run's shape survives: names, run types, tree structure, latency, errors and token usage are still recorded, so cost and performance work with no content at all. For residency rather than content, point `LANGSMITH_ENDPOINT` at the EU region or a self-hosted install.
code
python · 16 linesfrom langsmith import Client, traceable
client = Client(hide_outputs=True)
def redact(inputs: dict) -> dict:
email = inputs.get("email", "")
return {"email_domain": email.split("@")[-1]}
@traceable(client=client, process_inputs=redact)
def lookup_customer(email: str) -> dict:
return {"tier": "gold", "email": email}
print(lookup_customer("[email protected]"))go deeper
Know that prompts and completions are recorded by default and that there are switches — hide_inputs and hide_outputs on the client, or the matching environment variables — to stop that.
Explain that the hooks run in-process before anything is sent, that the arguments accept callables for transformation rather than deletion, and that process_inputs/process_outputs scope it per function.
Weigh what redaction costs you — losing readable prompts and the ability to promote traffic into datasets — and propose a split such as hidden bulk traffic plus a consented full-fidelity slice in its own project. Verify by reading back a stored run.
Own the classification decision: what content class may leave the network at all, masking versus self-hosting versus regional endpoint, retention, and who signs off. Make the control enforceable by environment rather than by developers remembering a decorator argument.
## The default is the problem The reason to trace an LLM application is to see the exact prompt that produced a bad answer. That is also exactly what a healthcare, finance or HR application is often forbidden to copy into a third-party store. So the question is never "should we trace" — it is "how much of the payload leaves the process", and the answer has to be enforced in code, not in a policy document. Crucially, all the redaction hooks below run **in your process, before the run is enqueued**. Nothing is sent and then deleted server-side; the content never crosses the network. That is the property a security review will ask you to demonstrate. ## Lever 1 — client-level hiding ``` from langsmith import Client client = Client(hide_inputs=True, hide_outputs=True) ``` Boolean `True` removes the payload entirely from every run that client ships. The equivalent environment switches are `LANGSMITH_HIDE_INPUTS` and `LANGSMITH_HIDE_OUTPUTS` set to `"true"`, which is the deployment-time version of the same decision and the one to prefer when the setting must differ per environment — full content in development, hidden in production. Both arguments also take a **callable** receiving the payload and returning the version to record. That turns an all-or-nothing switch into a transformation: keep field names but blank values, replace an email with its domain, hash a customer id so runs are still correlatable without being identifying, truncate long free text. A callable is usually the better engineering answer than `True`, because a run with structure but no values is far more debuggable than a run with a hole in it. Remember that the traced code must actually use that client. `@traceable(client=client)` binds a decorated function to a specific client; otherwise the default global client applies, which is where the environment variables earn their keep — they cannot be bypassed by a code path that forgot to pass a client. ## Lever 2 — per-function processing ``` @traceable(process_inputs=redact, process_outputs=trim) ``` `process_inputs` receives the mapping of the call's arguments and returns what to record; `process_outputs` does the same for the return value. This is scoped surgery: the one function that takes a customer record redacts it, while retrieval and generation steps keep full fidelity. It composes with client-level hiding — you can hide globally and selectively re-add a safe summary. A practical warning: these callables run on your request path. Keep them cheap and total. An exception or a slow regex inside a redaction function is a production incident caused by your observability layer, so unit-test them like any other code and prefer explicit allow-lists ("record these three fields") over deny-lists ("remove anything that looks like an email"), which fail open. ## What survives redaction With payloads removed you still get: run names and types, the tree structure, timings, error status and stack information, tags and metadata, and token usage from provider wrappers. That means latency analysis, error-rate monitoring, cost attribution and the shape of an agent's tool-calling loop all keep working. What you lose is the ability to read what was said — and, downstream, the ability to promote real traffic into evaluation datasets, because a redacted run has no example to promote. That tradeoff is the whole conversation: teams typically resolve it by hiding content on the bulk of production traffic while keeping a consented or synthetic slice at full fidelity, routed to its own project. ## Residency, which is a different problem Redaction controls *what* is stored. Residency controls *where*. `LANGSMITH_ENDPOINT` points the SDK at a different deployment — the EU region, or a self-hosted install inside your own network. If the requirement is "data must not leave our infrastructure", no amount of field-level masking satisfies it and self-hosting is the answer; if the requirement is "no personal data in third-party systems", masking may be enough. Get the requirement stated precisely before choosing, because the two have very different operational costs. ## Don't forget metadata Redaction hooks cover inputs and outputs. Metadata is a separate channel that you populate yourself, and it is a common leak: a middleware that dumps the whole authenticated user object, or the raw request body, into run metadata puts personal data straight into the store while inputs are dutifully hidden. Whitelist the metadata keys you set, and review them alongside the redaction functions. ## Verify, don't assume After wiring any of this, generate a run with known synthetic PII and read the stored run — through the UI or by pulling it back with the client. Redaction that was configured on a client the traced code does not use, or a decorator that a later refactor dropped, fails silently and looks exactly like success from inside the application.
- Why prefer a callable over hide_inputs=True?Because a run with structure but blanked values is still debuggable — you can see which fields were present, how long the input was, and where the shape differs between a working and a failing call. `True` leaves a hole, and the first time you investigate an incident you will wish you had kept the keys. A callable also lets you retain a hashed identifier so runs stay correlatable without being identifying.
- What still works once you hide all prompt and completion content?Latency and error analysis, the trace tree and tool-call sequence, tags and metadata filtering, and token/cost attribution from provider wrappers, since usage numbers are not part of the payload. What stops working is reading what was actually said — and promoting that traffic into evaluation datasets, since there is no example content left to promote.
- A regulated team says data cannot leave their infrastructure at all. Is masking enough?No. Masking controls what is in the payload, not where the payload lands; run metadata, names, timings and any residual free text still go to the vendor's servers. That requirement is a deployment decision — point LANGSMITH_ENDPOINT at a self-hosted install so the data stays inside your network, and accept the operational cost of running it.
saying these in an interview costs you the question
- Assumes content is redacted server-side after upload
- Configures a client the traced code never uses
- Forgets metadata is a separate leak channel
- Thinks hiding inputs also removes token and cost data
- Treats field masking as satisfying a data-residency requirement