How do you keep API keys and credentials out of LLM prompt logs and traces?
answer
- the trace is a full copy of the context
- the model never needs the key itself
- redact before the write, not on read
- hash the match to keep it correlatable
- a redaction hit means fix something upstream
basics
~20 sKeep secrets out of the prompt in the first place by injecting them at the tool layer, then redact at the write path: match known credential formats before a trace record is stored, replace matches with a hash or placeholder, and restrict who can read the trace store.
solid answer
~50 sA trace stores the whole context window, so anything that reached the prompt is duplicated into a system whose readers are usually far broader than the source system's. Two controls, in order. First, prevention: a model never needs a credential to call a tool, because the tool runtime holds the key and attaches it to the outbound request, so the secret should never enter the context at all. Second, containment: run a redactor at the write path, before the record is persisted, matching the credential shapes you actually issue plus common third-party key formats and authorization headers. Replace each match with a fixed placeholder or a keyed hash, so records stay correlatable without holding a usable secret. Then restrict trace access and alert on redaction hits, because a hit means a secret reached a prompt and the real fix is upstream of the log.
code
python · 10 linesimport hashlib
import re
KEY_PATTERN = re.compile(r"\bsk_live_[A-Za-z0-9]{24,}\b")
def redact(text: str) -> str:
def replace(match: re.Match) -> str:
digest = hashlib.sha256(match.group(0).encode()).hexdigest()[:12]
return "[redacted-key:" + digest + "]"
return KEY_PATTERN.sub(replace, text)go deeper
Know that traces store the whole prompt, that the model never needs the actual key because the tool runtime attaches it, and that redaction happens before anything is written.
Explain what to match and what to write in its place, why a keyed hash preserves correlation without keeping a usable secret, and why the redaction hit rate is itself a signal to act on.
Cover the upstream fixes that stop secrets reaching the context, access scoping on the trace store, applying redaction again on eval exports, and rotating any credential that reached the store.
Set the policy that credentials never enter model context, own where redaction sits relative to every store that copies a trace, and make redaction hit rate a tracked signal rather than background noise.
## Why prompt logs are a distinct problem Every serious LLM application logs prompts and traces, because you cannot debug or evaluate a system whose behaviour you cannot replay. The consequence is a store that contains, for every request, the entire assembled context: system prompt, user message, retrieved documents, tool definitions, tool results and the model's output. It is the most complete copy of your application's working data that exists anywhere, and it typically lives in a tracing or observability tool where read access is granted to the whole engineering organisation rather than to the narrow group entitled to the underlying records. Credentials end up in that store more often than people expect. A user pastes an API key into a support chat while describing their problem. A tool result echoes back a request that included an authorization header. A configuration file retrieved as context contains a connection string. A stack trace captured as a tool error includes a token in a URL. None of these require an attacker; they are ordinary application behaviour, and each one writes a live secret into a searchable log. ## Prevention comes first The better control is upstream. A model does not need a credential to use a tool, because the model does not make the outbound request; the tool runtime does. Keep keys in the runtime's configuration or secret store, attach them when the request is built, and never place them in a tool definition, a system prompt or a tool argument. If a design requires the model to pass a token around, that design has an easy fix: give the tool a handle or session reference and let the runtime resolve it to the real credential. Similarly, tool results should be shaped before they enter the context. Returning a raw HTTP exchange, including headers, is convenient and is how authorization headers reach prompts. Return the fields the model needs. ## Redaction at the write path Containment still matters, because prevention is never complete and users will paste things. The rule is that redaction runs before the record is written, not on the way out to a dashboard and not in a nightly cleanup job. Once a live key is in the store it has been exposed to everyone with read access and to any downstream export, and a later scrub does not undo that. What to match: the credential formats you issue yourself, which is why issuing keys with a recognisable, high-entropy prefix pays for itself here; common third-party key shapes; authorization headers and bearer tokens; private key blocks; connection strings with embedded passwords. Pattern matching is imperfect in both directions, so pair it with a shape-based check such as high-entropy strings in fields that should not contain them. What to write instead: a fixed placeholder is simplest. A keyed hash of the matched value is often better, because it lets you correlate that the same secret appeared in fifty traces, or match against a known leaked value, without storing anything usable. Storing a truncated prefix alongside the hash helps humans identify which key it was without recovering it. ## Treat a redaction hit as a signal The most valuable output of the redactor is not the redacted log; it is the count. Every hit means a real credential reached a real prompt, which means the prevention layer failed somewhere: a tool returning raw headers, a document ingested with a config file in it, a user flow that invites pasting secrets. Alert on hits, attribute them to a code path, and fix upstream. A redactor whose hit rate is quietly climbing is a system leaking secrets into context faster than anyone is noticing. ## Access and blast radius Beyond redaction, the trace store deserves the access control its contents warrant: authenticated access, role-scoped rather than org-wide reads, and audit of who queried what. If traces are exported to an evaluation dataset or replayed against a third-party judge, that export is a second copy with its own boundary and needs the same redaction applied, not merely inherited. ## Rotation is the backstop If a live credential did reach the store before redaction existed, treat it as disclosed and rotate it. This is the point where a candidate either understands secrets or does not: a secret that has been written to a broadly readable system is compromised regardless of whether anyone is known to have read it, and the only remediation that restores the property you cared about is issuing a new one.
- Why not just scrub secrets out of the trace store with a nightly job?Because exposure happens at write time. Between the write and the scrub, the live credential is readable by everyone with access to the tracing tool and by any export, replication or alert that copied the record. A nightly job also cannot un-send what was already viewed. Redaction belongs in the write path, with the nightly sweep as a check that the write-path redactor is working.
- Your redactor's hit rate doubled after a release. What does that tell you?That a code path started putting credentials into prompts. Common causes are a tool that began returning raw HTTP exchanges including authorization headers, an ingestion job that swept configuration files into the retrievable corpus, or a new user flow that invites pasting keys. The redactor contained the symptom; the fix is upstream, and the hit rate is the metric that made it visible.
- Does a key that reached the trace store need rotating if nobody read it?Yes. A secret written into a broadly readable system is disclosed whether or not you can prove someone looked, and access logs rarely settle the question. Rotate it, then fix the path that put it there. Treating unverified non-exposure as safety is the reasoning that turns a small incident into a long one.
saying these in an interview costs you the question
- Passing API keys to the model so it can call a tool
- Redacting only when traces are displayed, not when written
- Relying on a nightly cleanup job over stored traces
- Believing a hash of a key is reversible or is encryption
- Leaving a leaked key in place because nobody is known to have read it