How do you keep a secret out of OPA's decision log when the policy input contains it?
answer
- the log is written after evaluation
- a second policy decides what is logged
- JSON pointers into the event
- remove or upsert a field
- the event lists what was erased
basics
~20 sWrite a mask policy, by default at data.system.log.mask, returning JSON pointers into the log event such as /input/... OPA erases or replaces those fields before the entry leaves the process, and records which paths it erased.
solid answer
~50 sThe decision log records `input` verbatim, so a rule that rejects secrets passed as plain environment variables ends up shipping those very values to the log sink. The fix is a **mask policy** — by default `data.system.log.mask`, overridable with the `mask_decision` setting. OPA evaluates it against the decision log event just before writing, and the rule returns JSON pointers into that event: either a bare string like `"/input/request/object/spec/containers"` to erase, or an object `{"op": "upsert", "path": ..., "value": ...}` to replace the value with a placeholder. Two traps: the event is the mask policy's input, so the original policy input sits at `input.input`; and a pointer matching nothing is silently a no-op, no error. Masking runs before buffering, so nothing unredacted reaches the sink, and OPA lists the removed paths in the event's `erased` field.
code
rego · 13 linespackage system.log
# Erase the container specs OPA was asked to judge for env-var secrets.
mask contains "/input/request/object/spec/template/spec/containers" if {
input.input.request.kind.kind == "Deployment"
}
# Or keep the shape and make the redaction visible to a reader.
mask contains {
"op": "upsert",
"path": "/input/credentials/token",
"value": "<redacted>",
}go deeper
Know that the decision log stores the input document as it was received, so anything sensitive in the request lands in the log unless you do something about it. Name the mask policy as the mechanism.
Be ready to write the rule: the default path, that the event is the mask policy's input so the original input is at input.input, and the difference between a bare pointer, a remove op and an upsert op.
Show you know it fails silently. Describe how you verify redaction against a real emitted event, and why an input schema you do not control makes that a test you keep rather than a one-off check.
Frame it as data classification: decide what the decision-log stream is allowed to contain, who reads it, and whether a control that must inspect secret material should be emitting per-decision records at all.
## Why this problem exists at all Take a concrete rule: deployments must not pass secrets as plain environment variables. To enforce it, OPA has to *see* the environment variables — the container spec, names and values, is exactly the material the rule inspects. And the decision log records `input` verbatim. So the moment you enable the plugin, the control designed to keep secret values out of manifests starts shipping those same values, decision by decision, to a central log service that a much larger set of people can read. The rule works and the logging defeats it. This is the canonical decision-log trap and it is why masking exists. Note that you cannot solve it by trimming the input upstream. On an admission path, OPA is handed the object under review; refusing to look at the env values means the rule cannot evaluate. ## The mask policy OPA evaluates a second, separate policy against the decision log event immediately before the event is written or buffered. By default it looks for `data.system.log.mask`; the `mask_decision` configuration setting points it elsewhere if you prefer another path. The crucial mechanical detail — and the one candidates most often get wrong — is **what the mask policy's input is**. It is not the original policy input. It is the whole decision log event. So the original input document appears at `input.input`, the result at `input.result`, and the paths the rule returns are JSON pointers rooted at the event: `/input/request/object/...`, not `/request/object/...`. Writing the pointer without the `/input` prefix is the standard bug, and it fails in the worst possible way — see below. ## Two shapes of mask entry The rule builds a set. Each member is one of: - **A plain JSON pointer string**, e.g. `"/input/credentials/token"`. OPA removes the value at that path. - **An object with an `op`**, e.g. `{"op": "remove", "path": "/input/credentials/token"}` or `{"op": "upsert", "path": "/input/credentials/token", "value": "<redacted>"}`. `remove` deletes; `upsert` sets the value at that path to what you supply, creating the field if it is not there. Which you want depends on the reader. `remove` leaves the field absent, which is unambiguous but can make a downstream consumer that expects the field misbehave. `upsert` with a fixed placeholder keeps the document shape intact and makes it visibly obvious to a human that redaction happened rather than the field never having been sent. Either way, OPA records the removed paths in the event's `erased` field, so a reader can tell the difference between "this was redacted" and "this was never present". Because the mask policy is Rego, it can be conditional: mask the container spec only for Pod-shaped objects, mask a token field only when the request came from a particular path. Keep those conditions simple. A mask policy that is itself intricate is a policy whose failure mode is a leaked secret. ## The silent-failure trap A JSON pointer that matches nothing in the event is not an error. OPA does not reject the mask policy at load, does not fail the decision, and does not warn. It writes the event unmasked. This means a typo, a `/input` prefix you forgot, or a schema that shifted under you all produce the same symptom: everything looks fine and the secret is in the log. So verify, and verify against a real event rather than by reading the rule. Enable the console sink in a non-production environment, evaluate a representative input, and look at the emitted JSON: is the value gone, and does `erased` name the path you intended? Treat that as a test you keep, because the input shape is not under your control — an upstream API version bump can move the field your pointer aims at. ## What masking does not do Masking affects **what is written to the decision log**, and nothing else. The policy still evaluates over the real value in memory; the value still arrived over the wire; anything else that can see the input — a `print` call left in a rule, a trace or explain output, a debugging endpoint — is unaffected. If your threat model includes the OPA process itself or its other outputs, masking is not the control you want. ## Masking versus dropping The adjacent knob is the **drop** policy (`data.system.log.drop` by default, settable with `drop_decision`): a rule that, when true for an event, suppresses that event entirely. The distinction matters and interviewers push on it. Masking removes a *value*, keeping the record that a decision with a given id, path, outcome, revision and timestamp happened. Dropping removes the *fact that the decision happened at all* — nothing in the sink distinguishes a dropped decision from one that was never made. Reach for masking when the problem is a sensitive field. Reach for dropping only when the problem is volume you have consciously decided you can afford to lose, and understand that you are choosing to have no record.
- What is the difference between erasing a field and dropping the whole event?Erasing keeps the entry: you still see the decision_id, the path, the outcome, the bundle revision and the timestamp, minus one value. A drop policy suppresses the event entirely, so nothing in the sink distinguishes a dropped decision from one that never happened. Mask when a field is sensitive; drop only when you have consciously decided you can lose the record.
- How would you verify a mask rule is actually firing?Turn on the console sink in a test environment, evaluate a representative input, and read the emitted JSON: the value should be gone and the `erased` list should name your path. Reading the rule is not verification — a pointer that matches nothing is a silent no-op, so a typo or a missing `/input` prefix looks identical to a working rule.
- Does masking protect the value inside OPA itself?No. The policy still evaluates over the real value in memory, and the value still arrived over the wire. Masking only changes what the decision log plugin writes. A `print` left in a rule, a trace or explain output, or any other debugging surface is unaffected, so masking is not a control against the OPA process or its other outputs.
saying these in an interview costs you the question
- Thinks masking changes what the policy evaluates
- Writes mask pointers without the /input prefix
- Assumes a non-matching pointer raises an error
- Drops the whole event to hide one field
- Relies on the downstream log sink to redact