skip to content

What fields should a policy violation record carry beyond the human-readable message?

level: middleimportance: should knowfreq 48%

answer

  1. the sentence is for the person, not the pipeline
  2. what a review bot needs to place a comment
  3. join key: rule id plus resource plus path
  4. grouping by message string splits one rule
  5. one record per violation, additive fields

basics

~20 s

A violation record needs discrete fields other systems can read: a stable rule identifier, severity, the outcome, the resource identity, the offending field path, and a remediation hint. Prose cannot be filtered, aggregated or placed on a line.

solid answer

~40 s

The message is for the person; the record is for everything else. A usable record carries a stable rule id (so findings can be counted, suppressed and looked up), a severity, the outcome the engine chose (denied, warned, exempted), the identity of the resource — kind, namespace and name, or file and line for a source document — the offending field path, the human message, and a remediation hint. Those fields are what let a pull-request bot annotate the exact line, a dashboard group by rule instead of by sentence, and metrics answer "which rule blocks people most". If the only machine-readable thing is the message, every consumer ends up regex-parsing prose, and the payload is not a contract — it is a rendering that other systems have quietly become dependent on.

code

json · 13 lines
json
{
  "ruleId": "container-memory-limit-required",
  "severity": "high",
  "outcome": "deny",
  "resource": {
    "kind": "Deployment",
    "namespace": "payments",
    "name": "checkout-api"
  },
  "path": "spec.template.spec.containers[2].resources.limits.memory",
  "message": "container 'log-shipper' declares no memory limit",
  "remediation": "set resources.limits.memory (team default: 512Mi)"
}

go deeper

for a junior

Know that a gate emits two different things — a sentence for you and a structured record for other tools — and be able to name a few fields that record should carry.

for a middle

Explain what each field buys a specific consumer: field path for a review bot's inline comment, rule id for grouping, outcome for measuring whether a rule is ready to enforce.

for a senior

Show you design the record as an interface with named consumers, keep it one record per violation, and add fields rather than repurposing them.

for a principal

Frame the record as the integration point the whole guardrail programme reports through; without stable fields, every question about coverage and impact turns into bespoke text parsing.

## The message is a rendering; the record is the contract When a policy engine denies a change it produces two things that are easy to conflate. One is the **message** — a sentence written for the person reading it. The other is the **violation record** — a structured object describing what failed, which is what every other system consumes. Teams that only ever look at their own terminal build the first and skip the second, and then discover that everything downstream has to reverse-engineer the sentence. ## The fields, and what each one buys **Rule identifier.** A short, stable id for the rule that fired — `container-memory-limit-required`, not "the memory rule". This is the key everything else joins on: counting how often a rule fires, filtering a report to one rule, correlating today's finding with last week's, and referring to it in a conversation. Without an id, the only handle anyone has is the sentence. **Severity.** A rating the emitting policy assigns. It lets a consumer triage — surface the high ones inline in the review, roll the low ones into a weekly report — without re-implementing that judgement per consumer. **Outcome.** What the engine actually did about this violation, which is not always "blocked". The same rule may be enforcing in one environment and only observing in another, and a record that does not say which one produced it cannot distinguish "this stopped a change" from "this would have stopped a change". Any counting of gate impact depends on this field. **Resource identity.** *Which* thing failed. For a cluster object: kind, namespace, name, and API group/version. For a source document: repository, file path and line. A change can carry dozens of objects; without identity the consumer cannot say which one, and cannot deduplicate across runs. **Field path.** *Where inside that thing.* `spec.template.spec.containers[2].resources.limits.memory` is what lets a review bot post its comment on the offending line rather than at the top of the file, and what lets a person fix a five-container manifest in one edit. This is the field most often left out and the one that most changes the experience. **Message.** The human sentence. Still needed — it is what the developer reads — but it is one field among several rather than the whole payload. **Remediation hint.** What to change, ideally naming the field and the value or where to get it. Machine-readable remediation is also what lets tooling offer a suggested fix rather than only a complaint. A record built from those fields looks like the example attached to this question. Note the shape: identity and location are discrete fields, and the prose is one of them. ## Why prose alone fails Every consumer of a denial wants something the sentence cannot give it. - A **pull-request bot** wants a file and line, or a resource and field path, so it can annotate the change where the problem is. From prose it can only paste a wall of text into one comment. - A **dashboard** wants to group by rule and severity to answer "which rules fire most and where". Grouping by message string splits one rule into as many buckets as it has message variants — because a good message embeds the specific container name, every occurrence is a distinct string. - A **metrics pipeline** wants counts by rule and outcome over time. Prose gives it one unusable dimension. - **Deduplication and suppression** want a stable key. The natural key is (rule id, resource identity, field path). Any consumer that has to build that key by parsing a sentence is building on sand. The deeper problem is that when prose is the only machine-readable thing, consumers *do* parse it. They write a regular expression against the message, it works, and now the wording of that sentence is load-bearing infrastructure that nobody documented as such. ## Two shapes worth getting right **One record per violation, not one per evaluation.** A change that trips the same rule on three containers should produce three records with three field paths, not one record with three sentences glued together. Consumers aggregate; they cannot un-glue. **Additive, self-describing fields.** Consumers will read the fields they know and ignore the rest, so adding a field is safe and removing or repurposing one is not. Emitting a version marker on the payload costs nothing and tells a consumer what it is looking at. ## The practical test Ask what a new consumer would have to do to answer a simple question — "show me every change this quarter that was blocked by the memory-limit rule, and which service it was in". If the answer involves a regular expression over message text, the record is missing fields. If it is a filter on two named fields, the payload is doing its job.

  • One rule fires on three containers in the same object — one record or three?
    Three, each with its own field path. Consumers aggregate records; they cannot split a record that has glued three findings into one sentence. Three records let a bot annotate three lines, let a dashboard count three occurrences of the rule, and let a person fix them in one pass because each one is located.
  • Why does the record need to say whether the violation actually blocked the change?
    Because the same rule is often enforcing in one place and merely observing in another. Without an outcome field, a report cannot distinguish "this stopped a deploy" from "this would have", which is precisely the number you need when deciding whether a rule is ready to be turned on for everyone or is still too noisy.
  • Is severity worth carrying if the gate blocks on everything anyway?
    Yes, because consumers other than the gate use it. It drives what gets surfaced inline versus rolled into a periodic report, what gets triaged first when a batch of findings lands, and how a backlog is prioritised. It also survives a change in enforcement: when a rule moves from observe to block, the severity is already there.

saying these in an interview costs you the question

  • Treats the message string as the whole payload
  • Omits the field path because the message mentions the container
  • Emits one record per evaluation instead of per violation
  • Expects consumers to regex-parse prose for identity
  • Repurposes an existing field's meaning instead of adding one
  • Leaves out whether the violation actually blocked the change

context