skip to content

Your rule requiring every image to descend from an approved base is correct, but a third of production images are vendor-built and carry no lineage metadata — do you ship it?

level: principalimportance: nice to knowfreq 31%

answer

  1. Correct is not the same as shippable
  2. Every denial needs an available fix
  3. Endless exceptions are just noise
  4. Scope to whoever can comply
  5. Shelve it and write down why

basics

~20 s

Not as written. A rule is only shippable if every population it denies has a fix someone can perform, and nobody can add lineage metadata to a vendor's image. Narrow it to your own builds, or shelve it.

solid answer

~50 s

I would not ship it against that population. Correctness is not the bar; implementability is. Every change a rule denies has to have a remediation the owning team can carry out, and nobody can stamp lineage metadata onto an image a third party built and will not change. Shipped as written, a third of the estate produces a permanent stream of denials whose only ending is an exception, which teaches everyone that the gate is paperwork. So either narrow the rule to the population that can comply — our own builds, scoped by how the image was produced rather than by the field that is missing — or, if that covers too little of the risk to be worth the weight, shelve it and record the decision. I would not ship it in a weaker mode as a compromise.

code

json · 18 lines
json
{
  "our-build": {
    "config": {
      "Labels": {
        "org.opencontainers.image.base.name": "registry.internal/base/jre:2026-07-14",
        "org.opencontainers.image.base.digest": "sha256:9f2c..."
      }
    }
  },
  "vendor-image": {
    "config": {
      "Labels": {
        "vendor": "ExampleCorp",
        "release": "7.4"
      }
    }
  }
}

go deeper

for a junior

Know that a rule can be correct and still be wrong to ship, because the people it would deny may have no way to satisfy it.

for a middle

Be able to explain what happens operationally when a rule denies a population with no fix: a permanent exception queue that drowns the findings that did have one.

for a senior

Show that you would narrow the rule to the population that can comply and scope it on how the image was produced rather than on the field that is absent.

for a principal

Own the decision itself — declining to ship, naming the gap and its owner, and recording the revisit trigger so the call is not quietly reversed later.

## The test the rule fails The rule is right. Images should descend from a base you have chosen and can patch. The shadow run does not dispute that; it tells you something different, which is that a third of the estate cannot satisfy it. Those images come from vendors, the lineage metadata was never stamped, and no amount of work by the team that deploys them will make it appear. That converts a correctness question into an implementability one, and the test is blunt: **for every population this rule denies, is there an action the owning team can take that turns the denial into a pass?** For your own builds the answer is yes — set the label, rebase onto an approved image. For vendor images it is no, and no is permanent. ## Why shipping it anyway is worse than not shipping A rule that denies changes with no available fix does not produce security outcomes. It produces exception requests, forever, at the rate the vendor images rebuild. Three things follow, and they compound: - The exception process, which exists so that the rare justified deviation is visible, fills with routine traffic and stops being a signal. - Teams learn that the correct response to this gate is escalation rather than remediation, and they carry that lesson to the next rule you ship. - The denial volume masks the true positives among your own builds, which are the findings that actually had a fix available. So the cost of shipping is not the inconvenience; it is the credibility of the whole apparatus, spent on a control that changes nothing about vendor images. ## The three real options **Narrow the scope to the population that can comply.** Make the rule apply to images your pipelines produced, and be explicit that vendor images are out of scope. The mechanical detail matters: scope on how the image was produced — the builder that emitted it — not on the presence of the lineage field, because keying on the missing field means an unlabelled build of your own quietly escapes the rule too. You then have a rule that is smaller, honest about its boundary, and enforceable without a standing exception queue. You also have a visible gap, which you name rather than hide, with an owner for the different control that covers vendor images. **Fix the world first, then revisit.** The underlying condition is that you consume images you do not build. That is addressable — mirroring and rebuilding vendor images onto an approved base where the licence and support model allow it, or making lineage metadata a requirement in vendor selection and contract renewal. Both are real programmes measured in quarters and owned by platform and procurement, not by the policy author. Naming that path is part of the answer; pretending a rule can substitute for it is not. **Shelve it.** If narrowing leaves a rule that covers so little of the risk that its operational weight is not repaid, do not ship a diminished version to show progress. Say the rule does not ship, why, and what would have to change. ## What declining costs you, and what you owe Declining is a decision, and unwritten decisions get relitigated every quarter and eventually ship by default when someone new finds the rule in a branch. Write it down: the measured impact from the shadow run, the named population that cannot comply and why, the risk that stays open, who owns the alternative control for vendor images, and a trigger for revisiting — mirrored rebuilds reaching some coverage, or a procurement clause landing. The underlying discipline generalises past base images. A guardrail is a contract with the people it constrains: you tell them what is not allowed, and you owe them a way to comply. When the shadow run shows a population for whom that second half does not exist, the finding is not that the teams are non-compliant. It is that the rule, as scoped, is not yet a rule anyone can follow — and saying so out loud is the job.

  • How would you scope the rule so it only reaches images that can comply?
    Key it on how the image was produced rather than on the field that is missing: images emitted by our own build pipelines are in scope by construction, everything else is out. Scoping on the absent field instead would let an unlabelled build of our own escape the rule as well, which is the opposite of what you want.
  • What do you owe the organisation if you decline to ship it?
    A written decision. The measured impact, the named population that cannot comply and why, the risk that stays open, who owns the alternative control for vendor images, and a trigger for revisiting — for example mirrored rebuilds reaching real coverage, or a lineage requirement landing in vendor contracts. Unwritten declines get relitigated and eventually ship by default.
  • The CISO wants the rule live this quarter to close an audit finding. What do you say?
    That the rule as written cannot close it: a control a third of the estate can never satisfy produces exceptions, not compliance, and an auditor reading a standing exception queue draws the obvious conclusion. I would offer the narrowed rule, which is genuinely enforceable this quarter, plus a named owner and date for the vendor-image gap.

saying these in an interview costs you the question

  • Ships the rule and lets exceptions absorb the gap
  • Calls the rule fine because it is technically correct
  • Blames vendor teams for missing metadata
  • Leaves the decision unwritten so the rule ships by default
  • Scopes the rule on the missing field itself

context