skip to content

A managed container platform re-resolves your image reference at every cold start and offers no hook - how do you respond?

level: principalimportance: should knowfreq 32%

answer

  1. a control you cannot enforce must be converted
  2. check whether digests are accepted at all
  3. shrink who can move the name
  4. preventive becomes detective, with an owner
  5. scope the claim, register the exception

basics

~20 s

Check first whether the platform accepts a digest reference at all. If not, stop calling this a preventive control: shrink who can move the tag, detect divergence at each cold start with a named owner, and scope the compliance claim honestly.

solid answer

~50 s

The engineering question is small and the organisational one is large. Technically: confirm whether a digest reference is accepted anywhere in the platform's configuration, because if it is, the problem disappears. If it is not, you cannot enforce identity at pull time, so you convert the control rather than pretend it holds - restrict push rights on the source repository to the release pipeline alone, forbid remapping existing tags where the registry can, and add a detective check that compares the digest the platform reports for each cold-started instance against the digest approved for that release, with an alert, an owner and a rollback path. Then decide by asset value whether the workload belongs there at all. Finally, fix the claim: an estate described as fully pinned is now inaccurate, and the exception needs a register entry, an owner and a review date rather than silence.

go deeper

for a junior

Take away the core fact: if something other than you resolves the image name at start time, the identity of what runs is decided later and outside your control.

for a middle

Be able to explain why a build-time or deploy-time check does not compensate for a resolution that happens at cold start, and what evidence you would need to detect a mismatch.

for a senior

Show the practical ladder: confirm whether digests are accepted, restrict who can move the name, then build a mismatch alert with an owner, a detection target and a rehearsed response.

for a principal

Own the organisational half - converting an unenforceable preventive control into an owned detective one, scoping the compliance claim to what you can defend, and deciding workload placement by what the asset is worth rather than by policy purity.

## Name the situation precisely The platform resolves the image reference itself, at cold start, from a name you supply, and gives you no place to intervene between resolution and execution. That means the artifact identity that executes is decided by the platform at an arbitrary future moment - possibly weeks after your release, possibly several times a day as instances scale from zero. Your compliance record says a digest; what ran is whatever the name resolved to at that instant. The asset at risk here is not only integrity but **non-repudiation**: you cannot truthfully state what executed. A control you cannot enforce is not a control. The principal-level move is to convert it deliberately and say so, rather than leaving a policy on the page that the platform quietly ignores. ## The ladder, in order **1. Verify the constraint.** Vendors document less than they support. Check whether a digest form is accepted in any configuration surface - the service definition, an API field, a deployment payload. Half the time the constraint is a habit in your own templates rather than a platform limit, and the entire problem evaporates. Establish this before designing around it. **2. Shrink the mutation surface.** If a name must be resolved, make that name hard to move. Push rights on the release repository go to the release pipeline identity only, human push is removed, existing tags are made non-remappable where the registry supports it, and every publish is attributable. You have not restored the pin; you have reduced the set of principals who could undermine it from dozens to one automated identity, which is a real and defensible reduction. **3. Convert preventive to detective, with teeth.** Capture the digest the platform reports for each running instance from its runtime metadata or logs, compare it against the digest approved for that release, and alert on mismatch. This only counts as a control if it has three things a policy document usually lacks: a named owner, a stated detection target you actually measure, and a pre-agreed response - because the response is what happens at 3am, and "someone will look at it" is not one. **4. Escalate externally.** File the gap with the vendor as a security requirement rather than a feature wish, and ask peer customers to do the same. Roadmap pressure is slow but it is the only path to closing the gap properly, and a dated request also strengthens your own exception record. **5. Decide placement by asset, not by principle.** Migrating a workload off a platform is expensive and can introduce more risk than it removes. Judge by what unreviewed code running there could reach: money movement, credentials, an authoritative data store or an audit-truth obligation justifies the move; a stateless, low-privilege endpoint with no secrets usually does not. Write down which test you applied, because the next person will ask why this workload stayed and that one left. ## The claim is the part people get wrong Most of the failures here are not technical. A team that has done excellent pinning work everywhere else will let "all production artifacts are digest-pinned and verified" stand in a customer questionnaire, because one platform is an embarrassing footnote. That is the worst outcome available. The first auditor or customer engineer who finds the exception does not conclude "one gap" - they conclude the whole answer was written optimistically, and every other control you claimed is re-opened. So state it as it is: which components resolve a digest, which one does not, what compensates on that path, who owns it, and when it is expected to close. A scoped, dated, owned exception is a sign of a mature programme. An unqualified claim with a known hole behind it is a finding. ## What not to do - Do not accept the risk with no detection and no owner; that is a decision only in name. - Do not drop the control estate-wide because one platform cannot honour it - the answer to a partial gap is never uniform weakening. - Do not substitute a build-time scan and call it equivalent; scanning the artifact you built says nothing about the artifact the platform later resolved. - Do not migrate reflexively before you have measured what the workload can actually reach.

  • What exactly does the detective control compare?
    The digest the platform reports for the instance that actually executed, per cold start, against the digest approved for that release. It needs the platform to expose the resolved reference in runtime metadata or logs, plus an alert with a named owner and a pre-agreed rollback path. Without those three it is telemetry, not a control.
  • How do you decide whether to migrate the workload off the platform?
    By asset, not by principle. Weigh migration cost against what unreviewed code running there could reach: money movement, credentials or an audit-truth obligation justify moving; a stateless low-privilege endpoint usually does not. Record the test you applied so the next placement decision is consistent rather than mood-driven.
  • The customer contract says artifacts are pinned and verified. What do you tell them?
    Exactly which components resolve a digest, which one does not, what compensates on that path, and when it closes. Overclaiming is worse than the gap itself: the first reviewer who finds an unqualified statement with a known hole behind it re-opens every other control you asserted.

saying these in an interview costs you the question

  • Accepts the risk with no detection and no owner
  • Claims the whole estate is pinned when one platform is not
  • Drops the control everywhere because one platform cannot honour it
  • Offers a build-time scan as an equivalent compensating control
  • Migrates immediately without weighing what the workload can reach

context