Signature verification can be staffed at only one choke point across forty clusters — which do you pick?
answer
- who controls the control
- self-service pipelines are team-owned
- coverage, authority, cost
- denials must reach the owning team
- drift across forty clusters
basics
~20 sPick the seat closest to execution: admission. It is the only one that sees everything which actually tries to start, whoever created it. Accept that feedback moves to deploy time and that build context must travel in signed attestations.
solid answer
~50 sI would make admission the mandatory seat and keep any pipeline gates as unenforced fast feedback. The deciding argument is not strength but control: in a self-service platform the pipeline definition lives in the team's own repository, so a gate there is a control owned by the party it constrains — a fork, an edit or a deploy from a laptop removes it, and its absence is invisible. Admission sits on a path nobody routes around, so it also covers images that never appear in any pipeline: operator-installed agents, injected sidecars, a vendor chart's DaemonSet. What I accept losing is real. Feedback moves from merge time, where the author has context, to deploy time, where the message reaches whoever is on call. Policies that depend on build facts now need those facts carried in signed attestations. And forty clusters means forty copies of the policy, so my main operational risk becomes drift.
go deeper
Know the three candidate seats and that a check inside a pipeline only covers work that used that pipeline. You are not expected to run this trade-off yourself yet.
Be able to compare the seats on coverage and on who can disable them, and to explain why an operator-installed workload never passes a build-pipeline gate.
Expect to defend a specific choice for a specific estate, including how denials reach the right engineer and how build context is carried to the enforcement point.
Own the framing that the mandatory control must sit outside the authority of the party it constrains, name the costs you are accepting, and plan for drift across many clusters as the dominant long-run risk.
## The question behind the question Asked with a staffing constraint, this is not really about cryptography — it is about **who controls the control**, and about what you can honestly claim to an auditor or a customer. Three properties decide it: coverage (what population passes through this seat), authority (can the constrained party remove it), and cost (feedback quality, and what the seat puts in your availability path). ## Why the pipeline gate loses as the mandatory seat In a self-service platform, pipeline definitions live with the teams. That makes a gate there structurally advisory: - **The constrained party owns it.** A team can fork the shared template, pin an older version of it, or simply add a step that bypasses it. Not usually maliciously — under deadline pressure, because it was failing. - **Its absence is silent.** A gate that did not run produces no signal. You would have to build a separate control that audits every pipeline for the presence of the gate, at which point you are maintaining two controls to get one. - **It cannot see non-pipeline creation.** Operators install their own agents; platform components inject sidecars; a vendor chart brings images nobody in your organisation ever named. These never traverse a build pipeline of yours, so a gate covers none of them. A gate remains genuinely valuable — it catches ordinary mistakes early, cheaply, and tells the right person — but as the *only* mandatory control it lets you claim only that the paved road is clean. ## Why admission wins on the properties that matter It sits on the path that everything which starts must traverse, regardless of origin, and it is operated by the platform team rather than by the teams it constrains. That is separation of duties expressed as topology: the enforcement point is outside the control of the party it governs. It also produces the artefact you need for an external claim — a denial record for anything that failed, and, with the right policy, a demonstrable statement about the population of workloads rather than about the population of pipeline runs. ## What you are accepting, stated honestly A principal-level answer names the costs rather than pretending they are small: - **Feedback moves late and sideways.** A denial at merge time reaches the author with the change in front of them. A denial at deploy time reaches whoever is deploying, possibly during a release window, in the vocabulary of the policy engine. Mitigate by making denial messages name the failing image and the concrete remediation, and by routing them to the owning team rather than to the platform on-call. - **Build context has to be plumbed.** The seat sees an image reference; policies about where a build came from require that context to travel as signed attestations bound to the digest. That is a real project, not a policy line. - **You take an availability dependency.** Workload creation now depends on the check and on the material it needs. That argues for local, replicated verification material and for verifying expensive things earlier and handing this seat a cheap local decision. - **Forty clusters means forty places to drift.** The realistic long-run failure is not an attacker defeating the check; it is one cluster where it was never installed or was quietly disabled. Treat presence-and-enforcing as a continuously verified inventory property of every cluster, alarmed like any other production invariant, and expect that to be more of your operational effort than the policy itself. ## The estate where this answer changes If part of the estate is a managed container service or a hosted function platform, there is no creation hook you can occupy at all. The only seat left is the deploy API call your automation makes — which is a pipeline-shaped control again, with all the authority weaknesses above. There you compensate on the identity side: very few principals may call deploy, those principals are automation you control rather than humans, and the deploy path itself performs the verification. You should also be explicit that the provider's own image-pull path is invisible to you, so what you can attest is that you only ever asked for verified artifacts — not that only verified artifacts ran. Saying that plainly, rather than overclaiming, is the difference between a defensible control narrative and one that falls apart under audit. ## The one-line version Put the mandatory check where the thing that would be refused cannot route around it and the constrained party cannot switch it off; keep the earlier check as feedback, not as the claim; and spend the saved effort proving the mandatory check is on everywhere rather than making it cleverer.
- Part of the estate is a managed platform with no creation hook. Where does the check go there?Into the deploy API call your automation makes, which is the only seat the customer occupies. That is a pipeline-shaped control, so compensate on identity: only a small number of automated principals may call deploy, humans cannot, and the verification happens in that path. Be explicit that the provider's own pull path is invisible to you — you can attest that you only requested verified artifacts, not that only verified artifacts ran.
- How do you know, six months on, that the check is still enforcing in all forty clusters?By treating presence-and-enforcing as a continuously verified property rather than an install-time fact. Something outside each cluster periodically asserts that the policy exists, is in enforcing mode, and actually refuses a known-bad test artifact, and alarms when it does not. A synthetic probe that expects a denial is worth more than reading configuration, because it tests the outcome rather than the intent.
- A team argues the gate belongs in CI because denials at deploy time are too disruptive. What is your answer?That they are describing a real cost and the wrong remedy. Keep the CI check so they get the early signal, but it cannot be the mandatory one, because the control would then be owned by the party it constrains and would miss everything that never touches their pipeline. The disruption is addressed by better denial messages, routing failures to the owning team, and making the early check accurate enough that a deploy-time denial is a surprise.
saying these in an interview costs you the question
- Picks the CI gate without noticing teams own the pipeline
- Claims admission and a pipeline gate cover the same artifacts
- Ignores that build context is unavailable at the chosen seat
- Treats installing the policy once as the end of the work
- Overclaims coverage on a managed platform with no hook