Resources keep appearing that your pipeline policy would have rejected — how do you make the pipeline the only path?
answer
- stop adding, start closing
- ask the surface, not the pipeline
- group change events by calling identity
- close, constrain, watch
- coverage, not pass rate
basics
~20 sStop adding checks and start closing routes. Inventory every identity that can write to the surface, reduce standing write access so the deploy identity is effectively the only one, and put detection on whatever routes you choose to leave open.
solid answer
~50 sBuild the inventory from the target's side, not the pipeline's: take the change events on that surface over a window, group them by calling identity, and everything that is not the deploy identity is a route you did not know you had — a console session, a laptop apply, a second pipeline, a vendor integration, a controller. Then reduce standing write permission so the deploy identity is effectively the only principal that can create that resource type, and push whatever rules you can into the surface's own authorization layer so the denial happens where the mutation happens rather than one hop upstream. Some routes will stay open for good reasons; put those behind detection and attribution instead of pretending the check covers them. Measure coverage as the share of changes to the surface that transited the gate, not the pass rate of the check.
go deeper
Understand that when changes appear outside the pipeline, the fix is usually about who holds permissions rather than about the wording of the rule.
Be able to describe building a route inventory from the surface's own change events grouped by calling identity, instead of reasoning from what the pipeline recorded.
Show the full sequence: inventory the principals, close standing human write access, constrain the automation identities that must keep it, and put detection on whatever stays open.
Own the tradeoff between a route you close and a team's ability to ship, and be explicit about which residual routes you accept, who owns each, and what metric replaces check pass rate.
### Reframe the problem before solving it The reflex when violations keep appearing is to add another check or tighten the rule. Both are work on the route that is already covered. The finding — resources exist that the policy would have rejected — is evidence about *routes*, so the work belongs there. ### Step one: inventory from the target, not the pipeline The pipeline cannot tell you what bypassed it; by construction it has no record of changes it never saw. The surface can. Pull the change or audit events for that resource type over a meaningful window — a month is usually enough to catch weekly and monthly rhythms — and group them by calling identity. What comes back is normally uncomfortable and always useful: - The deploy identity, which is the route you thought was the only one. - Individual human identities, from console sessions and laptop applies. - A second automation identity belonging to a pipeline in another repository, often owned by a team that predates the current platform. - A vendor or integration identity with broad permissions granted during an onboarding nobody revisited. - A controller or autoscaler creating resources as a legitimate downstream effect of something else. That list is the actual scope of work. Until it exists, every conversation about enforcement is guesswork. ### Step two: separate the routes into three buckets **Close.** Human standing write access to production resource types is usually the biggest bucket and the most closable. Removing it converts the deploy identity into the only route, which is what makes everything inside the pipeline meaningful. Expect pushback, and expect the honest version of it: people use direct access because something in the pipeline is slow or unreliable. Fix that in the same programme or the access will come back. **Constrain.** Some identities must keep write access — an integration, a controller, a platform tool. For these, narrow the permission itself to the shapes the rule allows, so the constraint is expressed where the call is authorized rather than checked one hop upstream. This is the strongest form available: the rule and the mutation happen at the same place, so there is no gap between them. **Watch.** Whatever remains open is a deliberate decision, and it should be written down as one. Attach a detective control: an evaluation of live resources for the same condition, and an alert on creations by identities outside the expected set. The value is not just catching violations — it is that the residual route stays visible instead of quietly becoming the norm. ### Step three: change what you measure A check's pass rate is a statement about the changes it received, so it goes *up* as more work routes around the gate. The metric that reflects reality is coverage: of all changes to this surface in the period, what share transited the gate? Both numbers come from the inventory you already built. A programme reporting pass rate and a programme reporting coverage will make different decisions within a quarter of each other. ### The failure mode this is guarding against The worst outcome in this area is not a gate that blocks the wrong change. It is a placement that everyone believed was enforcing and structurally never could: an advisory check reported upward as a control, a deploy step whose credentials three teams share, a rule scoped to a resource type nobody uses any more. Those show up as green dashboards and a steady trickle of resources that should not exist. The tell is always the same — the report describes what the pipeline did, and nobody has asked the surface what happened to it. ### What good sounds like in the interview Name the inventory step first, and say explicitly that it comes from the target's change events rather than the pipeline. Sort routes into close, constrain and watch. Say that some routes will remain open on purpose and that you will name them rather than absorb them. Finish on the metric change, because that is what stops the problem recurring after you rotate off the team. An answer that goes straight to *write a better rule* signals someone who has authored guardrails but never had to prove one was working.
- You find a vendor integration creating these resources directly. What do you do with it?Treat it as a route and pick a bucket deliberately. Either bring it onto the gated path, or narrow its permissions to exactly the shapes the rule allows so the constraint sits where the call is authorized, or accept it and put explicit detection and an owner on it. What you must not do is leave it out of the coverage story.
- How do you know your inventory of routes is complete?You do not, and you should say so. What you can do is derive it from the surface's own change events over a window rather than from documentation or memory, and re-derive it periodically. Anything that appears as a calling identity you cannot account for is a route that existed whether or not anyone knew about it.
- The team says direct access is the only reason they can ship on time. How do you respond?Take the claim seriously and measure it: how often is direct access used, and what was slow or broken in the pipeline each time. Closing the route without fixing the reason produces a worse system, because the pressure reappears as pipeline definitions being edited around the check. Fix the cause and the closure holds.
saying these in an interview costs you the question
- Adds a fourth check to the route already covered
- Reports check pass rate as coverage of the surface
- Assumes the list of who can write is already known
- Treats detection as failure rather than a deliberate posture
- Closes direct access without addressing why people used it