When the platform's default confinement filter blocks a team's workload, what standard would you set for relaxing it?
answer
- name the call, not the wall
- a ladder, cheapest rung first
- scope to one workload, one environment
- an owner, an expiry, a review
- never move the platform-wide default
basics
~20 sRequire the request to name the refused call, not the wall. Then permit that one call in confinement scoped to that workload, with an owner, an expiry and a review — and treat switching the filter off as an exception, not the next rung down.
solid answer
~50 sThe request that arrives is always 'we need confinement off'; the first job of the standard is to convert it into 'which call was refused'. With that in hand there is a ladder, cheapest first: change the workload so it stops needing the call, then permit that single call for that single workload, then name the one device it must reach. Only below all of those sits switching the wall off — and that is not a rung, it is an exception, so it carries a named owner who is not the requester, a stated scope, an expiry, and a review. Two things the standard must forbid outright: relaxing the platform-wide default to end a queue of tickets, which removes the wall from every workload at once, and open-ended exceptions, which outlive the bug that motivated them and are inherited by every team that copies the spec.
go deeper
Take away the habit rather than the policy: when confinement blocks you, find out what was refused and ask for that one thing, instead of asking for the wall to be removed.
Be able to state the ladder — remove the need, permit the one call, name the one device — and say why each rung is cheaper than the one below it.
Argue the scope and the lifetime: one workload, one environment, an expiry that lapses by default, and evidence of the refusal before anything is granted.
Own the asymmetry. The wall is cheap for nearly every workload and pays once; under delivery pressure nothing resists removing it, so the standard exists to make the narrow fix easy and the broad one visibly expensive.
## The request that arrives, and the one you need What lands on the platform team is 'the filter breaks our worker, please turn it off'. That is not a request anyone can evaluate, because it names the wall rather than the obstacle. The first clause of any workable standard is therefore procedural: **a relaxation request must name what was refused** — which kernel call, or which device — and show how that was established. It is the clause that does the most work, because roughly half of these requests dissolve the moment someone has to produce the call: the workload turns out to have a configuration switch that keeps it off the fast path, or a debugging tool left in the image, or a library that has a portable fallback nobody had enabled. ## A ladder, cheapest first 1. **Remove the need.** Configure the workload not to take the path that makes the call, drop the tool that was making it, or use the portable implementation. Nothing about confinement changes. 2. **Permit the single call**, in confinement scoped to that one workload, with a comment naming why. The default stays intact for everyone else. 3. **Name the single device or path** in that workload's reachability profile, with the access it actually needs and no more. Often required together with step 2. 4. **Give the workload a stronger boundary** than a shared kernel if it genuinely needs broad kernel access. A workload that needs a long list of administrative entry points is telling you it does not belong beside other tenants behind one kernel, and that is a placement decision rather than a profile decision. 5. **Switch the wall off.** Not a rung — an exception, and the standard should describe it as one. ## What an exception must carry - **What was refused**, and the evidence: a reproduction, not an assertion. - **Scope**: one workload, one environment. Never the platform default, and never a whole team's namespace of workloads. - **An owner**, named, who is not the person who asked. An exception with only a requester has nobody to answer for it later. - **An expiry and a review date**, with the default outcome on expiry being that it lapses rather than that it renews. - **What compensates** in the meantime, and an honest statement of what does not. ## The costs nobody prices | response to the ticket | blast radius | how long it lasts | |---|---|---| | workload changed so the call is not made | none | permanent, and free | | one call permitted for one workload | that workload's surface grows by one entry point | until the workload changes | | confinement switched off for that workload | the whole denied tail returns for that workload | until someone audits it, which is usually never | | platform default relaxed | every workload on the platform, including future ones | effectively forever | Three failure modes are worth naming explicitly, because each one is a rational-looking local decision: - **Profile drift.** Per-workload confinement is code that nobody owns. It breaks when a library changes which calls it makes on an upgrade, and the team that hits the breakage is rarely the team that wrote the profile. - **Exceptions spread by copying.** A relaxed spec is the most copied artefact in any platform, because it is the one known to work. The workload that inherits it never needed the relaxation and nobody re-examines it. - **Ending the queue by moving the default.** When enough tickets arrive, relaxing the platform default is the cheapest-looking fix and the most expensive real one: it removes the wall from every workload at once, silently, including the ones that were never blocked. ## The judgment a lead is actually being asked for The honest framing is that this wall costs almost nothing for almost every workload and pays off exactly once, in a scenario you hope never happens. That asymmetry is what makes the standard necessary — under delivery pressure the local decision is always to remove the wall, and nothing in the system pushes back on its own. What a lead owns is the default answer, the shape of the exception, and the fact that someone looks at the exceptions again. The standard is not really about kernel calls; it is about making the narrow fix the path of least resistance and the broad one visibly expensive.
- The team says they cannot identify the call and the release is today. What do you do?Grant the narrowest thing that unblocks them with the shortest life you can justify — confinement relaxed for one workload in one environment, expiring in days, with the owner and the follow-up recorded before it is applied. The mistake is not granting it; the mistake is granting it with no end date, because nothing will revisit it afterwards.
- Ten teams file the same request in a quarter. Is relaxing the default now the right call?No — it is the most expensive response available, because it removes the wall from every workload including those that never asked. Ten identical requests are evidence about the workloads, so publish a reviewed profile for that workload shape that teams opt into, and leave the default where it is.
- How do you stop a granted exception spreading by copy-paste?Make it visible where it is used and make it expire. A relaxation that carries its owner and its expiry in the spec announces itself in review, and one that lapses forces the inheriting workload to justify itself rather than inherit silently. Publishing a reviewed alternative for the common case removes most of the reason to copy.
saying these in an interview costs you the question
- Treats switching the filter off as simply the next step after permitting a call.
- Relaxes the platform-wide default to clear a queue of tickets.
- Grants an exception with no expiry and no named owner.
- Accepts 'we need it off' without establishing what was refused.
- Assumes a relaxed spec stays with the workload that requested it.