How do you model a cluster add-on that holds full control-plane privilege?
answer
- Ask what it is inside of
- Not crossing the boundaries — erasing them
- Turn it round: what reaches it
- Low-privilege input, high-privilege actor
- Confused deputy with cluster privilege
basics
~20 sModel it as a component that sits inside every namespace boundary at once, not behind one. Its compromise does not cross boundaries, it erases them, so the threats worth listing are the ones that reach the component: its operators, and the tenant-writable input it acts on.
solid answer
~50 sA component with cluster-wide control-plane privilege is not a box behind a trust boundary; it is inside all of them simultaneously. So the useful modeling question is not `what can it reach` — everything — but `what reaches it`. Three paths matter. Its operators: a small platform team's credentials become the credentials of the whole cluster, which makes a compromised or coerced operator a first-class entry point. Its inputs: if the component reads fields on objects any tenant can write, then a low-privilege tenant is feeding a cluster-privileged consumer, which is a confused deputy and a straight elevation of privilege with no exploit required. And its position in the path: a component that reviews or mutates every workload change is a single point of failure, so denial of service against the platform belongs in the model too. I would also check attribution — if its actions are recorded under its own identity rather than the requester's, the audit trail loses non-repudiation.
go deeper
Be ready to say why a component with cluster-wide privilege matters more than an application container, and that its compromise affects every team rather than one.
Explain the mechanics of the borrowed-privilege path: an ordinary tenant writes a field, a highly privileged component reads it and acts, and privilege moves without any exploit.
Show that you invert the analysis — reach is total, so you model entry points — and produce concrete asks about privilege narrowing, input scoping, operator credentials and attribution.
Own the organisational side: how much standing cluster-wide privilege the platform should hold at all, who reviews the components nobody reads, and what the business accepts when the platform itself is the single point of failure.
## Why the usual boundary drawing fails here Most cluster models are drawn tenant-first: a namespace per team, lines between them, threats crossing those lines. A platform add-on — a controller or gating component every team installs, that nobody outside the platform team reviews, running with privilege over the whole cluster — breaks that picture, because it does not live on one side of any of those lines. It is **inside every one of them at the same time**. If you draw it as a box in its own namespace with a boundary around it, you have drawn a lie, because nothing on that line constrains what it does to the other namespaces. The honest notation is to treat the **control plane and everything holding cluster-wide privilege as one privilege domain** and to shade every tenant boundary as depending on it. Then the model states plainly: these tenant boundaries exist only while this domain is intact. That is a far more useful sentence to hand a platform team than a list of generic threats. ## Turn the question around: what reaches it? Since the reach of a compromise is total by construction, the analysis moves to entry points. **1. Operators.** A small platform team upgrades and administers the component. Their credentials are, transitively, the cluster's credentials. Model a compromised or coerced operator as a legitimate attacker position — not a suspicious one, just the highest-value credential in the estate. Threats: elevation of privilege (violating authorization) and tampering with cluster state (violating integrity). Useful mitigations are all organisational shape rather than product features: how many people hold it, whether use is intentional and reviewed, whether the credential is standing or acquired for a task. **2. Tenant-writable input.** This is the finding senior candidates get and others miss. A cluster-privileged controller usually acts on fields of objects that ordinary tenants create — a label, an annotation, a spec field, a reference to another object. If any of those fields decides *what the controller does* or *what it does it to*, a tenant with the lowest privilege in the cluster is now driving a component with the highest. That is the classic **confused deputy**: the privilege is not stolen, it is borrowed through a legitimate interface. It is elevation of privilege and needs no memory-safety bug, no escape, no stolen token. In the model, draw the data flow from tenant object to controller and put a boundary on it, because trust changes there even though both ends live inside the cluster. The corresponding design question is scope: does the action the controller takes stay confined to the requester's own namespace, or can a field name a target elsewhere? A controller that will act on any object named in an input is a cross-tenant primitive. **3. Its upgrade path.** New versions arrive from somewhere and run with the same privilege. Name it as an entry point and stop there — the build-and-deploy path is its own system to model, with its own actors and its own credentials, and it deserves that treatment rather than a footnote here. ## Categories that actually dominate Working STRIDE against this component rather than against tenant boundaries: - **Elevation of privilege** (authorization) — the headline, via operators and via borrowed privilege from tenant input. - **Tampering** (integrity) — cluster objects are the system's control state; altering them changes what runs where, which is more powerful than altering any single application's data. - **Denial of service** (availability) — a component in the path of every workload change or every scheduling decision is a common-mode failure. If it wedges, deploys stop and self-healing stops. Platform availability is a real asset and models often forget it because the attacker did not get anything. - **Repudiation** (non-repudiation) — if actions land in the audit trail under the component's own identity, you can see that something happened but not who caused it. Carrying the requesting principal through into the record is what preserves attribution. - **Information disclosure** (confidentiality) — a component that can read anywhere can read every secret; worth stating explicitly so the risk is owned rather than assumed away. ## What to do with the finding The output is a small number of concrete asks, in rough order of value: narrow the component's privilege to the object kinds and verbs it truly needs; treat tenant-writable fields as untrusted input and constrain any action to the requester's own scope; make the operator credential non-standing and its use reviewable; keep attribution in the audit record; and plan for the component being unavailable, because that failure mode arrives without any attacker. The reason to model this at all is that platform teams tend to review the workloads and skip the platform. The add-on nobody reads is the component with the most privilege in the building.
- A tenant can set a field on their own object that the controller turns into a cluster-wide action. How do you rate that?As elevation of privilege, high, because it needs no vulnerability at all — the privilege is borrowed through a supported interface. The rating should reflect that the attacker position is any authenticated tenant, the cheapest position in the cluster. The fix is to treat the field as untrusted input and to bind the resulting action to the requester's own scope.
- Where does availability enter this model?A component sitting in the path of every workload change or scheduling decision is a common-mode failure. If it stalls or rejects everything, deploys and recovery stop cluster-wide, which is denial of service against the platform rather than against any one application. It also arrives without an attacker, so it must be modelled and tested as an ordinary failure mode.
- The audit record names the controller, not the tenant whose request triggered it. Which STRIDE category is that?Repudiation, the threat against non-repudiation. You can prove an action occurred but not attribute it to its originator, so investigation and accountability both fail exactly where privilege is highest. The remedy is to carry the requesting principal into the record the controller writes, so the chain from request to effect stays intact.
A caretaker with a master key is not behind any of the building's doors; they stand inside all of them at once, so the only questions worth asking are who becomes the caretaker and who gets to tell them what to open.
saying these in an interview costs you the question
- Draws the add-on inside its own namespace boundary
- Says its privilege is fine because the team is trusted
- Ignores tenant-writable fields as an input path
- Lists no availability threat against the platform itself
- Assumes the audit trail attributes actions to the requester