In a shared Kubernetes cluster, which of namespace, pod and node is a real trust boundary?
answer
- Ask what enforces each line
- Some lines are administrative only
- Above the API versus below it
- Two pods, one kernel
- Node credential reaches its pods' secrets
basics
~20 sThe node is the hard line, because pods on one node share its kernel, its disk and its credential. A namespace is an administrative scope the control plane enforces above the API, so it stops nothing that travels below it.
solid answer
~60 sA trust boundary needs two things: a change in trust level, and a mechanism that enforces it. In a cluster those differ per candidate line. A namespace is enforced by the control plane when it authorizes API calls and scopes names, so it holds against someone acting through the API and against nothing underneath it. A node is enforced by the host kernel, so it holds against workloads on other nodes but not against a workload sitting beside you. A pod is a boundary only as far as its isolation settings genuinely hold. Take a platform cluster where a `payments` namespace and a `marketing-site` namespace land on the same node pool: compromise the marketing workload and the attacker is a process on a node that also runs payment pods, with that node's filesystem, its network position and its credential — which can read the secrets mounted into pods scheduled there. So I draw the namespace line, but I draw the node line too, and I rate what crosses the node line.
go deeper
Be ready to say what a trust boundary is in one sentence and to point at the node as the line that the kernel enforces. Knowing that a namespace is an administrative scope rather than isolation already puts you ahead.
Explain the mechanics: which mechanism stands on each candidate line, why namespace separation is decided at the API and node separation underneath it, and what a node's own credential can reach for the pods scheduled on it.
Show the judgment on a real design: place the boundaries on a co-tenanted cluster, name the assumption an escape would break, and rate the resulting threats against the business asset rather than against the diagram elements.
Own the tradeoff the boundary work exposes — dedicated node pools, compensating detection, or moving a tenant out — and be able to argue the cost of each to people who chose shared infrastructure to save money.
## What a trust boundary actually is On a data-flow diagram a trust boundary is a line crossed by data or control where the level of trust changes **and something enforces the change**. Both halves matter. A line with a trust difference and no enforcement is a wish drawn in ink, and a line with enforcement but no trust difference is clutter. Modeling a cluster goes wrong when people reach for the org chart's lines (team, namespace, project) instead of asking, for each line, *what mechanism stands on it, and which threats does that mechanism see?* ## The candidate lines inside a cluster | Candidate line | Enforced by | Holds against | |---|---|---| | Container / pod | Kernel isolation on the node: separate process and filesystem views, dropped capabilities, syscall filtering | A workload staying inside its own view of the host — only as far as those settings are actually applied | | Namespace | The control plane, when it authorizes API requests and scopes object names | An identity acting through the API; nothing that travels below the API | | Node | The host kernel, plus the node's own identity and disk | Workloads running on other nodes; not workloads on the same node | | Control plane | Authentication and authorization at the API server, and access to its datastore | Everyone who is not an administrator of the cluster | | Cluster edge | The ingress path and the surrounding cloud network or account | The internet and the rest of the estate | The crucial observation is that the namespace line and the node line face **opposite directions relative to the API**. Namespace separation is decided by the control plane on the way in through the API. Node separation is decided by the kernel, underneath every API call. An attacker who is already executing code inside a pod does not need the API at all, so a model that only draws namespace lines has drawn none of the boundaries their path crosses. ## Worked example: the payments and marketing namespaces An internal platform cluster runs a `payments` namespace and a `marketing-site` namespace, scheduled onto one shared node pool because it was cheaper. The threat to model is not an anonymous internet attacker stealing a database; it is the low-value workload as the foothold and payment-flow integrity as the asset. Assume the marketing workload is compromised through a flaw in its own application code. The attacker is now a process on a node that also runs payment pods. Without ever authenticating to the API as payments, they are adjacent to: - the **host kernel**, and through it every other pod on that node if isolation fails; - the **node filesystem**, including material mounted for the other pods scheduled there; - the **node's credential**, which is normally limited to the objects the node's own pods need — but those include the secrets mounted into every pod scheduled on it, payments among them; - the node's **network position** inside the cluster. So the namespace line, drawn as though it separated the two tenants, does not sit anywhere across the attacker's path. That is the finding, and it is worth more than a page of generic threats. ## Escape as the named assumption Attach one sentence to the node boundary: *we assume the container runtime and host kernel keep workloads on a node from reading or influencing each other.* Writing it down does two useful things. It becomes reviewable — a privileged pod, a host path mount, a shared device, or a runtime change all invalidate it, and the review is now a one-line diff rather than a re-read of the whole model. And it makes the consequence explicit: a container escape is not one more threat in the list, it **deletes a boundary**, so every threat you suppressed because of that boundary comes back at once. ## What the model says next Work STRIDE per crossing rather than per box. At the node line, elevation of privilege (violating authorization) and information disclosure (violating confidentiality) dominate; the marketing-to-payments direction is lateral movement, so the ratings that matter are tampering with payment flow and disclosure of payment data. Rate against the **asset**, not against the element: the risk is money and transaction integrity, not an abstract cluster. Then offer the choices the boundary work exposes: keep the arrangement and add compensating detection on the node line; move payments to its own node pool so the line falls where the kernel is; or move the tenant out entirely. Each is a different cost, and the model's job is to make the cost visible rather than to pick for the organisation. ## Common mistakes Drawing a box around every pod and stopping there hides both the node and the control plane, which is where the interesting privilege lives. Treating the cluster's internal network as trusted because it is internal reintroduces every spoofing and disclosure threat you thought you had removed. And treating the control plane as somebody else's infrastructure leaves the one component that is inside all of your boundaries out of the diagram.
- If the two namespaces move to separate node pools, which threats does that actually remove?It removes the shared kernel and the shared node credential, so same-node escape and secret pickup from a neighbour's mounts drop out. What survives is everything above the node: the shared control plane, shared cluster networking and storage, shared platform components installed for every team, and the shared operators who administer both pools. The boundary has moved up one level, not disappeared.
- How do you record the assumption that the runtime isolates co-scheduled workloads?Write it on the diagram as a labelled assumption attached to the node boundary, together with the threats it suppresses. That gives you a review trigger: when a workload gains a host mount, a privileged setting or a different runtime, you re-open exactly the threats that sentence was holding shut, instead of re-deriving the whole model.
- A retail chain runs one small node per store in a back room staff and contractors can reach. What changes in the model?The node boundary now has a physical face, so the attacker position shifts to someone with local network access or hands on the box. Disk-at-rest and any cached cluster credential become reachable, and the asset shifts from customer data to availability of checkout plus the credentials that node holds. Threats that assumed a locked data centre need re-rating, not re-listing.
Namespaces are separate mailboxes in one building's lobby; the node is the building's outer wall. Sorting the mail differently does not help against someone already inside the building.
saying these in an interview costs you the question
- Namespaces isolate tenants, so the tenants are separated
- A container is an isolation boundary as strong as a VM
- Draws a box per pod and calls the diagram finished
- Treats the cluster network as trusted because it is internal
- Leaves the control plane out as somebody else's infrastructure
- Rates threats against elements rather than against assets