skip to content

You are asked to introduce pod-level network segmentation across an existing multi-tenant Kubernetes cluster that today runs entirely unrestricted. How would you plan and sequence that programme so you end up with default-deny everywhere without causing outages?

level: principalimportance: should knowfreq 36%

answer

  1. Prove enforcement before trusting policy
  2. Observe real flows, do not guess from diagrams
  3. Platform allowances -> allow rules -> deny ingress -> deny egress
  4. Namespaces born closed via provisioning
  5. CI must assert the deny path, not just the allow path

basics

~20 s

Verify the CNI actually enforces policy, then observe real traffic before restricting it. Roll out per namespace: ship DNS and platform allowances first, derive workload allowances from observed flows, then apply default-deny ingress, then egress. Make the default-deny part of namespace provisioning so new namespaces are born closed, and test deny paths in CI.

solid answer

~60 s

I would treat it as a migration with a measurable end state, not a YAML exercise. **Prove enforcement first.** Confirm the CNI implements NetworkPolicy and demonstrate it in a scratch namespace - policies that are stored but ignored give false assurance, which is worse than no policy. **Observe before restricting.** Use flow logs or the CNI's policy audit mode to build the real communication graph. Hand-written guesses miss the batch job that runs monthly. **Sequence by blast radius.** Per namespace, in order: platform allowances (DNS, kubelet probes, API server, metrics scraping), then workload-specific allow rules derived from observed flows, then default-deny **ingress**, then soak, then default-deny **egress**. Egress last because it breaks the most and needs the DNS and external-destination story settled. **Make closed the default.** Bake default-deny plus the platform allowances into namespace provisioning, so new tenants start segmented and the programme does not regress. **Guardrail the semantics.** Because policies only add allowances and the peer OR/AND distinction is silently wrong-able, add lint or admission checks and CI tests that assert denied paths stay denied. Success metric: percentage of namespaces with enforced default-deny, tracked as a platform SLO with named owners.

code

yaml · 22 lines
yaml
# 1. platform allowances (applied first)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: {name: platform-allow-dns, namespace: tenant-a}
spec:
  podSelector: {}
  policyTypes: ["Egress"]
  egress:
    - to:
        - namespaceSelector:
            matchLabels: {kubernetes.io/metadata.name: kube-system}
          podSelector:
            matchLabels: {k8s-app: kube-dns}
      ports: [{protocol: UDP, port: 53}, {protocol: TCP, port: 53}]
---
# 2. default-deny (applied only after allowances exist)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: {name: default-deny-ingress, namespace: tenant-a}
spec:
  podSelector: {}
  policyTypes: ["Ingress"]

go deeper

for a junior

Understand that segmentation is rolled out gradually, namespace by namespace, and that DNS must be allowed before egress is denied.

for a middle

Lay out the ordering concretely and name the traffic that always breaks: DNS, kubelet probes, metrics scraping, API-server access.

for a senior

Drive it operationally - observe-mode data to derive rules, per-namespace soak with a fast rollback, CI tests that assert denied paths.

for a principal

Own the end state and its governance: born-closed namespaces, an egress gateway so policy survives changing third-party IPs, authoring-time guardrails against the additive model, a coverage SLO, and a clear line between L3/L4 policy and mesh-level identity.

## What makes this hard The end state is easy to describe - every namespace deny-by-default with explicit allowances - and hard to reach, for reasons specific to this API: - **Restriction is invisible until it bites.** Nothing tells you which flows a policy will break; you find out when a monthly job fails. - **Policies only add.** No team can narrow another's over-broad allowance, so a single sloppy policy silently defeats the segmentation around it. - **Enforcement may not exist.** The API accepts objects the plugin ignores, so "we have policies" is not evidence of anything. - **Namespaces are the unit.** There is no cluster-wide default in the portable API, so new namespaces are born open unless provisioning closes them. ## Phase 0 - prove the substrate Confirm the CNI enforces NetworkPolicy, and prove it empirically: scratch namespace, two pods, default-deny, show the connection fails, add an allow, show it succeeds. Record that as a recurring conformance test, because a CNI upgrade or a cluster rebuild can change it. Decide at this point whether you will use only the portable API or also vendor CRDs or AdminNetworkPolicy for cluster-wide guardrails - that choice affects portability and should be made deliberately, not by drifting into it. ## Phase 1 - build the traffic map Deriving allowances from architecture diagrams reliably misses reality. Use the CNI's flow logs, its observability tooling, policy audit mode, or a service mesh's telemetry to record actual pod-to-pod and pod-to-external flows over a period long enough to capture weekly and monthly jobs. The output is a per-namespace list of required peers and ports, which becomes the draft allow set. This phase is also where you discover the surprises: the shared cache everyone talks to, the debugging tool with cluster-wide reach, the job that calls a third-party API from a pod nobody owns. ## Phase 2 - platform allowances first Before any deny, ship the policies that everything needs, ideally as a templated bundle applied to every namespace: - DNS egress (UDP **and** TCP 53) to the DNS pods, selected by label. - Ingress for kubelet probes and for the metrics scraper. - Egress to the API server for workloads that use it. - Ingress from the ingress controller's namespace for public-facing workloads. Shipping these first means the later deny can never leave a gap. ## Phase 3 - ingress default-deny, namespace by namespace Start with a namespace whose owners are engaged and whose failure is survivable. Apply the derived allow rules, then the default-deny ingress, then soak with monitoring on application error rates and readiness. Repeat, widening. Keep a documented rollback (delete the deny policy) that any on-call engineer can execute in seconds. ## Phase 4 - egress default-deny Egress last, because it depends on decisions you should have already made: how third-party destinations are reached. The durable answer is an **egress gateway or proxy** - workloads get egress only to the proxy by label, and the proxy enforces allowed hostnames. That keeps NetworkPolicy stable while destination IPs churn, gives you one place for logging, and avoids scattering CIDR rules or committing to a vendor's DNS-aware CRD. ## Phase 5 - make it the default and keep it - Namespace provisioning creates the default-deny plus platform allowances automatically; a namespace cannot exist unsegmented. - **Governance of who may write policies.** Since anyone with create rights can widen access, either restrict NetworkPolicy creation in shared namespaces via RBAC or require the policy set to ship through the same reviewed pipeline as the workload. - **CI tests assert deny paths.** An over-broad peer list still passes every allow test; only a test that asserts a forbidden connection fails will catch the classic separate-dash mistake. - **Track coverage as an SLO**: percentage of namespaces with enforced default-deny, with named owners for the stragglers, reported like any other platform metric. ## Framing for the interview The judgement being probed is whether you treat this as a socio-technical migration. Strong answers name observation before enforcement, ordering by blast radius, DNS and probe traffic as the known breakages, born-closed provisioning, and governance of the additive model. They also state the limits: NetworkPolicy is L3/L4 blast-radius reduction, not identity - if the requirement is authenticated service-to-service communication, mTLS through a mesh is the complementary control, and choosing both means deciding where policy lives so the two do not contradict each other.

  • How do you stop one team's over-broad policy from silently undoing the segmentation around it?
    Two levers. Governance: restrict who may create NetworkPolicy objects in shared namespaces via RBAC, and require policies to ship through the same reviewed pipeline as the workload manifests. Automation: lint policies in CI for over-broad peers - empty selectors, separate-dash OR where an AND was meant, ipBlock 0.0.0.0/0 - and run tests asserting that specific forbidden connections still fail. The additive model gives no runtime protection, so the control has to be at authoring time.
  • Where does a service mesh with mTLS fit relative to this programme?
    They are complementary layers. NetworkPolicy is L3/L4 and reduces blast radius at the network level, enforced by the CNI regardless of what the application does. A mesh adds cryptographic workload identity and L7 authorization, which NetworkPolicy cannot express. Running both is normal, but you must decide where each rule lives so the two policy sets do not drift into contradicting each other, and remember that a mesh's sidecar can be bypassed at the network level unless NetworkPolicy backs it.

saying these in an interview costs you the question

  • Applying cluster-wide default-deny in one change without observing existing traffic
  • Assuming policies are enforced without proving the CNI supports them
  • Rolling out egress restrictions before settling DNS and external-destination handling
  • Leaving new namespaces unrestricted because default-deny is namespaced and was never automated
  • Testing only that allowed traffic works, which never catches an over-broad policy

context