Pods in one Kubernetes namespace suddenly cannot reach a Service in another namespace, and the failure presents as a hang rather than a connection refusal. How do you determine whether a NetworkPolicy is responsible, and what mistakes commonly cause this?
answer
- Hang = dropped (policy); refused = reached a pod
- No policy selects the pod → allow all; one does → deny by default, that direction only
- Additive only — no deny rules, no ordering
- Both client egress and server ingress must allow
- Policy ports = pod/targetPort, not Service port
basics
~20 sSilent drops point at NetworkPolicy. List policies in both namespaces and see which select the client and server pods: once any policy selects a pod for a direction, everything not explicitly allowed in that direction is denied. Check ingress and egress separately, and confirm the CNI enforces policy at all.
solid answer
~60 sA hang, rather than a refusal, means packets are being dropped rather than rejected — NetworkPolicy is the prime suspect. Method: 1. **List policies on both sides**: `kubectl get netpol -n <client-ns>` and `-n <server-ns>`, then check which ones select the pods via their `podSelector`. 2. **Reason per direction.** A policy with `policyTypes: [Ingress]` selecting the server means all ingress not allowed is denied; a policy with `Egress` selecting the client means all its outbound traffic must be allowed explicitly. Both must permit the flow. Policies are purely additive — there is no deny rule and no ordering. 3. **Check the cross-namespace clause.** Reaching another namespace needs `namespaceSelector` on the server-side ingress rule; a `podSelector` alone only matches pods in the policy's own namespace. 4. **Check ports.** Rules match the **pod's** port (the targetPort), not the Service port. 5. **Do not forget DNS.** A default-deny egress policy blocks port 53 first, so everything looks broken. Confirm the CNI enforces policy — with a plugin that ignores it, the objects exist and do nothing.
code
yaml · 17 linesapiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-egress
namespace: frontend
spec:
podSelector: {}
policyTypes: [Egress]
egress:
- to:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: kube-system}
podSelector:
matchLabels: {k8s-app: kube-dns}
ports:
- {protocol: UDP, port: 53}
- {protocol: TCP, port: 53}go deeper
Know that a NetworkPolicy only allows traffic, that selecting a pod makes that direction deny-by-default, and that you list policies with kubectl get netpol.
Explain per-direction semantics, additive rules, cross-namespace selectors, and the DNS egress exception.
Use the drop-versus-reject symptom, test pod IPs versus Service IPs, read the AND/OR selector subtlety correctly, and use CNI flow tooling to name the dropping policy rather than guessing.
Treat policy as a platform contract: default-deny bundled with DNS allows and templated rules, flow observability as a prerequisite for enforcement, and a rollout path that does not silently break tenants.
## Why the symptom is diagnostic TCP failure modes carry information. A **refusal** (RST) means a host received the SYN and nothing was listening — you got all the way to a pod. A **timeout** means the packet vanished: no route, no endpoint to DNAT to, or a firewall that drops rather than rejects. NetworkPolicy enforcement drops, so "it just hangs" plus "it used to work" plus "a policy was added recently" is a strong signature. ## The model you must hold NetworkPolicy is a namespaced object that selects pods and describes **allowed** traffic. Four rules govern everything: 1. **Default is allow.** A pod not selected by any policy accepts and initiates everything the CNI permits. 2. **Selection flips the default for that direction only.** The moment one policy with `policyTypes: [Ingress]` selects a pod, that pod's ingress becomes deny-by-default and only rules in policies selecting it are permitted. Egress is independent: a pod can be ingress-restricted and egress-unrestricted at the same time. A policy that lists `policyTypes: [Ingress, Egress]` with an empty `egress: []` is a full egress deny. 3. **Policies are additive.** Multiple policies form a union of allowances. There is no deny rule, no priority, no ordering — so you cannot "override" an allow, you can only stop allowing it. 4. **Both ends must permit the flow.** Client egress must allow the destination *and* server ingress must allow the source. This is where cross-namespace failures usually live: each team writes a policy for their own namespace and neither anticipates the other. ## Cross-namespace selection specifics In an ingress rule, `from` entries can carry `podSelector`, `namespaceSelector`, `ipBlock`, or a combination. The distinction that bites: ```yaml from: - podSelector: {matchLabels: {app: web}} # pods in THIS namespace only - namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: frontend}} podSelector: {matchLabels: {app: web}} # app=web pods in namespace frontend - namespaceSelector: {...} - podSelector: {...} # TWO separate entries = OR, much broader ``` A `namespaceSelector` and `podSelector` inside the **same** list item is an AND; as two list items it is an OR, which accidentally allows all pods in the selected namespace plus all local pods with that label. Also note `kubernetes.io/metadata.name` is applied automatically to every namespace by the API server, so you can select a namespace by name without adding custom labels. ## The DNS trap The single most common self-inflicted outage: a namespace adopts a deny-by-default egress policy for security, and every application immediately fails because it cannot resolve names. Egress rules must explicitly allow UDP **and** TCP port 53 to the cluster DNS pods: ```yaml egress: - to: - namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: kube-system}} podSelector: {matchLabels: {k8s-app: kube-dns}} ports: - {protocol: UDP, port: 53} - {protocol: TCP, port: 53} ``` The tell is that *everything* breaks at once, including calls that previously worked to unrelated destinations. ## Ports are pod ports NetworkPolicy operates on pods and their real ports. Service `port: 80` with `targetPort: 8080` means the rule must allow **8080**, because DNAT happens before policy evaluation on the receiving side. Writing `port: 80` in the policy produces a policy that looks right and drops everything. ## Enforcement is the CNI's job NetworkPolicy objects are inert unless the network plugin implements them. Calico, Cilium, Antrea and the cloud CNIs generally do; plain Flannel does not. The failure mode is opposite in each direction: with a non-enforcing CNI, policies you wrote for security do nothing (and you should know that); with an enforcing one, a policy applied by a platform team silently breaks a team that has never heard of NetworkPolicy. ## Diagnosis sequence 1. Rule out the simpler causes first — does the Service have ready endpoints, is the targetPort right, does DNS resolve? A hang has several sources and policy is only one. 2. `kubectl get netpol -A` and identify which policies select the client and the server: `kubectl describe netpol <name> -n <ns>` prints the resolved pod selector and rules. 3. `kubectl get pods --show-labels` on both sides and check them against those selectors by hand. 4. Test the layers: from a debug pod in the client namespace, try the server's **pod IP** on its real port. If the pod IP fails too, it is not the Service abstraction. Try a destination in the same namespace and a destination outside the cluster — the pattern of what works narrows which direction is restricted. 5. Use the CNI's own tooling where available — Cilium's `hubble observe --verdict DROPPED` or Calico's flow logs name the policy that dropped the packet, which turns an inference into a fact. 6. As a controlled experiment, add a temporary narrowly-scoped allow policy (specific pod labels and port, short-lived) and see whether traffic flows. Never do this by deleting the platform's default-deny policy. ## Prevention Ship default-deny with a DNS allow as a bundle, template the common allow rules, and test policy changes in a namespace with representative traffic. Because policies are additive and silent, the only reliable feedback loop is observing dropped flows — so a CNI with flow visibility is worth a great deal in a policy-enforcing cluster.
- Why is a hang a stronger indicator of NetworkPolicy than a connection refusal?Because policy enforcement drops packets rather than rejecting them, so the client's SYN goes unanswered and the connection sits until timeout. A refusal means a SYN actually reached a host whose kernel replied with a reset, which proves the network path worked and points instead at a wrong port or an application not listening. Distinguishing the two early saves you from auditing policies when the real problem is a targetPort.
- In a NetworkPolicy ingress rule, what is the difference between listing namespaceSelector and podSelector as one item versus two?Inside a single 'from' list item they are ANDed: the traffic must come from a pod matching the podSelector that lives in a namespace matching the namespaceSelector. As two separate list items they are ORed: any pod in the selected namespaces, plus any pod in the policy's own namespace matching the podSelector. The two-item form is far broader than intended and is one of the most common ways an ostensibly tight policy ends up permitting a whole namespace.
- A team applies a NetworkPolicy and nothing changes — traffic they meant to block still flows. What would you check?Whether the CNI actually enforces NetworkPolicy. The API accepts and stores the objects regardless, so with a plugin that does not implement them (plain Flannel, for example) the policy is inert. Confirm the installed CNI and its policy support, then verify that a policy does select the intended pods by comparing its podSelector with the pods' labels, and remember that a policy restricting only Ingress leaves egress entirely unrestricted.
NetworkPolicy is a guest list, not a bouncer with a grudge: while no list exists everyone walks in, but the moment one is posted at a door, anyone not written on it is quietly turned away — and the door on the other side of the corridor has its own separate list.
saying these in an interview costs you the question
- Expecting NetworkPolicy to have deny rules or evaluation order, rather than being purely additive allows
- Writing the Service port in a policy instead of the pod's targetPort
- Applying default-deny egress without an explicit DNS allow, then blaming CoreDNS
- Assuming a podSelector in an ingress rule can match pods in another namespace without a namespaceSelector
- Assuming policies are enforced without checking that the CNI implements them