skip to content

After applying a Kubernetes NetworkPolicy that denies all egress by default in a namespace, pods start failing with name-resolution errors and long hangs even for destinations you explicitly allowed. Explain the cause and the correct remedy.

level: seniorimportance: must knowfreq 54%

answer

  1. Egress deny kills DNS first
  2. Allow UDP 53 and TCP 53
  3. Select CoreDNS pods by label, not ClusterIP
  4. Hangs then NXDOMAIN = resolver retries with ndots 5
  5. Probes are ingress from the node, not egress

basics

~20 s

Egress default-deny also blocks DNS. Pods cannot reach the cluster DNS service on UDP and TCP port 53, so every hostname lookup times out before the allowed connection is ever attempted. Add an egress rule permitting port 53 to the DNS pods in the kube-system namespace, matching both UDP and TCP.

solid answer

~60 s

Egress policies govern **all** outbound traffic from the selected pods, including their DNS queries to the cluster DNS service. With default-deny egress and no DNS allowance, resolution fails; the resolver retries and only then errors, which is why the symptom is slow hangs rather than instant failures - and why it looks like the allowed destination is broken when the destination was never contacted. The remedy is an explicit egress rule to the DNS pods: ```yaml egress: - to: - namespaceSelector: matchLabels: {kubernetes.io/metadata.name: kube-system} podSelector: matchLabels: {k8s-app: kube-dns} ports: - {protocol: UDP, port: 53} - {protocol: TCP, port: 53} ``` Include **TCP** as well as UDP - large or truncated responses fall back to TCP. Match the DNS pods by label rather than by the Service ClusterIP: policy is evaluated against the backend pod after DNAT, so an `ipBlock` holding the ClusterIP generally does not match. The same class of surprise hits egress to external endpoints whose IPs change, and to the API server - both need deliberate allowances.

code

yaml · 21 lines
yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-dns-egress
  namespace: shop
spec:
  podSelector: {}
  policyTypes: ["Egress"]
  egress:
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: kube-system
          podSelector:
            matchLabels:
              k8s-app: kube-dns
      ports:
        - protocol: UDP
          port: 53
        - protocol: TCP
          port: 53

go deeper

for a junior

Know that denying egress also denies DNS, and that a rule allowing port 53 to the cluster DNS pods is required.

for a middle

Write the rule correctly - both protocols, label-selected CoreDNS pods, applied to all pods in the namespace - and explain why an ipBlock on the ClusterIP fails.

for a senior

Diagnose by isolating name resolution from connectivity, sequence the rollout so the allowance precedes the deny, and cover API-server and probe traffic.

for a principal

Design the egress story: a gateway or proxy for third-party destinations so policy stays portable, observe-mode rollout to derive real allowances, and CI tests asserting the deny path.

## Why DNS is the first casualty Almost every outbound connection a pod makes starts with a name lookup against the cluster DNS Service (CoreDNS, usually reachable as `kube-dns` in `kube-system`). When you flip a namespace to egress default-deny, that lookup is outbound traffic like any other and is denied. The application never gets an address, so it never even attempts the connection you carefully allowed. The symptom is characteristic: **slow failures**, not fast ones. The pod's resolver works through `/etc/resolv.conf`, and Kubernetes injects `ndots: 5` plus a search-domain list, so a single lookup for `payments.internal` can generate several queries, each retried and each timing out. Latency of several seconds followed by "name or service not known" is the tell. ## Writing the DNS allowance correctly The DNS rule needs three things right: 1. **Both protocols.** UDP 53 handles most queries, but responses over 512 bytes (or exceeding the EDNS0 buffer) set the truncated flag and the resolver retries over TCP 53. Allowing only UDP produces intermittent failures for large record sets - a horrible, hard-to-reproduce bug class. 2. **Select the DNS pods by label, not by IP.** Policy is enforced at the destination pod after kube-proxy has DNATed the Service VIP to a CoreDNS pod IP, so a rule written as `ipBlock: {cidr: 10.96.0.10/32}` (the Service ClusterIP) typically never matches. Use `namespaceSelector` on `kube-system` combined with the DNS pods' label - conventionally `k8s-app: kube-dns` for CoreDNS - **in the same peer entry**, so it is an AND rather than "all of kube-system". 3. **Apply it to every pod that needs it.** Because policies are additive and namespaced, a single policy with `podSelector: {}` allowing DNS egress is the clean pattern: apply it alongside the default-deny in every namespace so DNS is never the thing that breaks. ## The other egress traps - **External endpoints with changing IPs.** Upstream NetworkPolicy can only express egress destinations as CIDRs, so allowing a SaaS API means allowing a range you do not control and that changes. Options: allow a broad CIDR (weak), route external traffic through an egress gateway or proxy you *can* name by pod label (strong), or use a CNI's DNS-aware policy CRD (effective but non-portable). - **The Kubernetes API server.** Workloads using the API (operators, controllers, anything with a client-go loop) need egress to it. Its address is often outside the pod network entirely, so it needs an `ipBlock`, and in managed clusters that endpoint can change - another argument for an egress proxy. - **Node-local services.** Some setups run node-local DNS caches, in which case the DNS peer is on the node IP rather than a pod, changing what you must allow. - **Health checks and probes are not egress from the pod** - kubelet probes originate on the node and are *ingress* to the pod. Blocking them accidentally with an ingress default-deny is the mirror-image mistake, and typically presents as pods flapping to NotReady after a policy rollout. - **Return traffic is not egress.** Policies are stateful in every practical implementation: allowing an inbound connection permits its replies without a matching egress rule. Candidates who add mirrored rules "for responses" misunderstand the model. ## Rolling it out without an outage 1. Ship the DNS-allow policy **before** the default-deny, so the ordering can never leave a gap. 2. Roll out per namespace, starting with a low-risk one, and watch application error rates and DNS metrics. 3. If the CNI supports policy logging or audit mode, run in observe-only first and derive the required allowances from real traffic rather than guessing. 4. Add a synthetic test in CI that runs a pod in the namespace and asserts both an allowed and a denied destination behave correctly - the denied assertion is what catches over-broad rules later. ## Diagnosing it in the moment From a debug pod in the affected namespace, a lookup of `kubernetes.default` distinguishes "DNS blocked" from "destination blocked" immediately. Testing a destination by IP that you allowed will succeed while the hostname form hangs - that asymmetry is essentially diagnostic on its own.

  • Why is allowing only UDP port 53 insufficient?
    Responses larger than the client's UDP buffer come back truncated, and the resolver then retries the same query over TCP port 53. Large record sets and zone transfers trigger this. Allowing only UDP therefore yields intermittent, size-dependent resolution failures that are painful to reproduce, so both protocols belong in the rule.
  • After applying an ingress default-deny, pods start flipping to NotReady. What happened?
    Kubelet readiness and liveness probes originate on the node, not from another pod, so they arrive as ingress to the pod and are blocked by the default-deny. Most CNIs special-case node-sourced probe traffic, but not all do. The fix is an ingress allowance covering the node's address range or the CNI's host-traffic construct, and the lesson is to roll ingress policies namespace by namespace while watching pod readiness.
  • How would you handle egress to a third-party SaaS API whose IPs change frequently?
    Do not chase CIDRs. Route that traffic through an egress proxy or gateway running in the cluster, allow egress only to that proxy by pod label, and let the proxy enforce the allowed hostnames. That keeps the NetworkPolicy stable and portable, gives you a single place for logging and TLS termination policy, and avoids committing to a vendor-specific DNS-aware policy CRD that would not survive a CNI change.

It is like cutting the phone line to directory enquiries while insisting your outbound calls to head office are still permitted: the number you are allowed to dial is one you can no longer look up.

saying these in an interview costs you the question

  • Forgetting that egress default-deny also blocks the pod's DNS queries
  • Allowing UDP 53 only and treating the resulting intermittent failures as a CoreDNS bug
  • Writing the DNS rule against the kube-dns Service ClusterIP with ipBlock instead of selecting CoreDNS pods by label
  • Adding mirrored egress rules for reply traffic, misunderstanding that enforcement is stateful
  • Applying the default-deny before the DNS allowance and creating a self-inflicted outage window

context