You own cluster DNS for a large, busy multi-tenant Kubernetes cluster. How do you make name resolution scale and stay reliable, and how do you decide how much capacity it needs?
answer
- tier-0: every request starts with a lookup
- QPS = pods × lookups × amplification ÷ cache
- node-local cache first, replicas last
- requests + priority class + anti-affinity + PDB
- split internal resolution from external forwarding
basics
~20 sTreat cluster DNS as tier-0: model query volume from pods times lookups per request times search-list amplification, add a node-local cache to absorb repeats, scale and protect the DNS Deployment with requests, anti-affinity, a critical priority class and a disruption budget, and constrain tenants from bypassing or inflating it.
solid answer
~60 sStart from the fact that almost every request in the cluster begins with a lookup, so DNS is a **tier-0 dependency** whose failure is indistinguishable from a total outage. **Capacity model.** Queries per second ≈ active pods × lookups per request × amplification from `ndots` expansion and parallel A/AAAA queries, divided by whatever caching absorbs. Measure the real ratio rather than assuming; the amplification factor is frequently the dominant term. **Structural levers, in order of leverage:** a node-local cache on every node (absorbs repeats, removes the cross-node hop and the UDP query race, forwards upstream over TCP); sensible cache TTLs and stale-serving during upstream failure; then scaling the central DNS Deployment proportionally to cluster size. **Protection.** Resource requests so it is never CPU-throttled, a system-critical priority class, anti-affinity across nodes and zones, a disruption budget, and a rollout that never removes a large share of capacity at once. **Governance.** Default sensible `ndots` per workload class through admission, and constrain workloads that bypass cluster DNS. **Verification.** Latency and response-code objectives, plus a game day that kills replicas and watches what the applications actually do.
code
yaml · 29 linesapiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: coredns
namespace: kube-system
spec:
minAvailable: 2
selector:
matchLabels:
k8s-app: kube-dns
---
# excerpt from the CoreDNS Deployment
spec:
template:
spec:
priorityClassName: system-cluster-critical
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
k8s-app: kube-dns
containers:
- name: coredns
resources:
requests:
cpu: 200m
memory: 128Migo deeper
Know that cluster DNS runs as a normal deployment that can be overloaded, and that more pods means more lookups.
Be able to name the levers — caching, more replicas, lowering ndots — and explain why caching close to the client helps most.
Bring a capacity model with real measurements, order the interventions by leverage, and cover protection: requests, priority class, spread, disruption budget and safe rollouts.
Argue the whole position: DNS as a tier-0 dependency with an explicit objective, blast-radius separation between internal resolution and external forwarding, admission-level governance of tenant resolver behaviour, and game-day validation of client failure handling.
## Framing: DNS is on every request path In a Kubernetes cluster, service-to-service calls, database connections, object-store access and outbound API calls all start with a name lookup. That makes cluster DNS a dependency of essentially every request, with no natural fallback. Two consequences follow: its availability objective must be at least as strong as anything that depends on it, and its failure modes must be *partial-tolerant*, because a partial DNS failure presents to teams as random unexplained errors everywhere at once — the most expensive kind of incident to triage. ## Build a capacity model, do not guess A usable model: ``` QPS ≈ active_pods × lookups_per_request × requests_per_second_per_pod × amplification_factor ÷ cache_absorption ``` - **lookups per request** is 1 only for applications that reuse connections; naive clients resolve per request. - **amplification_factor** is the killer. Under the cluster default of `ndots:5`, an external hostname walks the search list before the correct query, and the A and AAAA pair roughly doubles the packet count. Six-to-eight-fold amplification for egress-heavy workloads is ordinary. - **cache_absorption** is everything you deploy in front: node-local caches and in-process resolver caches. Measure the real numbers from the DNS server's own request metrics and the share of responses that are NXDOMAIN. A large NXDOMAIN share is direct evidence that the amplification term dominates, and it says the cheapest capacity win is reducing queries rather than adding replicas. ## Lever 1 — a node-local cache The highest-leverage change is a per-node caching resolver that pods query at a link-local address. It: - serves repeats from the node, removing the cross-node network hop and the address translation to a DNS pod; - absorbs the repeated negative answers from search-list expansion; - forwards upstream over TCP, which removes the UDP query-collision race that produces multi-second resolution stalls; - keeps working through brief central-DNS disruptions, converting a hard failure into a cache-warm degradation. Its cost is one more per-node component in the critical path, with its own upgrade and failure story — a node-local cache that dies takes DNS on that node with it, so it needs the same operational rigour as the kubelet. ## Lever 2 — caching policy Inside the DNS server, decide TTLs deliberately. Longer caching of external names cuts upstream traffic but slows propagation of real changes. Serving **stale entries** when the upstream is unreachable is usually the right trade for a platform: a slightly outdated address beats a resolution failure. Negative-answer caching bounds the damage from misconfigured clients hammering names that do not exist. ## Lever 3 — scale and protect the central deployment Only after caching should you buy capacity with replicas: - **Scale proportionally** to nodes or cores rather than pinning a fixed count, since clusters grow silently past their original sizing. - **Requests, not just limits.** A DNS server throttled at its CPU limit drops queries in bursts, which is the hardest symptom to attribute. - **System-critical priority class**, so it is never the thing evicted under node pressure. - **Anti-affinity across nodes and zones**, so no single node or zone failure takes a large share of capacity. - **A disruption budget**, so drains and maintenance cannot remove replicas below a safe level. - **A rollout that preserves capacity**: with a small replica count, a rolling update removes a large fraction at once. Surge and readiness gating must keep enough serving capacity throughout. ## Lever 4 — separate blast radii At scale, consider splitting responsibilities: one path resolves in-cluster names from the API-server-derived view, another handles forwarding to external resolvers. External resolvers fail, rate-limit and slow down for reasons entirely outside the cluster; when they do, you want in-cluster resolution to be unaffected. Keeping the two on separate instances or separate forwarding policies stops an outside dependency from taking down internal service discovery. ## Lever 5 — governance in a multi-tenant cluster Tenants can hurt shared DNS in three ways: applications that resolve on every request; workloads left at a default `ndots` while making heavy external calls; and workloads that bypass cluster DNS entirely by pinning their own resolver, which escapes node-local caching and any forwarding policy or visibility the platform relies on. The platform answers are admission-time: default resolver options per workload class, constrain or require review for pods that supply their own resolver configuration, and publish the expectation that clients reuse connections. Attribute query volume per namespace so the conversation with a heavy tenant is evidence-based. ## Verification and objectives Define what "working" means: a latency objective for resolution and a ceiling on server-failure responses, measured continuously. Then test the failure modes rather than assuming them — a game day that deletes replicas, blocks the upstream resolver, and drains a node running a cache, while watching what applications actually do. The findings are usually about clients, not servers: services that cache resolutions for their entire lifetime and keep dialling dead addresses, or that treat a resolution failure as fatal instead of retrying. ## The judgement to communicate Order matters. Reduce queries, then cache close to the client, then add capacity, then protect what you have from your own operations. Adding replicas first is the common and least effective response, because it scales the symptom rather than the cause — and a cluster whose DNS load is mostly failed search-suffix attempts is paying for capacity it did not need.
- A team proposes tripling the DNS replica count to fix latency. What would you want to see before agreeing?Evidence of where the queries come from: the share of responses that are NXDOMAIN, per-namespace query attribution, and whether the servers are actually CPU-saturated or merely throttled at their limits. If most traffic is search-suffix expansion for external names, replicas scale the waste rather than removing it, and a node-local cache plus resolver-option defaults gives a far bigger reduction for less permanent cost.
- What is the risk of caching aggressively at the node level, and how do you bound it?Stale answers. A pod can keep receiving an address for a workload that has moved or been removed until the cached entry expires, which matters most for headless Services where the answer is the membership list. Bound it with short TTLs for in-cluster records while allowing longer ones for stable external names, and reserve stale-serving for the case where the upstream is unreachable, so staleness is a deliberate degradation rather than the steady state.
- Why separate in-cluster resolution from forwarding to external resolvers?They have unrelated failure modes. External resolvers can slow down, rate-limit or fail for reasons entirely outside the cluster, and if the same instances serve both, that backpressure delays in-cluster service discovery too. Splitting them — separate instances or separate forwarding policy with strict timeouts — keeps an outside dependency from turning into an internal service-discovery outage.
It is the switchboard of the building: nobody notices it until it stops, and then every conversation in the building fails at once. You add extension directories on each floor before hiring more operators.
saying these in an interview costs you the question
- Adding replicas as the first response without measuring query composition
- Setting tight CPU limits on the DNS pods and not noticing throttling-induced drops
- Leaving the DNS deployment without a priority class or disruption budget, so drains take it out
- Assuming client applications retry resolution failures gracefully without ever testing it
- Letting tenants pin their own resolvers, silently bypassing node-local caching and forwarding policy