skip to content

Services and Networking

How traffic reaches a pod: ClusterIP, NodePort and LoadBalancer Services, CoreDNS names, Ingress and Gateway API routing, NetworkPolicy, plus the kube-proxy and CNI datapaths. Interviewers start here because 'connection refused' is the standard exercise.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

Kubernetes replaced the older Endpoints API with EndpointSlices for tracking the backends of a Service. What problem did that solve, and how does the newer object differ in structure?

level: middleimportance: should knowfreq 48%

basics

~20 s

The old Endpoints object crammed every backend of a Service into one resource, so one pod change rewrote the whole object and pushed it to every watcher - painful at thousands of endpoints. EndpointSlices shard the same data into capped chunks (100 endpoints by default), so a change updates one small slice, and they add per-endpoint conditions and topology fields.

open as a page

How do you configure the cloud load balancer behind a Kubernetes LoadBalancer Service, for example internal-only, and what does spec.loadBalancerClass change?

level: middleimportance: should knowfreq 46%

basics

~20 s

The Service spec has only a few portable load-balancer fields, so provider-specific settings such as internal-only, balancer flavour or TLS are set with the provider's annotations on the Service. spec.loadBalancerClass hands the Service to a different load-balancer implementation instead of the default one.

open as a page

What does the cloud-controller-manager's service controller do with a Kubernetes type=LoadBalancer Service, and what role do the NodePorts underneath play?

level: middleimportance: should knowfreq 52%

basics

~20 s

The service controller adds a cleanup finalizer, has the provider build a balancer for the eligible nodes, writes the address into status, and updates or deletes the balancer as things change. Classic balancers forward to each node's NodePort.

open as a page

What does setting clusterIP: None on a Kubernetes Service do, and when would you choose that over a Service with a normal virtual IP?

level: middleimportance: should knowfreq 52%

basics

~20 s

It makes the Service headless: no virtual IP is allocated and no load balancing happens. Cluster DNS instead returns the individual pod IPs for the Service name, letting clients see and address each backing pod directly.

open as a page

Compare an encapsulated overlay pod network (for example VXLAN) with a routed pod network (for example one that advertises pod routes over BGP): how does each move a packet between two nodes, and what would make you choose one over the other?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An overlay wraps each pod packet inside a node-to-node UDP packet, so the underlay never needs to know pod addresses — but it costs header bytes (MTU) and hides traffic. A routed network advertises each node's pod range as real routes, so packets travel unwrapped at native MTU, but the underlay must accept those routes.

open as a page

A newly scheduled pod stays in ContainerCreating and its events show a failure to set up the sandbox network. How would you diagnose it, and what are the likely causes?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Read the pod events and note the node, then check that node: is the network plugin's DaemonSet pod healthy there, is a valid config present in /etc/cni/net.d, are the binaries in /opt/cni/bin, does the node have a pod CIDR assigned, and is its address pool exhausted or full of stale allocations?

open as a page

Workloads in a cluster intermittently fail name resolution: most lookups succeed, but a small fraction time out after roughly five seconds. How would you investigate, and what causes are plausible?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Roughly five seconds is the C library resolver's retry timeout, so a query is being dropped rather than answered slowly. Look for dropped UDP queries from parallel A and AAAA lookups colliding in connection tracking, and for overloaded or CPU-throttled DNS pods. Confirm with DNS server metrics, then deploy a node-local cache.

open as a page

In a multi-tenant Kubernetes cluster using the Gateway API, how does an HTTPRoute in one namespace get attached to a Gateway in another, and what controls does the cluster operator have over which namespaces are allowed to attach?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Attachment is two-sided. The route names the Gateway in parentRefs (with its namespace); the Gateway's listener declares allowedRoutes.namespaces as Same, All, or Selector plus permitted route kinds. Both must agree, and hostnames must intersect, or the route's status shows it was not accepted.

open as a page

A Kubernetes StatefulSet's replicas report Ready only after forming a quorum, and they never do. Why, and what does publishNotReadyAddresses on its headless Service change?

level: seniorimportance: should knowfreq 39%

basics

~20 s

Cluster DNS publishes only ready endpoints, so replicas waiting on a quorum cannot resolve each other: a deadlock. publishNotReadyAddresses makes the EndpointSlice controller mark every endpoint ready, so peer names resolve as soon as pods have IPs.

open as a page

Much real Kubernetes Ingress behaviour — path rewriting, timeouts, body-size limits, sticky sessions — is configured through controller-specific annotations rather than fields in the Ingress spec. Why is that, and what problems does it create in a shared cluster?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The Ingress spec only models host and path routing plus TLS, so everything else had to go somewhere. Annotations are untyped metadata strings: unvalidated, undiscoverable, controller-specific, and in a shared proxy some of them let one tenant influence configuration affecting everyone.

open as a page

A cluster runs two ingress controllers, one internal and one internet-facing. How does a given Kubernetes Ingress object end up served by exactly one of them, and what goes wrong if that selection is left ambiguous?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Each controller installs an IngressClass whose spec.controller holds its own identifier. An Ingress selects one via spec.ingressClassName. If it is omitted, only a class annotated as default applies; with no default nothing serves the object, and with two defaults or a mismatch you get silent non-service or accidental public exposure.

open as a page

Requests to a hostname served by a Kubernetes ingress controller come back with 404 from the controller itself rather than from the application. Walk through how you would diagnose it.

level: seniorimportance: should knowfreq 44%

basics

~20 s

A 404 from the controller means no rule matched. Check that the Ingress has the right class and was claimed, that the Host header matches spec.rules[].host exactly, that the path and pathType actually cover the request, and that the backend Service and its EndpointSlices resolve. Confirm the request reached the expected controller.

open as a page

A Kubernetes Service exposed outside the cluster has externalTrafficPolicy set to Cluster by default, and it can be set to Local instead. What changes when you switch it, and what are the tradeoffs?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Cluster forwards external traffic to a pod on any node, adding a hop and rewriting the source IP to the node's. Local only serves traffic from pods on the receiving node, preserving the real client IP but dropping traffic on pods-less nodes and risking uneven load.

open as a page

You are asked to introduce pod-level network segmentation across an existing multi-tenant Kubernetes cluster that today runs entirely unrestricted. How would you plan and sequence that programme so you end up with default-deny everywhere without causing outages?

level: principalimportance: should knowfreq 36%

basics

~20 s

Verify the CNI actually enforces policy, then observe real traffic before restricting it. Roll out per namespace: ship DNS and platform allowances first, derive workload allowances from observed flows, then apply default-deny ingress, then egress. Make the default-deny part of namespace provisioning so new namespaces are born closed, and test deny paths in CI.

open as a page

In Kubernetes, how do a plain Pod's spec.hostname and spec.subdomain fields give it its own DNS name, and what else must exist?

level: juniorimportance: nice to knowfreq 27%

basics

~20 s

Set spec.hostname to a DNS label and spec.subdomain to the name of a headless Service in the same namespace that selects the pod. Once the pod is ready, cluster DNS answers <hostname>.<subdomain>.<namespace>.svc.<cluster-domain> with its IP.

open as a page

Give an overview of Flannel, Calico and Cilium as Kubernetes pod network implementations: what each is built on, and what would push you toward one rather than another?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

Flannel is the minimal option — usually a VXLAN overlay, simple, no policy enforcement. Calico offers routed (BGP) or encapsulated modes with a mature policy engine on iptables or eBPF. Cilium builds its datapath on eBPF, adds identity-aware and layer-7 aware policy, can replace kube-proxy, and brings its own flow observability.

open as a page

How do you make a Kubernetes Service send repeated requests from the same client to the same backing pod, and what are the limitations of that mechanism?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Set spec.sessionAffinity: ClientIP, optionally tuning sessionAffinityConfig.clientIP.timeoutSeconds (default 10800). It hashes the client source IP to a pod. It is L4 only, breaks behind NAT or a proxy, and gives no stickiness when the pod dies or the pod set changes.

open as a page

In Kubernetes, how does a Service's internalTrafficPolicy: Local differ from trafficDistribution: PreferSameNode when keeping calls on the caller's node?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

internalTrafficPolicy: Local is a hard rule: ClusterIP traffic goes only to endpoints on the caller's node, and is dropped if there are none. trafficDistribution: PreferSameNode is a preference: it uses a same-node endpoint when one is ready and otherwise falls back to any ready endpoint.

open as a page

On a bare-metal Kubernetes cluster, how does MetalLB give LoadBalancer Services an external IP, and how do you choose between its Layer 2 and BGP modes?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

MetalLB's controller assigns an address from an IPAddressPool to each LoadBalancer Service. Its speakers then attract traffic for that address. In Layer 2 mode one node answers ARP/NDP for it; in BGP mode nodes advertise it to routers, which spread traffic across nodes with ECMP.

open as a page

You own cluster DNS for a large, busy multi-tenant Kubernetes cluster. How do you make name resolution scale and stay reliable, and how do you decide how much capacity it needs?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Treat cluster DNS as tier-0: model query volume from pods times lookups per request times search-list amplification, add a node-local cache to absorb repeats, scale and protect the DNS Deployment with requests, anti-affinity, a critical priority class and a disruption budget, and constrain tenants from bypassing or inflating it.

open as a page

You run a multi-tenant Kubernetes platform heading toward tens of thousands of Services and very high pod churn. How would you decide between keeping kube-proxy in iptables mode, moving to IPVS or nftables, or removing kube-proxy in favour of an eBPF dataplane?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Decide on measured convergence and failure behaviour, not benchmarks. Keep iptables while endpoint-sync latency stays inside your rollout error budget; move to nftables or IPVS when sync time is the bottleneck; drop kube-proxy for eBPF only when you accept coupling the Service dataplane to one CNI and can staff kernel-level debugging.

open as a page

showing 31–51 of 51