skip to content

Kubernetes

8 roadmaps448 questionsupdated

Declaring workloads and letting controllers reconcile the cluster toward that spec, exposing them through Services and Ingress, and diagnosing it when the cluster disagrees with you. It is the default platform question in devops, SRE and backend loops.

on this pageshow

guide

overview

~2 min

Kubernetes runs containers by comparing what you declared with what exists and letting controllers close the gap. Interviewers treat it as the default platform subject because most backend, SRE and DevOps roles touch it, and because the distance between "I wrote a Deployment" and "I know why this pod is Pending" is easy to hear. A strong answer names the object involved, the controller that acts on it, and the signal you would read when the two disagree. The hub follows a workload from declaration to traffic. [Pods and workloads](/topics/cloud-kubernetes-pods-workloads) covers what gets scheduled and which controller keeps it alive, and [API object conventions](/topics/cloud-kubernetes-api-objects) is the shared grammar every other object uses: labels, metadata, apply and ownership. [Services and networking](/topics/cloud-kubernetes-services-networking), [config and secrets](/topics/cloud-kubernetes-config-secrets) and [storage and volumes](/topics/cloud-kubernetes-storage-volumes) are what a running pod needs from outside itself: traffic, settings and data that outlive it. [Scaling and scheduling](/topics/cloud-kubernetes-scaling-scheduling) decides where pods land and how many there are. Underneath sits [cluster architecture](/topics/cloud-kubernetes-architecture), the control plane and node agents that make the loop turn, and [operators and CRDs](/topics/cloud-kubernetes-operators-crds) extend that loop to your own types. [Security](/topics/cloud-kubernetes-security) and [tenancy and quotas](/topics/cloud-kubernetes-namespaces-tenancy) decide who may do what, and how much. The operator's side is [troubleshooting](/topics/cloud-kubernetes-troubleshooting) and [cluster lifecycle](/topics/cloud-kubernetes-cluster-lifecycle), and [adoption and migration](/topics/cloud-kubernetes-adoption) asks whether a workload belongs on a cluster at all. Junior rounds check the object model: what a pod is, how a Deployment differs from a StatefulSet, what each Service type exposes, what a failing probe triggers. Middle rounds move to behaviour under change, such as rollouts, autoscaling, config updates and resource pressure. Senior and staff rounds are usually one incident or one design: a Service that answers nothing, a node that keeps evicting, an upgrade plan, a tenancy model. Learn pods, controllers and labels first; nearly every later answer is a consequence of them. Networking and configuration come next, because a first deployment usually breaks there. Architecture makes sense once you have seen the objects it serves, and troubleshooting is where all of it is tested together.

primer

### You declare, controllers converge Most objects have a spec you write and a status the cluster reports. Controllers are loops that watch objects, compare the two, act to narrow the difference, and try again later. Very little in the system is a one-shot command: a deleted pod comes back because a controller still wants it. A "why did Kubernetes do that" question is usually answered by naming which controller acted and what spec it was chasing. ### The API server is the only door Controllers, the scheduler, kubelets and `kubectl` do not talk to each other directly. They read and write objects through kube-apiserver, which alone persists them in etcd, and they learn about changes by watching. That single path is why authentication, authorisation, admission and audit can live in one place, and why losing the API server freezes change without stopping the containers already on the nodes. ### Pods are disposable; controllers give them a job A pod is never moved, only replaced, and a replacement gets a new IP and a fresh filesystem. The workload controller around it carries the intent: interchangeable replicas, stable identities with their own storage, one copy per node, or work that runs to completion. Choosing the wrong one rarely fails on day one; it fails during a rollout, a node loss or a scale-down. ### Labels are the wiring Objects seldom hold direct references to each other. Services, ReplicaSets, NetworkPolicies, PodDisruptionBudgets and affinity rules find their targets by label selector, so a label typo silently disconnects things instead of raising an error. Ownership runs the other way, through owner references that decide what is cleaned up when a parent is deleted. ### Requests schedule, limits enforce The scheduler places pods by their declared requests against a node's allocatable capacity, not by what they actually use. At runtime the node enforces limits, and the two resources fail differently: a container over its memory limit is killed, one over its CPU limit is slowed. The same numbers shape which pods go first when a node comes under pressure, so sizing is a reliability decision as much as a cost one. ### Declaring is not doing Several objects are only records until something acts on them. An Ingress needs an ingress controller, a `LoadBalancer` Service needs a cloud or bare-metal implementation, a custom resource needs its operator, and a PersistentVolumeClaim needs a provisioner or a matching volume. Many "it does nothing" incidents are a missing or unhealthy controller, not a wrong manifest. ### Networking is virtual, and readiness steers it A Service is a stable name and virtual address in front of a changing set of pods, and only ready pods are placed behind it, so probes shape traffic as much as they report health. DNS, the Service proxy layer and the CNI plugin each own one layer; debugging a connection means checking them in order instead of guessing. ### Namespaces organise; they do not isolate A namespace scopes names, access bindings, quotas and policies. It does not give a tenant its own nodes, kernel or control plane, and many resources are cluster-scoped. Real separation is layered: RBAC for the API, NetworkPolicy for traffic, Pod Security for what a container may do on its node, and sometimes separate node pools or clusters.

Pod
The smallest unit Kubernetes schedules: one or more containers placed together on a node, sharing a network identity and volumes. It is replaced, never moved.
Controller
A control loop that watches objects through the API server and acts to move actual state toward the declared spec, retrying until the two agree.
Deployment
A workload controller for interchangeable replicas. It manages one ReplicaSet per pod template version, so changes roll out gradually and can be rolled back.
StatefulSet
A workload controller that gives each replica a stable ordinal name, its own persistent volume claim, and ordered start, update and shutdown.
Label selector
A query over labels that one object uses to find others. Services, ReplicaSets, NetworkPolicies and disruption budgets all pick their targets this way.
Service
A stable virtual address and DNS name in front of the ready pods its selector matches, with types that expose it inside or outside the cluster.
EndpointSlice
The object listing a Service's current backend addresses and their readiness. Proxies read it, and it changes as pods start, stop or fail probes.
Ingress
An object describing HTTP host and path routing to Services, acted on by a separately installed ingress controller. Gateway API is its role-oriented successor.
kube-apiserver
The control plane's REST front end. It authenticates, authorises, admits and validates every request, and is the only component that reads and writes etcd.
etcd
The consistent, replicated key-value store holding all cluster state. Losing it without a backup means losing the cluster's record of everything declared.
kubelet
The agent on each node that runs the pods assigned to it through the container runtime, executes their probes, and reports node and pod status.
Requests and limits
Per-container CPU and memory figures: requests are what scheduling reserves, limits are the runtime ceiling. Together they also set the pod's QoS class.
PersistentVolumeClaim
A pod-side request for storage of a given size, access mode and class, bound to a PersistentVolume whose lifetime is independent of any pod.
CustomResourceDefinition
An object that registers a new resource type with the API server, giving it storage, endpoints and RBAC; behaviour comes from a controller you supply.
Operator
A controller plus custom resources that encode how to run one specific application, including day-2 work such as upgrades, backups and failover.
Admission control
The stage after authorisation where built-in plugins, webhooks or policies may modify or reject an object before it is stored.
RBAC
Role-based access control: Roles and ClusterRoles list allowed verbs on resources, and bindings grant them to users, groups or ServiceAccounts. Rules only ever add permissions.
ServiceAccount
The identity a pod uses when it calls the API server, usually presented as a short-lived token mounted into the pod and authorised through RBAC.
PodDisruptionBudget
A cap on how many replicas of an application voluntary disruptions, such as node drains, may take down at once. Hardware or node failures ignore it.
Namespace
A scope for names, access bindings, quotas and policies inside one cluster. It groups resources but does not isolate their workloads at runtime.

### One apply, followed through Follow one `kubectl apply` of a Deployment. The request reaches kube-apiserver, which authenticates the caller, checks RBAC, runs admission and stores the object in etcd. The Deployment controller sees it through a watch and creates a ReplicaSet for the current pod template; the ReplicaSet controller creates pods to match the replica count. The scheduler notices pods with no node, picks one using requests, affinity, taints and spread rules, and records the binding. The kubelet on that node sees the assignment, asks the container runtime to pull images and start containers, has the CNI plugin give the pod an address, mounts its volumes, config and secrets, and starts the probes. Traffic arrives through the same machinery. Once a pod is ready it joins the Service's EndpointSlices; kube-proxy on every node, or a CNI datapath that replaces it, steers the Service address to those backends, and CoreDNS gives the Service a name. An Ingress or Gateway controller watches its own objects and configures the proxy that brings outside traffic in. ### Where the other sections attach - **Autoscalers** edit fields other controllers already act on: the HPA changes a workload's replica count, and node autoscalers add capacity when pods stay unschedulable. - **Operators** repeat the Deployment controller's pattern for your own types, registered through CRDs and often guarded by admission webhooks. - **Troubleshooting** walks the chain backwards: object, events, scheduling, kubelet, container, endpoints, network. - **Cluster lifecycle** keeps the chain itself healthy: certificates, etcd backups, and upgrades in the order the skew rules allow. ### The wiring in one manifest The objects meet through labels rather than references. Below, the Deployment's selector must match its own template, the Service finds the pods by the same label, readiness decides when each one receives traffic, and requests place the pod while the memory limit caps it on the node: ```yaml kind: Deployment # apps/v1 metadata: {name: checkout} spec: replicas: 3 selector: {matchLabels: {app: checkout}} template: metadata: {labels: {app: checkout}} spec: containers: - name: app image: registry.example.com/checkout:1.4.2 resources: {requests: {cpu: 250m, memory: 256Mi}, limits: {memory: 256Mi}} readinessProbe: {httpGet: {path: /ready, port: 8080}} --- kind: Service # v1 metadata: {name: checkout} spec: {selector: {app: checkout}, ports: [{port: 80, targetPort: 8080}]} ``` Change the Service's label to `app: check-out` and every object stays valid, `kubectl apply` succeeds, and the Service has nothing behind it.

  1. Pods and Workloads →

    What Kubernetes actually runs and which controller keeps it alive; every other section refers to pods and workload types.

  2. API Object Conventions →

    The shared grammar of labels, selectors, apply and ownership that ties objects together without direct references.

  3. Services and Networking →

    How traffic finds a changing set of pods, and where the first connection failures of most deployments come from.

  4. Config and Secrets →

    How settings and credentials reach a container, and why changing them does not always reach the running process.

  5. Cluster Architecture →

    The control plane and node agents that turn the loop, easier to follow once you know the objects they serve.

  6. Troubleshooting →

    Where every earlier section is tested together: reading a broken workload from its events down to the network.

  • Pointing a liveness probe at a database or downstream service: when that dependency fails, the kubelet restarts every healthy replica and turns a partial outage into a full one.

  • Debugging an unreachable Service inside the pod before checking its endpoints; an empty list means a selector or readiness problem — see Service Connectivity Debugging.

  • Treating a namespace as a security boundary: without RBAC, NetworkPolicy and Pod Security, pods in different namespaces share nodes, kernel and network freely.

  • Calling Secrets safe because their values are encoded; protection comes from RBAC scoping and encryption at rest, and neither is tight by default.

  • Answering a Pending pod with more nodes before reading its events, which name the filter — resources, taints, affinity or volume zone — that ruled every node out.

  • Expecting an Ingress, a LoadBalancer Service or a custom resource to work because kubectl apply succeeded; each needs a controller that may not be installed.

  • Blaming the application for latency spikes that a tight CPU limit causes through throttling, which never shows as a restart — see CPU Throttling and Latency.

  • Upgrading worker nodes ahead of the control plane: no kubelet may run a newer minor version than the API server, so the control plane goes first.

  • Assuming a StatefulSet makes a database highly available; it supplies stable names and volumes, while replication, failover and backups stay with you or an operator.

This guide assumes a current Kubernetes release, 1.33 or later. Interviewers still ask about the changes along the way, because older clusters and older tutorials remain common: - **1.19** made Ingress stable under `networking.k8s.io/v1`; **1.21** did the same for EndpointSlices and **1.22** for server-side apply. - **1.24** removed dockershim, so the kubelet talks to runtimes such as containerd or CRI-O through the CRI; images built with Docker still run unchanged. The same release stopped auto-creating long-lived token Secrets for ServiceAccounts, leaving short-lived projected tokens as the norm. - **1.25** removed PodSecurityPolicy; Pod Security Admission, which enforces the Pod Security Standards per namespace, replaced it. - **1.28** widened the kubelet skew allowance to three minor versions behind the API server. - **1.30** made ValidatingAdmissionPolicy stable, so many validation rules can run in-process as CEL expressions instead of through a webhook. - **1.33** made native sidecar containers stable: init containers that keep running beside the main ones. Gateway API is versioned separately from Kubernetes and installed as CRDs; it reached v1.0 in 2023. When an answer depends on any of these, such as how a node runs containers, how pod security is enforced or how a workload gets its token, say which era you are describing.

Kubernetes is rarely used alone, and interviewers expect you to place it. Helm and Kustomize package and customise manifests, and GitOps controllers such as Argo CD and Flux keep a cluster in line with a Git repository. Prometheus is the usual metrics stack, service meshes such as Istio and Linkerd add mutual TLS and traffic control between services, and containerd or CRI-O run the containers underneath. Managed offerings such as EKS, GKE and AKS run the control plane for you; they remove etcd and API server operations, not the workload, networking and security decisions. As a platform choice it competes with simpler orchestrators such as HashiCorp Nomad and Amazon ECS, and with serverless container services and PaaS products that hide the cluster entirely. The trade that separates them is control against operating cost: Kubernetes offers an extensible, portable API and a large ecosystem, and asks for a team that understands it. A defensible answer names the workload: many services, custom scheduling or operators favour a cluster, while a handful of stateless apps often do not. [Adoption and Migration](/topics/cloud-kubernetes-adoption) covers that judgement in depth.

explore

report an issue with this guide →

questions

448 · 13 sections

In Kubernetes, what does a DaemonSet guarantee that a Deployment with a fixed replica count does not, and what kind of workload is it the right controller for?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A DaemonSet keeps one pod on every node matching its scope. It has no replica count: a new node automatically gets a pod, a deleted node loses its pod. Use it for per-node agents such as log shippers, metrics collectors, CNI and CSI plugins.

open as a page

In Kubernetes, describe what a Deployment, a ReplicaSet and a Pod each own, and walk through what the cluster actually does when you change the container image in a Deployment's pod template.

level: juniorimportance: must knowfreq 80%
basics
~20 s

A Deployment manages ReplicaSets; a ReplicaSet manages Pods. Changing the pod template creates a new ReplicaSet, which is scaled up while the old one is scaled down. Old ReplicaSets stay at zero replicas so you can roll back.

open as a page

How do you watch, pause and roll back an in-flight Kubernetes Deployment rollout from the command line, and what limits how far back a rollback can go?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Use kubectl rollout status to watch, rollout pause/resume to hold it, rollout history to list revisions and rollout undo (optionally --to-revision) to go back. How far back you can go is bounded by revisionHistoryLimit, default 10.

open as a page

In a Kubernetes Pod spec, what does the `initContainers` list do, and what happens if one of those containers exits with a non-zero status?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Init containers run to completion one at a time, in list order, before any app container starts. A non-zero exit means the kubelet retries that init container, or fails the Pod if restartPolicy is Never. Later containers never start.

open as a page

What is a Kubernetes Job, how does it differ from a Deployment, and which values may the `restartPolicy` field of a Job's Pod template take?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A Job runs Pods until a set number of them terminate successfully, then stops. A Deployment keeps a set of Pods running forever, restarting them whenever they exit. A Job's Pod template must use restartPolicy Never or OnFailure — Always is rejected.

open as a page

In Kubernetes, what is the difference between a label and an annotation, and how do you decide which one to use?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Labels are short identifying key/value pairs that selectors match, so Services, Deployments and kubectl -l find objects by them. Annotations carry non-identifying data of any shape, capped at 256 KiB in total, that tools read by key but nothing can select on.

open as a page

In Kubernetes, how do `kubectl create -f`, `kubectl replace -f` and `kubectl apply -f` differ when you manage an object from a YAML file?

level: juniorimportance: must knowfreq 74%
basics
~20 s

kubectl create only makes new objects and fails if the name exists. kubectl replace overwrites an existing object with the whole file. kubectl apply creates or updates by merging the file into the live object and leaves fields it never set alone.

open as a page

In Kubernetes, what are labels, and how do equality-based and set-based label selectors decide which objects match?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Kubernetes labels are key/value pairs on an object's metadata. A label selector lists requirements, all of which must match: equality (=, !=) or set-based (in, notin, exists). Services, ReplicaSets and kubectl use selectors to find objects.

open as a page

In Kubernetes, what happens to a Deployment's ReplicaSets and Pods when you run kubectl delete deployment, and how does --cascade=orphan change that?

level: juniorimportance: must knowfreq 72%
basics
~20 s

kubectl delete deployment removes the Deployment, then the garbage collector deletes its ReplicaSets and their Pods, which point at their owners through ownerReferences. With --cascade=orphan, those references are removed instead, and the ReplicaSets and Pods keep running.

open as a page

How does client-side `kubectl apply` combine the last-applied configuration, the live Kubernetes object and your file to decide what to change?

level: middleimportance: must knowfreq 58%
basics
~20 s

Client-side kubectl apply runs a three-way merge. Fields in your file are set on the live object, fields in the last-applied record but gone from the file are deleted, and fields in neither are left alone.

open as a page

Kubernetes requires every pod to have its own IP address on a flat network. What exactly does that model require, and how do containers inside a single pod reach each other?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Every pod gets its own IP. Any pod can reach any other pod's IP directly, on any node, with no NAT. Containers inside one pod share a single network namespace, so they talk over localhost and share one port space.

open as a page

How does a workload running in a Kubernetes cluster resolve another Service by name? Describe what component answers the query and the DNS name forms available.

level: juniorimportance: must knowfreq 70%
basics
~20 s

CoreDNS runs in the cluster and its Service address is written into every pod's /etc/resolv.conf. A Service resolves as <service>.<namespace>.svc.cluster.local to its ClusterIP. Inside the same namespace the bare name works; across namespaces use <service>.<namespace>.

open as a page

What does a Kubernetes Ingress object actually do, and why does creating one have no effect until something else is installed in the cluster?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An Ingress is just declarative configuration: hostname and path rules mapping to Services, plus TLS. Kubernetes ships no code that acts on it. An ingress controller must be installed; it watches Ingress objects and programs a real proxy or load balancer to match.

open as a page

A Kubernetes Service of type ClusterIP gets an IP address that no network interface owns and that nothing answers ARP for. Explain what actually happens to a packet sent to that IP, and which component makes it work.

level: juniorimportance: must knowfreq 78%
basics
~20 s

A ClusterIP is a virtual IP with no machine behind it. kube-proxy runs on every node and programs kernel rules that rewrite the destination of packets aimed at that IP to one of the Service's backing pod IPs (DNAT). A backend is chosen once per new connection; connection tracking rewrites the replies back.

open as a page

Why does a Kubernetes LoadBalancer Service show EXTERNAL-IP <pending> in kubectl get svc, and which component is supposed to replace it?

level: juniorimportance: must knowfreq 76%
basics
~20 s

A type=LoadBalancer Service only records a request. The cloud-controller-manager's service controller, or another load-balancer implementation such as MetalLB, must create the balancer and write its address into status. <pending> means nothing has done that yet.

open as a page

In Kubernetes, what are the ways a Pod can consume a ConfigMap, and how does injecting values as environment variables differ from mounting the ConfigMap as a volume?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Three ways: envFrom to import all keys as env vars, valueFrom.configMapKeyRef to map one key to one env var, or a volume where each key becomes a file. Env values freeze at container start; mounted files are refreshed while the Pod runs.

open as a page

How can a container in Kubernetes learn its own Pod name, namespace, node name and Pod IP without calling the Kubernetes API server?

level: juniorimportance: must knowfreq 52%
basics
~20 s

Use the downward API: in the Pod spec, set env[].valueFrom.fieldRef with fieldPath metadata.name, metadata.namespace, spec.nodeName or status.podIP. The kubelet fills the values in when it creates the container — no API credentials or network call needed.

open as a page

A teammate says Kubernetes Secrets are safe because "the values are base64-encrypted". What is actually stored, and which protections do and do not apply?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Base64 is encoding, not encryption — anyone can decode it. A Secret's value is stored as-is in etcd, plaintext by default, and readable by anyone with API read access to it. Real protection comes from RBAC, encryption at rest, and limiting who and what can read it.

open as a page

You edit a Kubernetes ConfigMap that a running Deployment already consumes. Which of its consumers pick up the new values without a Pod restart, which never do, and roughly how long does propagation take?

level: middleimportance: must knowfreq 64%
basics
~20 s

Files from a configMap volume are refreshed by the kubelet, typically within about a minute. Environment variables never update, and a volume mounted with subPath never updates. Even a refreshed file only helps if the application re-reads it.

open as a page

How does the External Secrets Operator turn a value in an external secret manager into a Kubernetes Secret, and what does an ExternalSecret's refreshInterval control?

level: middleimportance: must knowfreq 68%
basics
~20 s

A SecretStore or ClusterSecretStore defines provider access. An ExternalSecret names remote keys and a target Secret. The operator fetches the values, writes the Secret, and re-reads the provider every refreshInterval, which defaults to one hour.

open as a page

A PersistentVolumeClaim in Kubernetes declares accessModes such as ReadWriteOnce, ReadOnlyMany, ReadWriteMany or ReadWriteOncePod. What does each of those mean, and what is the unit they are scoped to?

level: juniorimportance: must knowfreq 62%
basics
~10 s

ReadWriteOnce: mounted read-write by one node. ReadOnlyMany: mounted read-only by many nodes. ReadWriteMany: mounted read-write by many nodes. ReadWriteOncePod: exactly one pod cluster-wide. Except for ReadWriteOncePod, the scope is the node, not the pod.

open as a page

What is the Container Storage Interface (CSI) in Kubernetes, and why were storage drivers moved out of the Kubernetes core codebase onto it?

level: juniorimportance: must knowfreq 50%
basics
~20 s

CSI is a standard gRPC contract for storage drivers. A vendor ships the driver as ordinary pods in the cluster instead of code compiled into Kubernetes, so drivers install, upgrade and get fixed independently of the Kubernetes release.

open as a page

What does the `volumeClaimTemplates` field on a Kubernetes StatefulSet do, and how does it differ from listing a PersistentVolumeClaim under a Deployment's pod volumes?

level: juniorimportance: must knowfreq 58%
basics
~10 s

volumeClaimTemplates makes the StatefulSet controller create one PersistentVolumeClaim per replica, named <template>-<statefulset>-<ordinal>. A Deployment's pod spec references one existing PVC, so every replica shares the same volume.

open as a page

What is a Kubernetes StorageClass, and how does a PersistentVolumeClaim that references one end up with real storage mounted into a Pod?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A StorageClass is a named storage tier: a provisioner plus its parameters. When a PersistentVolumeClaim names that class, the provisioner creates real storage and a matching PersistentVolume, which is bound to the claim. No admin pre-creates the volume.

open as a page

In a Pod spec you can mount an emptyDir volume or mount a PersistentVolumeClaim. Explain the lifecycle difference between the two and when each is the right choice.

level: juniorimportance: must knowfreq 72%
basics
~20 s

An emptyDir is created when the Pod is placed on a node and deleted with the Pod, so it survives container restarts but not rescheduling. A PersistentVolumeClaim points at storage whose lifetime is independent of the Pod, so data survives deletion and rescheduling.

open as a page

What does node affinity do for a Kubernetes Pod, and how does it differ from the simpler nodeSelector field in the Pod spec?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Both steer a Pod onto nodes carrying particular labels. nodeSelector is exact key=value matching and every entry must match. Node affinity adds an expression language (In, NotIn, Exists, Gt, Lt), OR-ed alternatives, and a soft preferred form the scheduler may ignore.

open as a page

In Kubernetes, how does a pod request a GPU, and why can it request 0.35 CPU but not 0.35 GPU?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A pod requests a GPU as an extended resource, such as nvidia.com/gpu: 1, in its container limits. Extended resources are whole units the scheduler counts, not shares. So Kubernetes rejects fractions, requires the request to equal the limit, and never overcommits them.

open as a page

What does a Kubernetes HorizontalPodAutoscaler object do, what must be true of the workload for it to work, and what does it deliberately not do?

level: juniorimportance: must knowfreq 70%
basics
~20 s

It watches a metric (usually CPU) for a Deployment or StatefulSet and rewrites that workload's replica count, clamped between minReplicas and maxReplicas, to keep the metric near a target. Needs metrics-server plus resource requests on the pods. It adds and removes pods only — it never resizes a pod and never adds nodes.

open as a page

What is KEDA in a Kubernetes cluster, and what does a KEDA ScaledObject declare about the workload it scales?

level: juniorimportance: must knowfreq 62%
basics
~20 s

KEDA is a Kubernetes operator that scales workloads on external event signals such as Kafka lag, queue depth or a Prometheus query. A ScaledObject names the target workload, its replica bounds and its triggers. KEDA turns that into an HPA it owns.

open as a page

In a Kubernetes cluster with a spot node pool, which workloads belong on those nodes, and how do you keep everything else off them?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Spot nodes suit replicated, stateless or checkpointing work that can lose a node at short notice. Label the pool by capacity type and taint it, so only pods that tolerate the taint and select the label land there.

open as a page

In a Kubebuilder project built on controller-runtime, what does the framework provide, and what do you actually write inside Reconcile(ctx, req)?

level: juniorimportance: must knowfreq 58%
basics
~20 s

controller-runtime supplies a manager that runs shared caches, clients, work queues, leader election and metrics; Kubebuilder scaffolds the API types, markers and Makefile. You write Reconcile: fetch the object named by req, converge actual state to spec, then update status.

open as a page

What is a Kubernetes CustomResourceDefinition, and what exactly do you get once you apply one to a cluster?

level: juniorimportance: must knowfreq 65%
basics
~20 s

A CustomResourceDefinition registers a new resource type with the Kubernetes API server. You immediately get REST endpoints for it, storage in etcd, kubectl support, watches, RBAC and labels — but nothing acts on the objects until you run a controller for them.

open as a page

A team wants automated backups, version upgrades and failover for a stateful database running on Kubernetes, all driven through the Kubernetes API. Explain what the Kubernetes Operator pattern is and how it delivers that.

level: juniorimportance: must knowfreq 68%
basics
~20 s

An Operator is a custom Kubernetes resource type plus a controller running in the cluster. Users declare the desired database in a custom object; the controller continuously drives the real system to match that spec, automating install, backup, upgrade and failover.

open as a page

In Kubernetes, a request that creates an object can be intercepted by webhooks registered through both a MutatingWebhookConfiguration and a ValidatingWebhookConfiguration. In what order does the API server run them relative to each other and to the rest of its request handling, and why does that order matter?

level: middleimportance: must knowfreq 60%
basics
~20 s

The API server authenticates and authorizes, then runs mutating admission (webhooks may patch the object), then schema validation, then validating admission (accept or reject only), then writes to etcd. Mutation runs first so validators judge the final stored object.

open as a page

A CustomResourceDefinition in `apiextensions.k8s.io/v1` requires an OpenAPI v3 validation schema. What does that schema do to the objects users submit, what happens to fields the schema does not mention, and how would you deliberately allow arbitrary content in one part of the object?

level: middleimportance: must knowfreq 50%
basics
~20 s

The schema validates types, required fields and constraints at admission, applies declared defaults, and prunes — silently strips — any field it does not describe. To keep arbitrary content, mark that subtree with x-kubernetes-preserve-unknown-fields: true.

open as a page

In Kubernetes, when and how does a namespace's LimitRange apply its default and defaultRequest values, and what happens to Pods that already exist?

level: juniorimportance: must knowfreq 62%
basics
~10 s

The LimitRanger admission plugin in kube-apiserver copies defaultRequest and default into any container missing a request or limit at Pod creation, then rejects Pods outside min and max. Running Pods are never changed.

open as a page

What does a Kubernetes namespace actually scope, and which resources, such as Nodes, PersistentVolumes and CRDs, sit outside every namespace?

level: juniorimportance: must knowfreq 72%
basics
~10 s

A Kubernetes namespace scopes object names and most workload objects: Pods, Services, Secrets, ConfigMaps, PVCs. Nodes, PersistentVolumes, StorageClasses, ClusterRoles, CRDs and Namespaces themselves are cluster-scoped. Run kubectl api-resources --namespaced=false to list them.

open as a page

What is a Kubernetes ResourceQuota, and which kinds of consumption can its spec.hard field cap for one namespace?

level: juniorimportance: must knowfreq 64%
basics
~20 s

A ResourceQuota is a namespaced object whose spec.hard caps the namespace's total CPU and memory requests and limits, storage requests and object counts. The API server rejects any create that would push usage past a cap.

open as a page

A Kubernetes namespace has been stuck in Terminating for 19 minutes; how do you find what is blocking the namespace controller's purge, and fix it?

level: middleimportance: must knowfreq 72%
basics
~20 s

Read the namespace's status.conditions. They usually point to objects held by finalizers whose controller is gone, or to an unavailable aggregated API that makes discovery fail. Restore that controller or API, or remove the stale finalizer or APIService.

open as a page

In Kubernetes, what is the difference between soft and hard multi-tenancy, and what do tenants still share when each gets only a namespace?

level: middleimportance: must knowfreq 68%
basics
~20 s

Soft multi-tenancy gives each trusted team a namespace fenced by RBAC, quota and network policy. Hard multi-tenancy assumes hostile tenants and separates clusters or nodes, because namespaces still share the kernel, nodes, control plane and cluster-scoped objects.

open as a page

In Kubernetes, what does metrics-server provide for `kubectl top` and the HorizontalPodAutoscaler, and what is it deliberately not designed to be?

level: juniorimportance: must knowfreq 72%
basics
~20 s

metrics-server scrapes current CPU and memory usage from every kubelet, keeps only the latest points in memory, and serves them as the metrics.k8s.io API. It is not a monitoring system: no history, no custom metrics, no alerting.

open as a page

In a Kubernetes control plane, what is kube-apiserver responsible for, and why is it the only component that talks to etcd directly?

level: juniorimportance: must knowfreq 75%
basics
~20 s

kube-apiserver is the cluster's front door: a REST API that authenticates, authorises, validates and persists every object, and streams changes to watchers. Routing all writes through it gives one place for auth, validation, admission and audit.

open as a page

In a Kubernetes cluster, what does the kube-scheduler actually do when a new Pod is created, and what does it not do?

level: juniorimportance: must knowfreq 78%
basics
~20 s

kube-scheduler watches for Pods with an empty spec.nodeName, picks a suitable node, and writes that choice back to the API server (a binding). It does not start the container — the kubelet on the chosen node does that.

open as a page

In a Kubernetes control plane, what is etcd, what exactly is stored in it, and which components are allowed to talk to it?

level: juniorimportance: must knowfreq 70%
basics
~20 s

etcd is the cluster's only persistent database: a distributed, strongly consistent key-value store holding every API object (Pods, Deployments, Services, ConfigMaps, Secrets, RBAC, node state). Only kube-apiserver connects to it; everything else reads and writes through the API server.

open as a page

If every kube-apiserver in a Kubernetes cluster becomes unreachable, what keeps running on the nodes and what stops working?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Running Pods keep running: the kubelet keeps their containers alive and kube-proxy's installed Service rules keep routing. Anything that needs a write or fresh state stops working: scheduling, scaling, rollouts, rescheduling off failed nodes and kubectl.

open as a page

Walk through what happens to a request such as 'kubectl apply -f pod.yaml' inside the Kubernetes API server, from the TLS connection to the object being stored in etcd.

level: juniorimportance: must knowfreq 62%
basics
~20 s

The API server terminates TLS, then runs three gates in order: authentication (who are you — cert, service-account token, or OIDC token), authorization (may this identity do this verb on this resource — usually RBAC), and admission (mutating plugins/webhooks may edit the object, then validating ones may reject it). Only then is the object schema-validated and written to etcd.

open as a page

In a Kubernetes Pod spec, what do the securityContext fields runAsNonRoot and runAsUser do, and why does it matter whether a container process runs as UID 0?

level: juniorimportance: must knowfreq 62%
basics
~20 s

runAsUser sets the numeric UID the container process runs as. runAsNonRoot: true makes the kubelet refuse to start the container if it would run as UID 0. Root inside the container is real root on the host kernel, so any container escape starts privileged.

open as a page

In Kubernetes RBAC, what are the three parts of a policy rule — apiGroups, resources and verbs — and how would you write a Role that grants read-only access to Pods in a single namespace?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A rule says which API groups, which resource types, and which verbs are allowed. Read-only pods is apiGroups: [""] (core group), resources: ["pods"], verbs: ["get","list","watch"]. A Role holds the rules; a RoleBinding attaches it to a user, group or ServiceAccount. RBAC is purely additive — there is no deny.

open as a page

A teammate argues that Kubernetes Secret objects are safe to share and commit to git because their values are base64-encoded. What is wrong with that reasoning, and what actually protects Secret data in a cluster?

level: juniorimportance: must knowfreq 65%
basics
~20 s

Base64 is an encoding, not encryption: anyone decodes it in one command, no key needed. By default a Secret sits unencrypted in etcd and is readable by anyone whose RBAC allows reading Secrets. Real protection is least-privilege RBAC, encryption at rest, and keeping plaintext manifests out of git.

open as a page

When a process running inside a Pod calls the Kubernetes API, what identity does it present and where does that credential come from? Why do security reviews object to leaving every workload on the namespace's default ServiceAccount?

level: juniorimportance: must knowfreq 60%
basics
~20 s

It presents a ServiceAccount: a namespaced identity named by spec.serviceAccountName. The kubelet projects a signed JWT for it into the container at /var/run/secrets/kubernetes.io/serviceaccount/token. Leaving everything on default means all workloads share one identity, so any permission granted to it is granted to all of them and nothing is attributable.

open as a page

When a Kubernetes cluster's API server becomes unreachable, what happens to pods that are already running, and what stops working?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Running pods keep running and existing Service traffic keeps flowing, because kubelets and kube-proxy work from state they already have. Anything that needs the API server stops: scheduling, rescheduling, scaling, rollouts, endpoint updates and CronJob runs.

open as a page

How do you run a command or open an interactive shell inside a container of a running Kubernetes pod with kubectl, and what are the limits of that approach?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Use kubectl exec -it POD -c CONTAINER -- sh. It starts an extra process inside an already-running container, streamed through the API server and the kubelet. The binary must exist in the image and the container must be running.

open as a page

What does a Kubernetes container's imagePullPolicy control, what is its default, and how does it interact with images already cached on the node?

level: juniorimportance: must knowfreq 68%
basics
~20 s

imagePullPolicy tells the kubelet when to check the registry: Always on every start, IfNotPresent only when the image is not cached, Never not at all. Unset, it defaults to Always for :latest or images with no tag and no digest, otherwise IfNotPresent.

open as a page

How do you read container logs with kubectl, including output from a container instance that has already exited, and why might kubectl logs return nothing at all?

level: juniorimportance: must knowfreq 70%
basics
~20 s

kubectl logs POD -c CONTAINER shows the current instance; --previous shows the last terminated one. Add -f to follow, --since and --tail to bound output. Nothing appears if the app logs to a file instead of stdout/stderr, the pod never started, or the logs rotated.

open as a page

In Kubernetes, after a pod is deleted and its Deployment replaces it, can you still read the old pod's container logs with kubectl, and why?

level: juniorimportance: must knowfreq 72%
basics
~20 s

No. Kubernetes keeps container logs only as files on the pod's node, tied to that pod; once the pod is deleted the kubelet removes its containers and log directory, and the replacement pod starts with empty logs.

open as a page

In a managed Kubernetes service, which parts of the cluster does the provider operate, and what stays the cluster owner's responsibility?

level: juniorimportance: must knowfreq 70%
basics
~10 s

The provider runs the control plane: kube-apiserver, etcd, the scheduler and controller managers, including their hosts, patching and availability. You still own worker nodes, workloads, RBAC, add-ons, network choices and deciding when to upgrade.

open as a page

In Kubernetes, what do `kubectl cordon`, `kubectl drain` and `kubectl uncordon` each do when you take a node out of service?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Cordon marks a node unschedulable so no new pods land there. Drain cordons it and then evicts its pods so their controllers recreate them elsewhere. Uncordon makes the node schedulable again without moving any pods back.

open as a page

When a Kubernetes cluster is lost, what does an etcd snapshot bring back, and what must come from volume backups or Git instead?

level: middleimportance: must knowfreq 62%
basics
~20 s

An etcd snapshot restores Kubernetes API objects (Deployments, Secrets, RBAC, PVC and PV objects, custom resources) but none of the bytes stored on persistent volumes. Volume data needs CSI snapshots or Velero; Git re-creates only what was committed.

open as a page

When a node runs kubeadm join with a bootstrap token and --discovery-token-ca-cert-hash, what does each value prove, and to whom?

level: middleimportance: must knowfreq 66%
basics
~20 s

The bootstrap token is a short-lived shared secret: the node checks a token-signed cluster-info and the API server accepts the new kubelet. The CA cert hash pins the cluster CA public key, so a leaked token cannot impersonate the cluster.

open as a page

Why do organizations running Kubernetes split workloads across many clusters instead of one large cluster, and what does each driver buy them?

level: middleimportance: must knowfreq 62%
basics
~10 s

A Kubernetes cluster is one failure and control domain, so teams run many clusters to limit blast radius, place workloads in specific regions, draw compliance boundaries and give tenants isolation that namespaces cannot provide.

open as a page

Before moving a ticket-booking checkout service onto Kubernetes, what must the application itself change to run well as a container?

level: juniorimportance: must knowfreq 70%
basics
~20 s

The app must read configuration from environment variables or mounted files, log to stdout and stderr, expose a readiness signal that reflects real ability to serve, shut down cleanly on SIGTERM, and size itself from its container limits.

open as a page

On a Kubernetes platform run by a platform team, what is a golden-path template, and why use it instead of hand-written manifests?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A golden-path template is a platform-maintained, pre-approved way to deploy a service: developers fill in a few values and the template produces the Deployment, Service, probes and resources, so every team gets safe defaults without writing raw Kubernetes manifests.

open as a page

When PostgreSQL runs in a Kubernetes StatefulSet with per-replica PersistentVolumeClaims, what does Kubernetes actually provide, and what is still left for you to build?

level: middleimportance: must knowfreq 64%
basics
~20 s

A StatefulSet gives each PostgreSQL pod a stable name, its own volume and ordered rollout. It knows nothing about which pod is primary, replication, failover, backups or safe upgrades, so all of that is still yours.

open as a page

When moving a service from VMs into Kubernetes pods, which host-level assumptions usually break, and what replaces each one?

level: middleimportance: must knowfreq 64%
basics
~20 s

Pods are disposable and restart anywhere with a new IP and an empty filesystem, so local disk, fixed IPs, in-memory sessions, host cron and log files break. They move to object storage or a PVC, a Service name, a shared session store, a CronJob and stdout.

open as a page

A team wants to run PostgreSQL and Kafka inside its Kubernetes cluster rather than use managed services. How do you decide, and what would make you say no?

level: principalimportance: must knowfreq 52%
basics
~20 s

Decide by ownership and risk, not by preference. In-cluster operators suit teams that can run databases, need portability or have no managed option. Otherwise a managed service is usually cheaper overall. Say no when nobody can own recovery.

open as a page