Kubernetes
Declaring workloads and letting controllers reconcile the cluster toward that spec, exposing them through Services and Ingress, and diagnosing it when the cluster disagrees with you. It is the default platform question in devops, SRE and backend loops.
on this pageshowhide
guide
overview
~2 minKubernetes runs containers by comparing what you declared with what exists and letting controllers close the gap. Interviewers treat it as the default platform subject because most backend, SRE and DevOps roles touch it, and because the distance between "I wrote a Deployment" and "I know why this pod is Pending" is easy to hear. A strong answer names the object involved, the controller that acts on it, and the signal you would read when the two disagree. The hub follows a workload from declaration to traffic. [Pods and workloads](/topics/cloud-kubernetes-pods-workloads) covers what gets scheduled and which controller keeps it alive, and [API object conventions](/topics/cloud-kubernetes-api-objects) is the shared grammar every other object uses: labels, metadata, apply and ownership. [Services and networking](/topics/cloud-kubernetes-services-networking), [config and secrets](/topics/cloud-kubernetes-config-secrets) and [storage and volumes](/topics/cloud-kubernetes-storage-volumes) are what a running pod needs from outside itself: traffic, settings and data that outlive it. [Scaling and scheduling](/topics/cloud-kubernetes-scaling-scheduling) decides where pods land and how many there are. Underneath sits [cluster architecture](/topics/cloud-kubernetes-architecture), the control plane and node agents that make the loop turn, and [operators and CRDs](/topics/cloud-kubernetes-operators-crds) extend that loop to your own types. [Security](/topics/cloud-kubernetes-security) and [tenancy and quotas](/topics/cloud-kubernetes-namespaces-tenancy) decide who may do what, and how much. The operator's side is [troubleshooting](/topics/cloud-kubernetes-troubleshooting) and [cluster lifecycle](/topics/cloud-kubernetes-cluster-lifecycle), and [adoption and migration](/topics/cloud-kubernetes-adoption) asks whether a workload belongs on a cluster at all. Junior rounds check the object model: what a pod is, how a Deployment differs from a StatefulSet, what each Service type exposes, what a failing probe triggers. Middle rounds move to behaviour under change, such as rollouts, autoscaling, config updates and resource pressure. Senior and staff rounds are usually one incident or one design: a Service that answers nothing, a node that keeps evicting, an upgrade plan, a tenancy model. Learn pods, controllers and labels first; nearly every later answer is a consequence of them. Networking and configuration come next, because a first deployment usually breaks there. Architecture makes sense once you have seen the objects it serves, and troubleshooting is where all of it is tested together.
primer
### You declare, controllers converge Most objects have a spec you write and a status the cluster reports. Controllers are loops that watch objects, compare the two, act to narrow the difference, and try again later. Very little in the system is a one-shot command: a deleted pod comes back because a controller still wants it. A "why did Kubernetes do that" question is usually answered by naming which controller acted and what spec it was chasing. ### The API server is the only door Controllers, the scheduler, kubelets and `kubectl` do not talk to each other directly. They read and write objects through kube-apiserver, which alone persists them in etcd, and they learn about changes by watching. That single path is why authentication, authorisation, admission and audit can live in one place, and why losing the API server freezes change without stopping the containers already on the nodes. ### Pods are disposable; controllers give them a job A pod is never moved, only replaced, and a replacement gets a new IP and a fresh filesystem. The workload controller around it carries the intent: interchangeable replicas, stable identities with their own storage, one copy per node, or work that runs to completion. Choosing the wrong one rarely fails on day one; it fails during a rollout, a node loss or a scale-down. ### Labels are the wiring Objects seldom hold direct references to each other. Services, ReplicaSets, NetworkPolicies, PodDisruptionBudgets and affinity rules find their targets by label selector, so a label typo silently disconnects things instead of raising an error. Ownership runs the other way, through owner references that decide what is cleaned up when a parent is deleted. ### Requests schedule, limits enforce The scheduler places pods by their declared requests against a node's allocatable capacity, not by what they actually use. At runtime the node enforces limits, and the two resources fail differently: a container over its memory limit is killed, one over its CPU limit is slowed. The same numbers shape which pods go first when a node comes under pressure, so sizing is a reliability decision as much as a cost one. ### Declaring is not doing Several objects are only records until something acts on them. An Ingress needs an ingress controller, a `LoadBalancer` Service needs a cloud or bare-metal implementation, a custom resource needs its operator, and a PersistentVolumeClaim needs a provisioner or a matching volume. Many "it does nothing" incidents are a missing or unhealthy controller, not a wrong manifest. ### Networking is virtual, and readiness steers it A Service is a stable name and virtual address in front of a changing set of pods, and only ready pods are placed behind it, so probes shape traffic as much as they report health. DNS, the Service proxy layer and the CNI plugin each own one layer; debugging a connection means checking them in order instead of guessing. ### Namespaces organise; they do not isolate A namespace scopes names, access bindings, quotas and policies. It does not give a tenant its own nodes, kernel or control plane, and many resources are cluster-scoped. Real separation is layered: RBAC for the API, NetworkPolicy for traffic, Pod Security for what a container may do on its node, and sometimes separate node pools or clusters.
- Pod
- The smallest unit Kubernetes schedules: one or more containers placed together on a node, sharing a network identity and volumes. It is replaced, never moved.
- Controller
- A control loop that watches objects through the API server and acts to move actual state toward the declared spec, retrying until the two agree.
- Deployment
- A workload controller for interchangeable replicas. It manages one ReplicaSet per pod template version, so changes roll out gradually and can be rolled back.
- StatefulSet
- A workload controller that gives each replica a stable ordinal name, its own persistent volume claim, and ordered start, update and shutdown.
- Label selector
- A query over labels that one object uses to find others. Services, ReplicaSets, NetworkPolicies and disruption budgets all pick their targets this way.
- Service
- A stable virtual address and DNS name in front of the ready pods its selector matches, with types that expose it inside or outside the cluster.
- EndpointSlice
- The object listing a Service's current backend addresses and their readiness. Proxies read it, and it changes as pods start, stop or fail probes.
- Ingress
- An object describing HTTP host and path routing to Services, acted on by a separately installed ingress controller. Gateway API is its role-oriented successor.
- kube-apiserver
- The control plane's REST front end. It authenticates, authorises, admits and validates every request, and is the only component that reads and writes etcd.
- etcd
- The consistent, replicated key-value store holding all cluster state. Losing it without a backup means losing the cluster's record of everything declared.
- kubelet
- The agent on each node that runs the pods assigned to it through the container runtime, executes their probes, and reports node and pod status.
- Requests and limits
- Per-container CPU and memory figures: requests are what scheduling reserves, limits are the runtime ceiling. Together they also set the pod's QoS class.
- PersistentVolumeClaim
- A pod-side request for storage of a given size, access mode and class, bound to a PersistentVolume whose lifetime is independent of any pod.
- CustomResourceDefinition
- An object that registers a new resource type with the API server, giving it storage, endpoints and RBAC; behaviour comes from a controller you supply.
- Operator
- A controller plus custom resources that encode how to run one specific application, including day-2 work such as upgrades, backups and failover.
- Admission control
- The stage after authorisation where built-in plugins, webhooks or policies may modify or reject an object before it is stored.
- RBAC
- Role-based access control: Roles and ClusterRoles list allowed verbs on resources, and bindings grant them to users, groups or ServiceAccounts. Rules only ever add permissions.
- ServiceAccount
- The identity a pod uses when it calls the API server, usually presented as a short-lived token mounted into the pod and authorised through RBAC.
- PodDisruptionBudget
- A cap on how many replicas of an application voluntary disruptions, such as node drains, may take down at once. Hardware or node failures ignore it.
- Namespace
- A scope for names, access bindings, quotas and policies inside one cluster. It groups resources but does not isolate their workloads at runtime.
### One apply, followed through Follow one `kubectl apply` of a Deployment. The request reaches kube-apiserver, which authenticates the caller, checks RBAC, runs admission and stores the object in etcd. The Deployment controller sees it through a watch and creates a ReplicaSet for the current pod template; the ReplicaSet controller creates pods to match the replica count. The scheduler notices pods with no node, picks one using requests, affinity, taints and spread rules, and records the binding. The kubelet on that node sees the assignment, asks the container runtime to pull images and start containers, has the CNI plugin give the pod an address, mounts its volumes, config and secrets, and starts the probes. Traffic arrives through the same machinery. Once a pod is ready it joins the Service's EndpointSlices; kube-proxy on every node, or a CNI datapath that replaces it, steers the Service address to those backends, and CoreDNS gives the Service a name. An Ingress or Gateway controller watches its own objects and configures the proxy that brings outside traffic in. ### Where the other sections attach - **Autoscalers** edit fields other controllers already act on: the HPA changes a workload's replica count, and node autoscalers add capacity when pods stay unschedulable. - **Operators** repeat the Deployment controller's pattern for your own types, registered through CRDs and often guarded by admission webhooks. - **Troubleshooting** walks the chain backwards: object, events, scheduling, kubelet, container, endpoints, network. - **Cluster lifecycle** keeps the chain itself healthy: certificates, etcd backups, and upgrades in the order the skew rules allow. ### The wiring in one manifest The objects meet through labels rather than references. Below, the Deployment's selector must match its own template, the Service finds the pods by the same label, readiness decides when each one receives traffic, and requests place the pod while the memory limit caps it on the node: ```yaml kind: Deployment # apps/v1 metadata: {name: checkout} spec: replicas: 3 selector: {matchLabels: {app: checkout}} template: metadata: {labels: {app: checkout}} spec: containers: - name: app image: registry.example.com/checkout:1.4.2 resources: {requests: {cpu: 250m, memory: 256Mi}, limits: {memory: 256Mi}} readinessProbe: {httpGet: {path: /ready, port: 8080}} --- kind: Service # v1 metadata: {name: checkout} spec: {selector: {app: checkout}, ports: [{port: 80, targetPort: 8080}]} ``` Change the Service's label to `app: check-out` and every object stays valid, `kubectl apply` succeeds, and the Service has nothing behind it.
- Pods and Workloads →
What Kubernetes actually runs and which controller keeps it alive; every other section refers to pods and workload types.
- API Object Conventions →
The shared grammar of labels, selectors, apply and ownership that ties objects together without direct references.
- Services and Networking →
How traffic finds a changing set of pods, and where the first connection failures of most deployments come from.
- Config and Secrets →
How settings and credentials reach a container, and why changing them does not always reach the running process.
- Cluster Architecture →
The control plane and node agents that turn the loop, easier to follow once you know the objects they serve.
- Troubleshooting →
Where every earlier section is tested together: reading a broken workload from its events down to the network.
Pointing a liveness probe at a database or downstream service: when that dependency fails, the kubelet restarts every healthy replica and turns a partial outage into a full one.
Debugging an unreachable Service inside the pod before checking its endpoints; an empty list means a selector or readiness problem — see Service Connectivity Debugging.
Treating a namespace as a security boundary: without RBAC, NetworkPolicy and Pod Security, pods in different namespaces share nodes, kernel and network freely.
Calling Secrets safe because their values are encoded; protection comes from RBAC scoping and encryption at rest, and neither is tight by default.
Answering a Pending pod with more nodes before reading its events, which name the filter — resources, taints, affinity or volume zone — that ruled every node out.
Expecting an Ingress, a
LoadBalancerService or a custom resource to work becausekubectl applysucceeded; each needs a controller that may not be installed.Blaming the application for latency spikes that a tight CPU limit causes through throttling, which never shows as a restart — see CPU Throttling and Latency.
Upgrading worker nodes ahead of the control plane: no kubelet may run a newer minor version than the API server, so the control plane goes first.
Assuming a StatefulSet makes a database highly available; it supplies stable names and volumes, while replication, failover and backups stay with you or an operator.
This guide assumes a current Kubernetes release, 1.33 or later. Interviewers still ask about the changes along the way, because older clusters and older tutorials remain common: - **1.19** made Ingress stable under `networking.k8s.io/v1`; **1.21** did the same for EndpointSlices and **1.22** for server-side apply. - **1.24** removed dockershim, so the kubelet talks to runtimes such as containerd or CRI-O through the CRI; images built with Docker still run unchanged. The same release stopped auto-creating long-lived token Secrets for ServiceAccounts, leaving short-lived projected tokens as the norm. - **1.25** removed PodSecurityPolicy; Pod Security Admission, which enforces the Pod Security Standards per namespace, replaced it. - **1.28** widened the kubelet skew allowance to three minor versions behind the API server. - **1.30** made ValidatingAdmissionPolicy stable, so many validation rules can run in-process as CEL expressions instead of through a webhook. - **1.33** made native sidecar containers stable: init containers that keep running beside the main ones. Gateway API is versioned separately from Kubernetes and installed as CRDs; it reached v1.0 in 2023. When an answer depends on any of these, such as how a node runs containers, how pod security is enforced or how a workload gets its token, say which era you are describing.
Kubernetes is rarely used alone, and interviewers expect you to place it. Helm and Kustomize package and customise manifests, and GitOps controllers such as Argo CD and Flux keep a cluster in line with a Git repository. Prometheus is the usual metrics stack, service meshes such as Istio and Linkerd add mutual TLS and traffic control between services, and containerd or CRI-O run the containers underneath. Managed offerings such as EKS, GKE and AKS run the control plane for you; they remove etcd and API server operations, not the workload, networking and security decisions. As a platform choice it competes with simpler orchestrators such as HashiCorp Nomad and Amazon ECS, and with serverless container services and PaaS products that hide the cluster entirely. The trade that separates them is control against operating cost: Kubernetes offers an extensible, portable API and a large ecosystem, and asks for a team that understands it. A defensible answer names the workload: many services, custom scheduling or operators favour a cluster, while a handful of stateless apps often do not. [Adoption and Migration](/topics/cloud-kubernetes-adoption) covers that judgement in depth.
explore
- Pods and Workloads49 questions
- Pod Model and Lifecycle5 questions
- Init and Sidecar Containers5 questions
- Deployments and ReplicaSets4 questions
- StatefulSets4 questions
- DaemonSets4 questions
- Jobs and CronJobs6 questions
- Liveness, Readiness, and Startup Probes5 questions
- Requests, Limits, and QoS Classes5 questions
- Graceful Shutdown3 questions
- Rollouts and Rollback4 questions
- Canary and Blue-Green4 questions
- API Object Conventions13 questions
- Labels & Selectors3 questions
- Annotations & Metadata3 questions
- Apply & Field Management4 questions
- Ownership & Garbage Collection3 questions
- Services and Networking51 questions
- Service Types6 questions
- kube-proxy and EndpointSlices5 questions
- Cluster DNS and CoreDNS6 questions
- Ingress and Ingress Controllers6 questions
- Gateway API Overview4 questions
- NetworkPolicies5 questions
- Pod Network Model and CNI5 questions
- Traffic Policy & Source IP4 questions
- LoadBalancer Provisioning4 questions
- Headless Services & Pod DNS3 questions
- Endpoints & Connection Draining3 questions
- Config and Secrets24 questions
- ConfigMaps5 questions
- Secrets6 questions
- Downward API4 questions
- Immutable ConfigMaps and Secrets4 questions
- External Secret Integration5 questions
- Storage and Volumes29 questions
- Ephemeral Volume Types4 questions
- Volumes vs PV/PVC4 questions
- StorageClasses and Dynamic Provisioning4 questions
- Access Modes and Volume Modes5 questions
- StatefulSet volumeClaimTemplates4 questions
- Expansion & Snapshots4 questions
- CSI Concept4 questions
- Scaling and Scheduling59 questions
- Scheduler Filter and Score Phases5 questions
- Node and Pod Affinity4 questions
- Taints and Tolerations5 questions
- Topology Spread Constraints5 questions
- Horizontal Pod Autoscaler5 questions
- VPA and Cluster Autoscaler Overview4 questions
- PodDisruptionBudgets4 questions
- Priority and Preemption5 questions
- Event-Driven Autoscaling4 questions
- Allocatable & Bin Packing4 questions
- Node Autoscaling5 questions
- Spot & Interrupted Capacity5 questions
- GPU & Device Plugins4 questions
- Operators and CRDs31 questions
- CustomResourceDefinitions6 questions
- Custom Controllers and Reconcile Loops5 questions
- The Operator Pattern4 questions
- Admission Webhooks5 questions
- CRD Versions & Conversion4 questions
- Status & Scale Subresources3 questions
- controller-runtime & Kubebuilder4 questions
- Tenancy & Quotas18 questions
- Namespaces & Scoping4 questions
- ResourceQuota4 questions
- LimitRange3 questions
- Multi-Tenancy Models4 questions
- Namespace Deletion3 questions
- Cluster Architecture43 questions
- kube-apiserver6 questions
- etcd4 questions
- Controller Manager and Scheduler Roles4 questions
- Node Components5 questions
- Desired-State Reconciliation4 questions
- Watch and Informer Mechanics5 questions
- API Groups & Versioning3 questions
- HA Control Planes5 questions
- API Priority & Fairness4 questions
- Aggregated APIs & Metrics3 questions
- Security41 questions
- AuthN, AuthZ, and Admission Flow5 questions
- RBAC4 questions
- Service Accounts and Workload Identity4 questions
- securityContext and Pod Security Standards4 questions
- Policy Enforcement via Admission5 questions
- Secrets Access Hygiene5 questions
- User Access & Kubeconfig5 questions
- API Audit Logging4 questions
- Cluster Hardening Baseline5 questions
- Troubleshooting46 questions
- kubectl Triage Workflow5 questions
- CrashLoopBackOff and ImagePullBackOff6 questions
- Pending and Unschedulable Pods5 questions
- Node Pressure and Evictions5 questions
- Debugging Running Containers6 questions
- Service Connectivity Debugging5 questions
- Image Pull Failures3 questions
- Logs and Events3 questions
- CPU Throttling and Latency4 questions
- Control-Plane Failures4 questions
- Cluster Lifecycle28 questions
- Cluster Bootstrap4 questions
- Upgrades & Version Skew5 questions
- Node Maintenance & Draining3 questions
- Cluster PKI & Certificates4 questions
- etcd Backup & Recovery4 questions
- Managed Control Planes4 questions
- Multi-Cluster Topology4 questions
- Adoption & Migration16 questions
- Migrating VM Workloads4 questions
- Databases & Data Services5 questions
- Application Readiness3 questions
- Developer Self-Service4 questions
questions
448 · 13 sectionsIn Kubernetes, what does a DaemonSet guarantee that a Deployment with a fixed replica count does not, and what kind of workload is it the right controller for?
basics
~20 sA DaemonSet keeps one pod on every node matching its scope. It has no replica count: a new node automatically gets a pod, a deleted node loses its pod. Use it for per-node agents such as log shippers, metrics collectors, CNI and CSI plugins.
In Kubernetes, describe what a Deployment, a ReplicaSet and a Pod each own, and walk through what the cluster actually does when you change the container image in a Deployment's pod template.
basics
~20 sA Deployment manages ReplicaSets; a ReplicaSet manages Pods. Changing the pod template creates a new ReplicaSet, which is scaled up while the old one is scaled down. Old ReplicaSets stay at zero replicas so you can roll back.
How do you watch, pause and roll back an in-flight Kubernetes Deployment rollout from the command line, and what limits how far back a rollback can go?
basics
~20 sUse kubectl rollout status to watch, rollout pause/resume to hold it, rollout history to list revisions and rollout undo (optionally --to-revision) to go back. How far back you can go is bounded by revisionHistoryLimit, default 10.
In a Kubernetes Pod spec, what does the `initContainers` list do, and what happens if one of those containers exits with a non-zero status?
basics
~20 sInit containers run to completion one at a time, in list order, before any app container starts. A non-zero exit means the kubelet retries that init container, or fails the Pod if restartPolicy is Never. Later containers never start.
What is a Kubernetes Job, how does it differ from a Deployment, and which values may the `restartPolicy` field of a Job's Pod template take?
basics
~20 sA Job runs Pods until a set number of them terminate successfully, then stops. A Deployment keeps a set of Pods running forever, restarting them whenever they exit. A Job's Pod template must use restartPolicy Never or OnFailure — Always is rejected.
In Kubernetes, what is the difference between a label and an annotation, and how do you decide which one to use?
basics
~20 sLabels are short identifying key/value pairs that selectors match, so Services, Deployments and kubectl -l find objects by them. Annotations carry non-identifying data of any shape, capped at 256 KiB in total, that tools read by key but nothing can select on.
In Kubernetes, how do `kubectl create -f`, `kubectl replace -f` and `kubectl apply -f` differ when you manage an object from a YAML file?
basics
~20 skubectl create only makes new objects and fails if the name exists. kubectl replace overwrites an existing object with the whole file. kubectl apply creates or updates by merging the file into the live object and leaves fields it never set alone.
In Kubernetes, what are labels, and how do equality-based and set-based label selectors decide which objects match?
basics
~20 sKubernetes labels are key/value pairs on an object's metadata. A label selector lists requirements, all of which must match: equality (=, !=) or set-based (in, notin, exists). Services, ReplicaSets and kubectl use selectors to find objects.
In Kubernetes, what happens to a Deployment's ReplicaSets and Pods when you run kubectl delete deployment, and how does --cascade=orphan change that?
basics
~20 skubectl delete deployment removes the Deployment, then the garbage collector deletes its ReplicaSets and their Pods, which point at their owners through ownerReferences. With --cascade=orphan, those references are removed instead, and the ReplicaSets and Pods keep running.
How does client-side `kubectl apply` combine the last-applied configuration, the live Kubernetes object and your file to decide what to change?
basics
~20 sClient-side kubectl apply runs a three-way merge. Fields in your file are set on the live object, fields in the last-applied record but gone from the file are deleted, and fields in neither are left alone.
Kubernetes requires every pod to have its own IP address on a flat network. What exactly does that model require, and how do containers inside a single pod reach each other?
basics
~20 sEvery pod gets its own IP. Any pod can reach any other pod's IP directly, on any node, with no NAT. Containers inside one pod share a single network namespace, so they talk over localhost and share one port space.
How does a workload running in a Kubernetes cluster resolve another Service by name? Describe what component answers the query and the DNS name forms available.
basics
~20 sCoreDNS runs in the cluster and its Service address is written into every pod's /etc/resolv.conf. A Service resolves as <service>.<namespace>.svc.cluster.local to its ClusterIP. Inside the same namespace the bare name works; across namespaces use <service>.<namespace>.
What does a Kubernetes Ingress object actually do, and why does creating one have no effect until something else is installed in the cluster?
basics
~20 sAn Ingress is just declarative configuration: hostname and path rules mapping to Services, plus TLS. Kubernetes ships no code that acts on it. An ingress controller must be installed; it watches Ingress objects and programs a real proxy or load balancer to match.
A Kubernetes Service of type ClusterIP gets an IP address that no network interface owns and that nothing answers ARP for. Explain what actually happens to a packet sent to that IP, and which component makes it work.
basics
~20 sA ClusterIP is a virtual IP with no machine behind it. kube-proxy runs on every node and programs kernel rules that rewrite the destination of packets aimed at that IP to one of the Service's backing pod IPs (DNAT). A backend is chosen once per new connection; connection tracking rewrites the replies back.
Why does a Kubernetes LoadBalancer Service show EXTERNAL-IP <pending> in kubectl get svc, and which component is supposed to replace it?
basics
~20 sA type=LoadBalancer Service only records a request. The cloud-controller-manager's service controller, or another load-balancer implementation such as MetalLB, must create the balancer and write its address into status. <pending> means nothing has done that yet.
In Kubernetes, what are the ways a Pod can consume a ConfigMap, and how does injecting values as environment variables differ from mounting the ConfigMap as a volume?
basics
~20 sThree ways: envFrom to import all keys as env vars, valueFrom.configMapKeyRef to map one key to one env var, or a volume where each key becomes a file. Env values freeze at container start; mounted files are refreshed while the Pod runs.
How can a container in Kubernetes learn its own Pod name, namespace, node name and Pod IP without calling the Kubernetes API server?
basics
~20 sUse the downward API: in the Pod spec, set env[].valueFrom.fieldRef with fieldPath metadata.name, metadata.namespace, spec.nodeName or status.podIP. The kubelet fills the values in when it creates the container — no API credentials or network call needed.
A teammate says Kubernetes Secrets are safe because "the values are base64-encrypted". What is actually stored, and which protections do and do not apply?
basics
~20 sBase64 is encoding, not encryption — anyone can decode it. A Secret's value is stored as-is in etcd, plaintext by default, and readable by anyone with API read access to it. Real protection comes from RBAC, encryption at rest, and limiting who and what can read it.
You edit a Kubernetes ConfigMap that a running Deployment already consumes. Which of its consumers pick up the new values without a Pod restart, which never do, and roughly how long does propagation take?
basics
~20 sFiles from a configMap volume are refreshed by the kubelet, typically within about a minute. Environment variables never update, and a volume mounted with subPath never updates. Even a refreshed file only helps if the application re-reads it.
How does the External Secrets Operator turn a value in an external secret manager into a Kubernetes Secret, and what does an ExternalSecret's refreshInterval control?
basics
~20 sA SecretStore or ClusterSecretStore defines provider access. An ExternalSecret names remote keys and a target Secret. The operator fetches the values, writes the Secret, and re-reads the provider every refreshInterval, which defaults to one hour.
A PersistentVolumeClaim in Kubernetes declares accessModes such as ReadWriteOnce, ReadOnlyMany, ReadWriteMany or ReadWriteOncePod. What does each of those mean, and what is the unit they are scoped to?
basics
~10 sReadWriteOnce: mounted read-write by one node. ReadOnlyMany: mounted read-only by many nodes. ReadWriteMany: mounted read-write by many nodes. ReadWriteOncePod: exactly one pod cluster-wide. Except for ReadWriteOncePod, the scope is the node, not the pod.
What is the Container Storage Interface (CSI) in Kubernetes, and why were storage drivers moved out of the Kubernetes core codebase onto it?
basics
~20 sCSI is a standard gRPC contract for storage drivers. A vendor ships the driver as ordinary pods in the cluster instead of code compiled into Kubernetes, so drivers install, upgrade and get fixed independently of the Kubernetes release.
What does the `volumeClaimTemplates` field on a Kubernetes StatefulSet do, and how does it differ from listing a PersistentVolumeClaim under a Deployment's pod volumes?
basics
~10 svolumeClaimTemplates makes the StatefulSet controller create one PersistentVolumeClaim per replica, named <template>-<statefulset>-<ordinal>. A Deployment's pod spec references one existing PVC, so every replica shares the same volume.
What is a Kubernetes StorageClass, and how does a PersistentVolumeClaim that references one end up with real storage mounted into a Pod?
basics
~20 sA StorageClass is a named storage tier: a provisioner plus its parameters. When a PersistentVolumeClaim names that class, the provisioner creates real storage and a matching PersistentVolume, which is bound to the claim. No admin pre-creates the volume.
In a Pod spec you can mount an emptyDir volume or mount a PersistentVolumeClaim. Explain the lifecycle difference between the two and when each is the right choice.
basics
~20 sAn emptyDir is created when the Pod is placed on a node and deleted with the Pod, so it survives container restarts but not rescheduling. A PersistentVolumeClaim points at storage whose lifetime is independent of the Pod, so data survives deletion and rescheduling.
What does node affinity do for a Kubernetes Pod, and how does it differ from the simpler nodeSelector field in the Pod spec?
basics
~20 sBoth steer a Pod onto nodes carrying particular labels. nodeSelector is exact key=value matching and every entry must match. Node affinity adds an expression language (In, NotIn, Exists, Gt, Lt), OR-ed alternatives, and a soft preferred form the scheduler may ignore.
In Kubernetes, how does a pod request a GPU, and why can it request 0.35 CPU but not 0.35 GPU?
basics
~20 sA pod requests a GPU as an extended resource, such as nvidia.com/gpu: 1, in its container limits. Extended resources are whole units the scheduler counts, not shares. So Kubernetes rejects fractions, requires the request to equal the limit, and never overcommits them.
What does a Kubernetes HorizontalPodAutoscaler object do, what must be true of the workload for it to work, and what does it deliberately not do?
basics
~20 sIt watches a metric (usually CPU) for a Deployment or StatefulSet and rewrites that workload's replica count, clamped between minReplicas and maxReplicas, to keep the metric near a target. Needs metrics-server plus resource requests on the pods. It adds and removes pods only — it never resizes a pod and never adds nodes.
What is KEDA in a Kubernetes cluster, and what does a KEDA ScaledObject declare about the workload it scales?
basics
~20 sKEDA is a Kubernetes operator that scales workloads on external event signals such as Kafka lag, queue depth or a Prometheus query. A ScaledObject names the target workload, its replica bounds and its triggers. KEDA turns that into an HPA it owns.
In a Kubernetes cluster with a spot node pool, which workloads belong on those nodes, and how do you keep everything else off them?
basics
~20 sSpot nodes suit replicated, stateless or checkpointing work that can lose a node at short notice. Label the pool by capacity type and taint it, so only pods that tolerate the taint and select the label land there.
In a Kubebuilder project built on controller-runtime, what does the framework provide, and what do you actually write inside Reconcile(ctx, req)?
basics
~20 scontroller-runtime supplies a manager that runs shared caches, clients, work queues, leader election and metrics; Kubebuilder scaffolds the API types, markers and Makefile. You write Reconcile: fetch the object named by req, converge actual state to spec, then update status.
What is a Kubernetes CustomResourceDefinition, and what exactly do you get once you apply one to a cluster?
basics
~20 sA CustomResourceDefinition registers a new resource type with the Kubernetes API server. You immediately get REST endpoints for it, storage in etcd, kubectl support, watches, RBAC and labels — but nothing acts on the objects until you run a controller for them.
A team wants automated backups, version upgrades and failover for a stateful database running on Kubernetes, all driven through the Kubernetes API. Explain what the Kubernetes Operator pattern is and how it delivers that.
basics
~20 sAn Operator is a custom Kubernetes resource type plus a controller running in the cluster. Users declare the desired database in a custom object; the controller continuously drives the real system to match that spec, automating install, backup, upgrade and failover.
In Kubernetes, a request that creates an object can be intercepted by webhooks registered through both a MutatingWebhookConfiguration and a ValidatingWebhookConfiguration. In what order does the API server run them relative to each other and to the rest of its request handling, and why does that order matter?
basics
~20 sThe API server authenticates and authorizes, then runs mutating admission (webhooks may patch the object), then schema validation, then validating admission (accept or reject only), then writes to etcd. Mutation runs first so validators judge the final stored object.
A CustomResourceDefinition in `apiextensions.k8s.io/v1` requires an OpenAPI v3 validation schema. What does that schema do to the objects users submit, what happens to fields the schema does not mention, and how would you deliberately allow arbitrary content in one part of the object?
basics
~20 sThe schema validates types, required fields and constraints at admission, applies declared defaults, and prunes — silently strips — any field it does not describe. To keep arbitrary content, mark that subtree with x-kubernetes-preserve-unknown-fields: true.
In Kubernetes, when and how does a namespace's LimitRange apply its default and defaultRequest values, and what happens to Pods that already exist?
basics
~10 sThe LimitRanger admission plugin in kube-apiserver copies defaultRequest and default into any container missing a request or limit at Pod creation, then rejects Pods outside min and max. Running Pods are never changed.
What does a Kubernetes namespace actually scope, and which resources, such as Nodes, PersistentVolumes and CRDs, sit outside every namespace?
basics
~10 sA Kubernetes namespace scopes object names and most workload objects: Pods, Services, Secrets, ConfigMaps, PVCs. Nodes, PersistentVolumes, StorageClasses, ClusterRoles, CRDs and Namespaces themselves are cluster-scoped. Run kubectl api-resources --namespaced=false to list them.
What is a Kubernetes ResourceQuota, and which kinds of consumption can its spec.hard field cap for one namespace?
basics
~20 sA ResourceQuota is a namespaced object whose spec.hard caps the namespace's total CPU and memory requests and limits, storage requests and object counts. The API server rejects any create that would push usage past a cap.
A Kubernetes namespace has been stuck in Terminating for 19 minutes; how do you find what is blocking the namespace controller's purge, and fix it?
basics
~20 sRead the namespace's status.conditions. They usually point to objects held by finalizers whose controller is gone, or to an unavailable aggregated API that makes discovery fail. Restore that controller or API, or remove the stale finalizer or APIService.
In Kubernetes, what is the difference between soft and hard multi-tenancy, and what do tenants still share when each gets only a namespace?
basics
~20 sSoft multi-tenancy gives each trusted team a namespace fenced by RBAC, quota and network policy. Hard multi-tenancy assumes hostile tenants and separates clusters or nodes, because namespaces still share the kernel, nodes, control plane and cluster-scoped objects.
In Kubernetes, what does metrics-server provide for `kubectl top` and the HorizontalPodAutoscaler, and what is it deliberately not designed to be?
basics
~20 smetrics-server scrapes current CPU and memory usage from every kubelet, keeps only the latest points in memory, and serves them as the metrics.k8s.io API. It is not a monitoring system: no history, no custom metrics, no alerting.
In a Kubernetes control plane, what is kube-apiserver responsible for, and why is it the only component that talks to etcd directly?
basics
~20 skube-apiserver is the cluster's front door: a REST API that authenticates, authorises, validates and persists every object, and streams changes to watchers. Routing all writes through it gives one place for auth, validation, admission and audit.
In a Kubernetes cluster, what does the kube-scheduler actually do when a new Pod is created, and what does it not do?
basics
~20 skube-scheduler watches for Pods with an empty spec.nodeName, picks a suitable node, and writes that choice back to the API server (a binding). It does not start the container — the kubelet on the chosen node does that.
In a Kubernetes control plane, what is etcd, what exactly is stored in it, and which components are allowed to talk to it?
basics
~20 setcd is the cluster's only persistent database: a distributed, strongly consistent key-value store holding every API object (Pods, Deployments, Services, ConfigMaps, Secrets, RBAC, node state). Only kube-apiserver connects to it; everything else reads and writes through the API server.
If every kube-apiserver in a Kubernetes cluster becomes unreachable, what keeps running on the nodes and what stops working?
basics
~20 sRunning Pods keep running: the kubelet keeps their containers alive and kube-proxy's installed Service rules keep routing. Anything that needs a write or fresh state stops working: scheduling, scaling, rollouts, rescheduling off failed nodes and kubectl.
Walk through what happens to a request such as 'kubectl apply -f pod.yaml' inside the Kubernetes API server, from the TLS connection to the object being stored in etcd.
basics
~20 sThe API server terminates TLS, then runs three gates in order: authentication (who are you — cert, service-account token, or OIDC token), authorization (may this identity do this verb on this resource — usually RBAC), and admission (mutating plugins/webhooks may edit the object, then validating ones may reject it). Only then is the object schema-validated and written to etcd.
In a Kubernetes Pod spec, what do the securityContext fields runAsNonRoot and runAsUser do, and why does it matter whether a container process runs as UID 0?
basics
~20 srunAsUser sets the numeric UID the container process runs as. runAsNonRoot: true makes the kubelet refuse to start the container if it would run as UID 0. Root inside the container is real root on the host kernel, so any container escape starts privileged.
In Kubernetes RBAC, what are the three parts of a policy rule — apiGroups, resources and verbs — and how would you write a Role that grants read-only access to Pods in a single namespace?
basics
~20 sA rule says which API groups, which resource types, and which verbs are allowed. Read-only pods is apiGroups: [""] (core group), resources: ["pods"], verbs: ["get","list","watch"]. A Role holds the rules; a RoleBinding attaches it to a user, group or ServiceAccount. RBAC is purely additive — there is no deny.
A teammate argues that Kubernetes Secret objects are safe to share and commit to git because their values are base64-encoded. What is wrong with that reasoning, and what actually protects Secret data in a cluster?
basics
~20 sBase64 is an encoding, not encryption: anyone decodes it in one command, no key needed. By default a Secret sits unencrypted in etcd and is readable by anyone whose RBAC allows reading Secrets. Real protection is least-privilege RBAC, encryption at rest, and keeping plaintext manifests out of git.
When a process running inside a Pod calls the Kubernetes API, what identity does it present and where does that credential come from? Why do security reviews object to leaving every workload on the namespace's default ServiceAccount?
basics
~20 sIt presents a ServiceAccount: a namespaced identity named by spec.serviceAccountName. The kubelet projects a signed JWT for it into the container at /var/run/secrets/kubernetes.io/serviceaccount/token. Leaving everything on default means all workloads share one identity, so any permission granted to it is granted to all of them and nothing is attributable.
When a Kubernetes cluster's API server becomes unreachable, what happens to pods that are already running, and what stops working?
basics
~20 sRunning pods keep running and existing Service traffic keeps flowing, because kubelets and kube-proxy work from state they already have. Anything that needs the API server stops: scheduling, rescheduling, scaling, rollouts, endpoint updates and CronJob runs.
How do you run a command or open an interactive shell inside a container of a running Kubernetes pod with kubectl, and what are the limits of that approach?
basics
~20 sUse kubectl exec -it POD -c CONTAINER -- sh. It starts an extra process inside an already-running container, streamed through the API server and the kubelet. The binary must exist in the image and the container must be running.
What does a Kubernetes container's imagePullPolicy control, what is its default, and how does it interact with images already cached on the node?
basics
~20 simagePullPolicy tells the kubelet when to check the registry: Always on every start, IfNotPresent only when the image is not cached, Never not at all. Unset, it defaults to Always for :latest or images with no tag and no digest, otherwise IfNotPresent.
How do you read container logs with kubectl, including output from a container instance that has already exited, and why might kubectl logs return nothing at all?
basics
~20 skubectl logs POD -c CONTAINER shows the current instance; --previous shows the last terminated one. Add -f to follow, --since and --tail to bound output. Nothing appears if the app logs to a file instead of stdout/stderr, the pod never started, or the logs rotated.
In Kubernetes, after a pod is deleted and its Deployment replaces it, can you still read the old pod's container logs with kubectl, and why?
basics
~20 sNo. Kubernetes keeps container logs only as files on the pod's node, tied to that pod; once the pod is deleted the kubelet removes its containers and log directory, and the replacement pod starts with empty logs.
In Kubernetes, what do `kubectl cordon`, `kubectl drain` and `kubectl uncordon` each do when you take a node out of service?
basics
~20 sCordon marks a node unschedulable so no new pods land there. Drain cordons it and then evicts its pods so their controllers recreate them elsewhere. Uncordon makes the node schedulable again without moving any pods back.
When a Kubernetes cluster is lost, what does an etcd snapshot bring back, and what must come from volume backups or Git instead?
basics
~20 sAn etcd snapshot restores Kubernetes API objects (Deployments, Secrets, RBAC, PVC and PV objects, custom resources) but none of the bytes stored on persistent volumes. Volume data needs CSI snapshots or Velero; Git re-creates only what was committed.
When a node runs kubeadm join with a bootstrap token and --discovery-token-ca-cert-hash, what does each value prove, and to whom?
basics
~20 sThe bootstrap token is a short-lived shared secret: the node checks a token-signed cluster-info and the API server accepts the new kubelet. The CA cert hash pins the cluster CA public key, so a leaked token cannot impersonate the cluster.
Why do organizations running Kubernetes split workloads across many clusters instead of one large cluster, and what does each driver buy them?
basics
~10 sA Kubernetes cluster is one failure and control domain, so teams run many clusters to limit blast radius, place workloads in specific regions, draw compliance boundaries and give tenants isolation that namespaces cannot provide.
Before moving a ticket-booking checkout service onto Kubernetes, what must the application itself change to run well as a container?
basics
~20 sThe app must read configuration from environment variables or mounted files, log to stdout and stderr, expose a readiness signal that reflects real ability to serve, shut down cleanly on SIGTERM, and size itself from its container limits.
On a Kubernetes platform run by a platform team, what is a golden-path template, and why use it instead of hand-written manifests?
basics
~20 sA golden-path template is a platform-maintained, pre-approved way to deploy a service: developers fill in a few values and the template produces the Deployment, Service, probes and resources, so every team gets safe defaults without writing raw Kubernetes manifests.
When PostgreSQL runs in a Kubernetes StatefulSet with per-replica PersistentVolumeClaims, what does Kubernetes actually provide, and what is still left for you to build?
basics
~20 sA StatefulSet gives each PostgreSQL pod a stable name, its own volume and ordered rollout. It knows nothing about which pod is primary, replication, failover, backups or safe upgrades, so all of that is still yours.
When moving a service from VMs into Kubernetes pods, which host-level assumptions usually break, and what replaces each one?
basics
~20 sPods are disposable and restart anywhere with a new IP and an empty filesystem, so local disk, fixed IPs, in-memory sessions, host cron and log files break. They move to object storage or a PVC, a Service name, a shared session store, a CronJob and stdout.
A team wants to run PostgreSQL and Kafka inside its Kubernetes cluster rather than use managed services. How do you decide, and what would make you say no?
basics
~20 sDecide by ownership and risk, not by preference. In-cluster operators suit teams that can run databases, need portability or have no managed option. Otherwise a managed service is usually cheaper overall. Say no when nobody can own recovery.