Pods and Workloads
Pods are what Kubernetes schedules, and a controller decides how many run and for how long: stateless replicas, stable identities, per-node agents, batch runs. The wrong controller only hurts during a deploy or an outage, which is why interviewers probe it.
part ofKubernetesoverview, primer and where to startread it →on this pageshowhide
explore
- Pod Model and Lifecycle5 questions
- Init and Sidecar Containers5 questions
- Deployments and ReplicaSets4 questions
- StatefulSets4 questions
- DaemonSets4 questions
- Jobs and CronJobs6 questions
- Liveness, Readiness, and Startup Probes5 questions
- Requests, Limits, and QoS Classes5 questions
- Graceful Shutdown3 questions
- Rollouts and Rollback4 questions
- Canary and Blue-Green4 questions
questions
page 2 of 2Kubernetes probes support exec, httpGet, tcpSocket and grpc handlers. How does each one work, and what are the failure modes of choosing the wrong one?
basics
~20 shttpGet: the kubelet calls the container's IP and port; status 200-399 is success. tcpSocket: success if a TCP connection opens - it proves only that a listener exists. exec: runs a command inside the container, exit 0 is success, but forks a process every period. grpc: calls the standard gRPC health checking service.
In Kubernetes, how do you run a canary with two Deployments behind one Service using replica counts, and where does that approach break down?
basics
~20 sPut stable and canary Deployments behind one Service that selects only their shared app label; the canary gets roughly its fraction of ready pods. That share moves in one-pod steps, is counted per connection, and drifts when either side scales.
A platform team wants every namespace to have sensible per-container resource defaults plus a hard cap on total consumption. Explain what the Kubernetes LimitRange and ResourceQuota objects each do, and how they interact.
basics
~20 sLimitRange is per-object: it defaults and bounds an individual container's or Pod's requests and limits at admission. ResourceQuota is per-namespace aggregate: it caps total requests, limits and object counts. A quota on a resource makes declaring it mandatory; LimitRange defaults satisfy that.
After `kubectl rollout undo deployment/transcoder --to-revision=4`, what exactly has Kubernetes reverted, and which revision number does the Deployment then report?
basics
~20 sOnly the pod template is reverted: kubectl copies revision 4's template back into the Deployment. The controller reuses that old ReplicaSet and renumbers it to the next revision, so the Deployment reports a new number, not 4.
A Kubernetes StatefulSet has a field called podManagementPolicy that accepts OrderedReady or Parallel. What does each setting change, and when would you choose Parallel?
basics
~20 sOrderedReady (the default) creates Pods one at a time in ordinal order, waiting for each to be Running and Ready, and deletes in reverse one at a time. Parallel starts and stops them all at once. Neither changes rolling-update ordering, which stays sequential.
How does a Kubernetes DaemonSet roll a new pod template out across a large fleet of nodes, and which settings control the speed and blast radius of that rollout?
basics
~20 sWith updateStrategy RollingUpdate (the default), the controller replaces pods node by node, keeping at most maxUnavailable (default 1) unavailable at a time; maxSurge can instead start the new pod before deleting the old. OnDelete updates a node only when you delete its pod manually.
A Kubernetes Deployment rollout has been stuck partway for ten minutes: some new pods are up, the old ones are still serving. How do you diagnose it, and what does the ProgressDeadlineExceeded condition on a Deployment mean?
basics
~20 sLook at the Deployment's conditions and events, then at the new ReplicaSet and its pods. ProgressDeadlineExceeded means the rollout made no progress within progressDeadlineSeconds (default 600); it is a status marker only — Kubernetes does not roll back for you.
A Kubernetes Job's Pod runs a batch process plus a log-shipping helper container. The batch process exits 0, yet the Pod stays in `Running` and the Job never reports completion. What is happening, and what is the correct fix?
basics
~20 sA Job's Pod is only complete when its regular containers have all exited, and the log shipper never exits. Fix it by declaring the helper as a native sidecar — an initContainers entry with restartPolicy: Always — so completion is judged on the batch container alone.
An application container in a Kubernetes Pod sends all of its outbound traffic through a proxy container in the same Pod, and it fails its first few requests every time the Pod starts. Explain why, and how you make the proxy fully up before the application container starts.
basics
~20 sContainers listed in a Pod's containers field start in parallel, so the app can beat the proxy. Fix it by declaring the proxy in initContainers with restartPolicy: Always and a startupProbe — the kubelet then waits for the probe to pass before starting the app.
A Kubernetes CronJob stopped firing after the cluster control plane was unavailable for several hours, and the controller logs mention too many missed start times. Explain the mechanism behind that, and what role `startingDeadlineSeconds` plays.
basics
~20 sThe CronJob controller counts every schedule point missed since the last run; if it finds more than 100 it gives up and logs an error instead of guessing. startingDeadlineSeconds limits how far back it looks and how late a missed run may still start — setting it below 100 intervals prevents the lock-up.
In Kubernetes Gateway API, how do HTTPRoute backendRefs weights give a search-autocomplete canary a 0.5% share, and which traffic do those weights never reach?
basics
~20 sList a stable and a canary Service in one HTTPRoute rule with weights 995 and 5; the gateway sends 0.5% of requests to the canary whatever its pod count. Pod-to-pod calls via the Service ClusterIP bypass those weights.
A CI job gates a Kubernetes transcoding Deployment's release on `kubectl rollout status`, and workers take about 7 minutes to turn Ready. How do you set progressDeadlineSeconds, minReadySeconds and --timeout so the gate fails correctly?
basics
~20 sprogressDeadlineSeconds measures the longest gap between progress events, not the whole rollout, so it must exceed one pod's time to Ready plus margin. minReadySeconds catches early crashes, and a --timeout above the total rollout time is the backstop.
Kubernetes StatefulSet rolling updates support a `spec.updateStrategy.rollingUpdate.partition` value. Explain how it works and what you would use it for.
basics
~20 sStatefulSet rolling updates go in reverse ordinal order, one Pod at a time. Setting partition to k means only Pods with ordinal greater than or equal to k are updated; lower ordinals stay on the old revision. Setting it to replicas minus one gives a single-Pod canary.
On a 3-node kubeadm Kubernetes cluster, why would each old payments-authorization API pod sit in Terminating for 47 seconds during a rollout, then exit with code 137 mid-request, and how do you fix it?
basics
~20 sAn exec preStop sleep longer than terminationGracePeriodSeconds consumes the whole budget; the kubelet then allows only a 2-second stop window, so SIGKILL follows SIGTERM. Shorten the sleep and size the grace period to sleep plus drain.
How do you decide whether two processes belong in the same Kubernetes Pod or in two separate Pods? Give the criteria you apply and what each choice costs you.
basics
~20 sSame Pod when they must share fate, scale 1:1, and need localhost or a shared filesystem. Separate Pods when they scale, release or fail independently. A Pod is the unit of scheduling, scaling, rollout and blast radius, so co-locating couples all four at once.
As the owner of a shared Kubernetes platform, when do hand-built canary and blue-green releases from Deployments, Services and HTTPRoute weights stop being enough, and what replaces them?
basics
~20 sHand-built releases suffice while few services release rarely under a watching person. When many teams need automatic bake-and-abort, precise splits and an audit trail, a progressive-delivery controller or mesh should own the steps; core Kubernetes never judges a release.
You own a shared, multi-tenant Kubernetes cluster. How do you decide what CPU and memory requests and limits to set for a service, and what is your policy on overcommitting CPU versus memory?
basics
~20 sSize requests from observed usage percentiles plus headroom, not guesses. Overcommit CPU deliberately — it is compressible and degrades gracefully. Do not overcommit memory: set memory request equal to limit so failures stay contained and attributable rather than node-wide.
A workload needs an auxiliary concern — fetching config at startup, renewing TLS certificates, shipping logs, or applying a database schema change. How do you decide between an init container, a long-running sidecar container, a separate workload object, or building the behaviour into the main image?
basics
~20 sDecide by lifetime and blast radius: one-shot precondition for this Pod means init container; continuous work bound to this Pod's lifetime means sidecar; work that must happen once per release or scale independently means its own Job or Deployment; universal, stable behaviour with no separate lifecycle belongs in the image.
Kubernetes documents that a CronJob may create a Job twice for one schedule point, or not at all. Given that guarantee, how would you design a nightly billing batch that must charge each customer exactly once?
basics
~20 sTreat the schedule as at-least-once and make exactly-once a property of the payload: derive a deterministic run key from the period, take a database lock or unique-constraint claim on it, make every charge idempotent with a per-invoice key, and alert on a heartbeat rather than on Job failures.
showing 31–49 of 49