skip to content

Pods and Workloads

Pods are what Kubernetes schedules, and a controller decides how many run and for how long: stateless replicas, stable identities, per-node agents, batch runs. The wrong controller only hurts during a deploy or an outage, which is why interviewers probe it.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

Kubernetes probes support exec, httpGet, tcpSocket and grpc handlers. How does each one work, and what are the failure modes of choosing the wrong one?

level: middleimportance: should knowfreq 50%

basics

~20 s

httpGet: the kubelet calls the container's IP and port; status 200-399 is success. tcpSocket: success if a TCP connection opens - it proves only that a listener exists. exec: runs a command inside the container, exit 0 is success, but forks a process every period. grpc: calls the standard gRPC health checking service.

open as a page

In Kubernetes, how do you run a canary with two Deployments behind one Service using replica counts, and where does that approach break down?

level: middleimportance: should knowfreq 55%

basics

~20 s

Put stable and canary Deployments behind one Service that selects only their shared app label; the canary gets roughly its fraction of ready pods. That share moves in one-pod steps, is counted per connection, and drifts when either side scales.

open as a page

A platform team wants every namespace to have sensible per-container resource defaults plus a hard cap on total consumption. Explain what the Kubernetes LimitRange and ResourceQuota objects each do, and how they interact.

level: middleimportance: should knowfreq 42%

basics

~20 s

LimitRange is per-object: it defaults and bounds an individual container's or Pod's requests and limits at admission. ResourceQuota is per-namespace aggregate: it caps total requests, limits and object counts. A quota on a resource makes declaring it mandatory; LimitRange defaults satisfy that.

open as a page

After `kubectl rollout undo deployment/transcoder --to-revision=4`, what exactly has Kubernetes reverted, and which revision number does the Deployment then report?

level: middleimportance: should knowfreq 42%

basics

~20 s

Only the pod template is reverted: kubectl copies revision 4's template back into the Deployment. The controller reuses that old ReplicaSet and renumbers it to the next revision, so the Deployment reports a new number, not 4.

open as a page

A Kubernetes StatefulSet has a field called podManagementPolicy that accepts OrderedReady or Parallel. What does each setting change, and when would you choose Parallel?

level: middleimportance: should knowfreq 38%

basics

~20 s

OrderedReady (the default) creates Pods one at a time in ordinal order, waiting for each to be Running and Ready, and deletes in reverse one at a time. Parallel starts and stops them all at once. Neither changes rolling-update ordering, which stays sequential.

open as a page

How does a Kubernetes DaemonSet roll a new pod template out across a large fleet of nodes, and which settings control the speed and blast radius of that rollout?

level: seniorimportance: should knowfreq 42%

basics

~20 s

With updateStrategy RollingUpdate (the default), the controller replaces pods node by node, keeping at most maxUnavailable (default 1) unavailable at a time; maxSurge can instead start the new pod before deleting the old. OnDelete updates a node only when you delete its pod manually.

open as a page

A Kubernetes Deployment rollout has been stuck partway for ten minutes: some new pods are up, the old ones are still serving. How do you diagnose it, and what does the ProgressDeadlineExceeded condition on a Deployment mean?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Look at the Deployment's conditions and events, then at the new ReplicaSet and its pods. ProgressDeadlineExceeded means the rollout made no progress within progressDeadlineSeconds (default 600); it is a status marker only — Kubernetes does not roll back for you.

open as a page

A Kubernetes Job's Pod runs a batch process plus a log-shipping helper container. The batch process exits 0, yet the Pod stays in `Running` and the Job never reports completion. What is happening, and what is the correct fix?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A Job's Pod is only complete when its regular containers have all exited, and the log shipper never exits. Fix it by declaring the helper as a native sidecar — an initContainers entry with restartPolicy: Always — so completion is judged on the batch container alone.

open as a page

An application container in a Kubernetes Pod sends all of its outbound traffic through a proxy container in the same Pod, and it fails its first few requests every time the Pod starts. Explain why, and how you make the proxy fully up before the application container starts.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Containers listed in a Pod's containers field start in parallel, so the app can beat the proxy. Fix it by declaring the proxy in initContainers with restartPolicy: Always and a startupProbe — the kubelet then waits for the probe to pass before starting the app.

open as a page

A Kubernetes CronJob stopped firing after the cluster control plane was unavailable for several hours, and the controller logs mention too many missed start times. Explain the mechanism behind that, and what role `startingDeadlineSeconds` plays.

level: seniorimportance: should knowfreq 38%

basics

~20 s

The CronJob controller counts every schedule point missed since the last run; if it finds more than 100 it gives up and logs an error instead of guessing. startingDeadlineSeconds limits how far back it looks and how late a missed run may still start — setting it below 100 intervals prevents the lock-up.

open as a page

In Kubernetes Gateway API, how do HTTPRoute backendRefs weights give a search-autocomplete canary a 0.5% share, and which traffic do those weights never reach?

level: seniorimportance: should knowfreq 42%

basics

~20 s

List a stable and a canary Service in one HTTPRoute rule with weights 995 and 5; the gateway sends 0.5% of requests to the canary whatever its pod count. Pod-to-pod calls via the Service ClusterIP bypass those weights.

open as a page

A CI job gates a Kubernetes transcoding Deployment's release on `kubectl rollout status`, and workers take about 7 minutes to turn Ready. How do you set progressDeadlineSeconds, minReadySeconds and --timeout so the gate fails correctly?

level: seniorimportance: should knowfreq 38%

basics

~20 s

progressDeadlineSeconds measures the longest gap between progress events, not the whole rollout, so it must exceed one pod's time to Ready plus margin. minReadySeconds catches early crashes, and a --timeout above the total rollout time is the backstop.

open as a page

Kubernetes StatefulSet rolling updates support a `spec.updateStrategy.rollingUpdate.partition` value. Explain how it works and what you would use it for.

level: seniorimportance: should knowfreq 35%

basics

~20 s

StatefulSet rolling updates go in reverse ordinal order, one Pod at a time. Setting partition to k means only Pods with ordinal greater than or equal to k are updated; lower ordinals stay on the old revision. Setting it to replicas minus one gives a single-Pod canary.

open as a page

On a 3-node kubeadm Kubernetes cluster, why would each old payments-authorization API pod sit in Terminating for 47 seconds during a rollout, then exit with code 137 mid-request, and how do you fix it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An exec preStop sleep longer than terminationGracePeriodSeconds consumes the whole budget; the kubelet then allows only a 2-second stop window, so SIGKILL follows SIGTERM. Shorten the sleep and size the grace period to sleep plus drain.

open as a page

How do you decide whether two processes belong in the same Kubernetes Pod or in two separate Pods? Give the criteria you apply and what each choice costs you.

level: principalimportance: should knowfreq 40%

basics

~20 s

Same Pod when they must share fate, scale 1:1, and need localhost or a shared filesystem. Separate Pods when they scale, release or fail independently. A Pod is the unit of scheduling, scaling, rollout and blast radius, so co-locating couples all four at once.

open as a page

As the owner of a shared Kubernetes platform, when do hand-built canary and blue-green releases from Deployments, Services and HTTPRoute weights stop being enough, and what replaces them?

level: principalimportance: should knowfreq 34%

basics

~20 s

Hand-built releases suffice while few services release rarely under a watching person. When many teams need automatic bake-and-abort, precise splits and an audit trail, a progressive-delivery controller or mesh should own the steps; core Kubernetes never judges a release.

open as a page

You own a shared, multi-tenant Kubernetes cluster. How do you decide what CPU and memory requests and limits to set for a service, and what is your policy on overcommitting CPU versus memory?

level: principalimportance: should knowfreq 44%

basics

~20 s

Size requests from observed usage percentiles plus headroom, not guesses. Overcommit CPU deliberately — it is compressible and degrades gracefully. Do not overcommit memory: set memory request equal to limit so failures stay contained and attributable rather than node-wide.

open as a page

A workload needs an auxiliary concern — fetching config at startup, renewing TLS certificates, shipping logs, or applying a database schema change. How do you decide between an init container, a long-running sidecar container, a separate workload object, or building the behaviour into the main image?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Decide by lifetime and blast radius: one-shot precondition for this Pod means init container; continuous work bound to this Pod's lifetime means sidecar; work that must happen once per release or scale independently means its own Job or Deployment; universal, stable behaviour with no separate lifecycle belongs in the image.

open as a page

Kubernetes documents that a CronJob may create a Job twice for one schedule point, or not at all. Given that guarantee, how would you design a nightly billing batch that must charge each customer exactly once?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Treat the schedule as at-least-once and make exactly-once a property of the payload: derive a deterministic run key from the period, take a database lock or unique-constraint claim on it, make every charge idempotent with a per-invoice key, and alert on a heartbeat rather than on Job failures.

open as a page

showing 31–49 of 49