In Kubernetes, what exactly is a Pod, what do the containers inside one share, and what is the 'pause' (sandbox) container for?
answer
- smallest schedulable unit, not the container
- one IP, one port space, localhost
- volumes shared, root filesystem not
- pause container owns the namespaces
- Pods are mortal, never rescheduled
basics
~20 sA Pod is Kubernetes' smallest schedulable unit: one or more containers placed on the same node, sharing one network namespace (one Pod IP, one port space, localhost), IPC and mounted volumes. A tiny 'pause' container holds those namespaces open.
solid answer
~50 sA Pod, not a container, is what Kubernetes schedules and assigns an IP to. Containers in a Pod always land on one node and share: - the **network namespace** - one Pod IP, one port space; they reach each other over `localhost`; - **IPC**, and optionally the PID namespace (`shareProcessNamespace: true`); - **volumes** declared in the Pod spec and mounted per container. They do *not* share a root filesystem: each container keeps its own image and mount namespace, so files move between them only through a shared volume. Underneath, the kubelet asks the runtime to create a *sandbox* first: the **pause** container, a near-empty process that sleeps forever while owning the namespaces. App containers join those namespaces, so a crashed app container can be restarted in place and keep the same Pod IP. Destroy the sandbox and the whole Pod is gone. Most Pods hold exactly one app container; extra containers belong there only when they must share fate, network and locality.
code
yaml · 22 linesapiVersion: v1
kind: Pod
metadata:
name: web
spec:
volumes:
- name: content
emptyDir: {}
containers:
- name: nginx
image: nginx:1.27
ports:
- containerPort: 80
volumeMounts:
- name: content
mountPath: /usr/share/nginx/html
- name: updater
image: alpine:3.20
command: ["sh", "-c", "while true; do date > /data/index.html; sleep 10; done"]
volumeMounts:
- name: content
mountPath: /datago deeper
Say the Pod is the smallest deployable unit, holds one or more containers on one node, and gives them a shared IP and shared volumes.
Name the specific namespaces that are shared (network, IPC, UTS) versus private (mount, and PID unless shareProcessNamespace is set), and explain the pause/sandbox container.
Draw the consequences: stable Pod IP across container restarts, port conflicts inside a Pod, Pods being mortal and never rescheduled, and why controllers rather than bare Pods run production workloads.
Frame the Pod as the platform's fate-sharing and co-scheduling boundary, and discuss when multi-container Pods are justified versus separate Pods that scale and release independently.
## The problem a Pod solves A container image packages one process tree. Occasionally two processes are so tightly coupled that they behave as one program: a server plus a log shipper reading the same directory, or an app plus a local proxy it talks to on `localhost`. Rather than teach every controller, scheduler and network plugin about groups of containers, Kubernetes introduced one wrapper object - the **Pod** - and made *it* the unit of scheduling, addressing and lifecycle. Everything above (Deployments, Jobs, Services) manipulates Pods, never containers directly. ## What is shared Containers in a Pod are guaranteed to run on the same node, and they share: - **Network namespace.** The Pod gets one IP address. All containers see the same interfaces and the same port space, so they can call each other on `localhost:<port>` with no service discovery - and they *conflict* if two of them bind the same port. - **IPC namespace**, so SystemV IPC and POSIX shared memory work between them. - **Volumes.** Volumes are declared once at Pod level (`spec.volumes`) and mounted independently in each container (`volumeMounts`). This is the only supported way to exchange files. - **UTS namespace** (hostname), plus the node, the scheduling constraints and the lifecycle. Optional: `shareProcessNamespace: true` puts all containers in one PID namespace so they can see and signal each other's processes. ## What is NOT shared Each container keeps **its own image, its own root filesystem and mount namespace, its own resource requests/limits, its own probes and its own restart accounting**. Writing a file in container A is invisible to container B unless the path is a shared volume. Environment variables are per container too. ## The pause container Namespaces need an owning process. If the app container were that owner, restarting it after a crash would tear the namespaces down - and the Pod would lose its IP mid-life. So the kubelet, through the CRI, first creates a **sandbox**: a minimal container (the `pause` image, a handful of kilobytes) whose entire job is to call `pause()` and hold the namespaces open, and to reap orphaned processes when the PID namespace is shared. The CNI plugin wires the Pod IP into that sandbox. App containers are then created *joining* the sandbox's namespaces. Consequences worth stating in an interview: an app container can crash and restart many times while the Pod IP stays constant; and if the sandbox itself dies, the whole Pod is recreated - for a controller-managed workload that means a replacement Pod with a new name and IP. ## Pods are not durable A Pod is a mortal, mostly immutable object. Once created you can change very little (container image, `activeDeadlineSeconds`, tolerations - and, where enabled in recent versions, in-place resource resize); you cannot repoint it to another node. A Pod is never rescheduled: if the node dies, that Pod object is marked for deletion or sits Unknown, and a controller must create a *new* Pod. That is why you almost never write bare Pod manifests in production. ## When to put two containers in one Pod Use one Pod when the containers must scale 1:1, start and stop together, and communicate over `localhost` or a shared volume. Use separate Pods when they scale independently, are released independently, or should not die together. "They are related" is not a reason; "they are useless apart and must share a filesystem or a port space" is. ## Inspecting it `kubectl get pod -o wide` shows the Pod IP and node. `kubectl describe pod` lists containers, volumes and sandbox events. `kubectl exec -c <container>` must select one container, precisely because they do not share a filesystem.
- Two containers in the same Pod both listen on port 8080. What happens?They share one network namespace and therefore one port space, so the second one fails to bind with 'address already in use' and crashes. You must give them different ports; there is no per-container port isolation inside a Pod.
- Do containers in a Pod share a filesystem?No. Each container has its own image and mount namespace, so its root filesystem is private. The only shared storage is a volume declared in spec.volumes and mounted into both containers, which then see the same files at possibly different paths.
A Pod is like a single apartment: the rooms (containers) have their own furniture but share one street address, one phone line and the hallway closet. The pause container is the lease that keeps the address reserved while you swap the furniture.
saying these in an interview costs you the question
- Saying a Pod is 'just another word for a container'
- Claiming containers in a Pod share the whole filesystem by default
- Believing a Pod is rescheduled to another node when its node fails
- Thinking each container in a Pod gets its own IP address
- Putting unrelated services in one Pod 'to save resources'