When moving a service from VMs into Kubernetes pods, which host-level assumptions usually break, and what replaces each one?
answer
- pods are replaced, not repaired
- new IP, empty writable layer
- disk, IP, session, cron, logs, processes
- object store, Service name, shared session store
- CronJob and stdout
basics
~20 sPods are disposable and restart anywhere with a new IP and an empty filesystem, so local disk, fixed IPs, in-memory sessions, host cron and log files break. They move to object storage or a PVC, a Service name, a shared session store, a CronJob and stdout.
solid answer
~40 sA VM is a long-lived machine; a Kubernetes pod is a replaceable unit that a node drain, eviction or rollout deletes and recreates elsewhere with a new IP and an empty writable layer. So I audit the VM for six assumptions. Files written to local disk go to object storage, or to a `PersistentVolumeClaim` when the app truly needs a filesystem. Callers that use the VM's IP switch to a Service's stable DNS name. In-memory sessions behind a sticky load balancer move to a shared store so any replica can serve any request. Host crontab entries become `CronJob` objects. Log files become stdout/stderr, which the kubelet captures and `kubectl logs` reads. And a supervisor running several daemons is split so each container has one main process.
code
yaml · 15 linesapiVersion: v1
kind: Pod
metadata:
name: shipment-tracking-label-render
spec:
containers:
- name: api
image: registry.example.com/shipment-tracking:3.14.2
volumeMounts:
- name: scratch
mountPath: /tmp/labels
volumes:
- name: scratch
emptyDir:
sizeLimit: 512Migo deeper
Remember the core fact: a pod can be deleted and recreated elsewhere at any time, with a new IP and a fresh filesystem. Most migration breakages follow from that.
Walk through each assumption and name its Kubernetes replacement: object storage or PVC, Service DNS, shared session store, CronJob, stdout, and one main process per container.
Show you would prove the migration by forcing churn: drain nodes and kill pods under load, and treat any lost file or session as a blocker before cutover.
Frame the audit as a cost decision: which assumptions to fix on the VMs before moving, and which justify leaving a component outside the cluster entirely.
## Why the VM model breaks On a VM, the machine is the unit of identity: it has a hostname, a fixed **IP address**, a disk that survives reboots, a crontab and a `/var/log` directory that an operator can SSH into. Applications quietly depend on all of that. In Kubernetes the unit is the **Pod**, and pods are designed to be replaced. A `kubectl drain` during node maintenance, a rolling update of a Deployment, kubelet node-pressure eviction or a node failure all delete the pod; the Deployment's ReplicaSet then creates a *new* pod, usually on a different node. That new pod gets a new IP, and its container starts from the image with an empty **writable layer**. Nothing the old container wrote to its own filesystem comes along. Take a **shipment-tracking API** that ran on three VMs and is moving to a **12-node cluster** that also serves GPU models. When the platform team rotates nodes, each drain of a busy node takes about **13 minutes**, and every pod on it is recreated somewhere else. Any assumption tied to one machine fails during that window. ## The six assumptions and their replacements | VM assumption | What breaks in Kubernetes | Replacement | |---|---|---| | Files on local disk persist | writable layer is discarded with the pod | object storage, or a `PersistentVolumeClaim` | | The host has a fixed IP | pod IPs change on every reschedule | a Service's DNS name; an egress gateway for outbound allowlists | | Sessions live in process memory | the next request may hit another replica | a shared session store or stateless tokens | | Host crontab runs jobs | there is no host to own the crontab | a `CronJob` object | | Logs are files in `/var/log` | files vanish with the pod and nobody collects them | write to stdout/stderr | | One box runs several daemons | the kubelet only watches each container's main process | one main process per container | ## Local disk The tracking API wrote proof-of-delivery photos to `/data`. In a pod with no volume, those files live in the container's writable layer and disappear when the pod is deleted. Choices, in order of preference: - **Object storage** for blobs: every replica can read every photo, and there is no volume to attach. - A **PersistentVolumeClaim** when the software insists on a POSIX filesystem. Remember that most block storage attaches to one node at a time, which constrains scaling. - An **`emptyDir`** volume for true scratch space, such as a PDF being rendered. It is deleted with the pod, and with `sizeLimit` set the kubelet evicts the pod if it grows past that size. ## Network identity and sessions Other systems often address the VM by IP: config files, `/etc/hosts` entries, firewall rules. Inside the cluster, callers should use a **Service** name, which stays stable while pods behind it change. Outbound allowlists at partners need a stable egress address, which is a separate problem. In-memory sessions only worked because a sticky load balancer kept each user on one VM. Service-level `sessionAffinity: ClientIP` or ingress cookie affinity can reduce the misses, but when the pod is deleted the session is gone. The durable fix is moving session state to a **shared store** or to signed, stateless tokens. ## Scheduled work and logs A crontab has no home: putting `crond` inside the image runs the job once *per replica*. The Kubernetes object for scheduled work is a **`CronJob`**, which creates a Job, and so a pod, at each schedule point. For logs, the container runtime captures **stdout and stderr** into files on the node. `kubectl logs` reads those streams, and a node-level log agent ships them. The kubelet rotates them, by default at `containerLogMaxSize` 10Mi with `containerLogMaxFiles` 5. A file the app writes inside the container is invisible to all of this. ## An audit order that works 1. Inventory the VM: `crontab -l`, open files and write paths, listening ports, supervisor configs, firewall rules that mention its IP. 2. Fix state first (disk, sessions), because those cause data loss rather than errors. 3. Replace cron and logging, which are mechanical. 4. Split processes and deploy. 5. Prove it by draining a node on purpose and watching nothing break.
- The shipment-tracking API renders label PDFs to a temp directory before uploading them. Does that directory need a PersistentVolumeClaim?No. Data that only has to live as long as the pod belongs in an `emptyDir` volume. It is shared by the pod's containers, deleted with the pod, and with `sizeLimit` set the kubelet evicts the pod if it grows past the limit. A PVC would add attach time and node constraints for data nobody needs after a restart.
- How do you find these host assumptions before the migration instead of in production?Inventory the VM: crontab entries, supervisor configs, paths the process writes to, config files with hard-coded hosts or IPs, and firewall rules naming the VM. Then run the service in pods on a staging cluster and deliberately delete pods and drain a node during a load test. Anything that loses data or sessions under that churn is an assumption you missed.
A VM is a house you own, where you can leave things in the attic; a pod is a hotel room, and whatever you leave behind is gone when housekeeping resets it for the next guest.
saying these in an interview costs you the question
- Pods keep their IP address when they are recreated on another node
- Files in the container filesystem survive because the pod restarts in place
- Sticky load balancing makes in-memory sessions safe in Kubernetes
- kubectl logs reads the application's log files inside the container
- Leaving crond in the image is fine because only one pod runs it