A suspect Kubernetes pod may reschedule at any moment - what do you capture, and what can wait?
answer
- sort by what the reschedule destroys
- the durable set can wait
- digest, not tag; UID, not name
- pod IPs get recycled
- stop the clock before collecting
basics
~20 sCapture only what dies with the container: the suspect process's memory, the container identity and image digest, the writable layer, live sockets and the pod-IP mapping, and the pod object. The image, audit records and shipped telemetry can wait.
solid answer
~50 sSort by what the next lifecycle event destroys, not by what looks interesting. Perishable: the suspect process's memory, because a memory-only loader has no other copy; the container identity - container id, pod UID, node, image reference and its `sha256` digest; the writable layer, which holds anything staged or dropped; the socket table and the pod-IP-to-pod mapping, since cluster pod IPs are recycled and flow records become uninterpretable without it; and the pod object with its mounts and projected tokens before it is garbage-collected. Durable, so leave it: the image, immutable and pullable by digest; Kubernetes audit events retained at the API server; runtime, proxy and flow telemetry already off the node. And the highest-value move is usually not collecting faster - pause the rollout, hold the autoscaler off the node, cordon it, so the deadline stops running at all.
go deeper
Be able to separate what exists only inside a running container and disappears on reschedule from what is stored elsewhere and can be fetched at leisure.
Explain the sorting rule - capture by what the next lifecycle event destroys - and defend the image digest and the pod-IP mapping as urgent rather than optional metadata.
Show that your first instinct is to remove the deadline: pause the rollout, hold the autoscaler off, cordon the node, then collect calmly. Interviewers watch for whether you race a clock you could have stopped.
Own whether this is achievable at all in practice. If collecting the perishable set depends on a specialist being awake, most intrusions will lose it, and the answer is what the platform captures automatically.
## The sorting rule Under a reschedule deadline the useful question is not "what evidence would I like" but **"what does the next container lifecycle event destroy, and what is stored somewhere that outlives it"**. Sort every candidate artefact into those two buckets and spend the window entirely on the first. Almost everyone under time pressure does the reverse, because the durable things are the easy ones to grab. ## The perishable set, and why each item earns its place **The suspect process's memory.** A loader that runs in memory has no on-disk copy anywhere: not in the image, not in the writable layer, not in a registry. If you take one thing, take this. It is also the artefact that answers the questions you will be asked afterwards — what the code did, what it was configured to talk to, what it had already decrypted or collected. **The container's identity.** Container id, pod name and **pod UID**, namespace, node, start time, the image reference *and its sha256 digest*, the command and arguments it was started with. This is small, fast, and it is what makes everything else interpretable months later. Tags are mutable — `payments-api:1.4.2` can point to different bytes next week — so a tag is a name, while a digest is the actual thing that ran. **The writable layer.** Whatever was dropped: staged archives, added tooling, scripts, partial output. In a container this is usually tiny compared with a host disk, so it is a cheap grab with a high hit rate. Memory tells you what was executing; the layer tells you what was staged. Neither implies the other. **Live network and process state.** The socket table with peers, the process tree with parents, open file descriptors. Plus the **pod IP to pod mapping**: cluster pod IPs come from a pool and are recycled, so a flow record for `10.42.3.17` becomes uninterpretable once the mapping is gone. That single line of metadata is what keeps all your network evidence attributable. **The pod object itself.** The live spec and status — mounted volumes, projected tokens, environment references, service account, node assignment, restart history. Once the object is garbage-collected, reconstructing what the workload had access to becomes archaeology. ## The set that can wait - **The image.** Immutable and pullable by digest whenever you like. Analysing it during the window is a common and expensive mistake. - **Kubernetes audit events.** Written at the API server, retained centrally, unaffected by the pod's life. They will tell you who called what, including whether an interactive `pods/exec` was requested — later. - **Runtime sensor and EDR telemetry, flow records, proxy and DNS logs.** Already off the node. Subject to retention, not to the pod's lifetime — check the retention window, but do not spend the deadline on them. - **Cloud control-plane logs** for whatever the pod's identity touched. Same reasoning. ## The move that beats capturing faster The single highest-value action is often not collection at all: **stop the thing that will destroy the pod**. Pause the Deployment's rollout so a reconcile does not replace it, exclude the node from autoscaler scale-down, cordon the node so a subsequent drain is less likely, and tell the on-call platform engineer what you are doing before they do the obvious thing. Cordoning marks the node unschedulable and leaves the running pods alone — it does not evict anything, which is exactly why it is safe to do immediately and why it is not, on its own, containment. Removing the deadline converts a five-minute scramble into an unhurried collection, and a truncated capture is worth far less than a complete one: a process image taken while the process is running has no single point in time to begin with, and a truncated one may hold only fragments. ## Doing this while the pod is still dangerous Capturing is not a reason to leave a live foothold untouched. The two are separable: deny the pod's egress with a network policy, revoke the service account's and any cloud credentials it holds, and block its known peers at the egress gateway. The container keeps running — and stays capturable — while the operator's reach is cut. Whether that is enough, and who authorises the wait, is the harder decision that follows. ## What this looks like when it is answered badly Weak answers start with the image because it is the familiar artefact, describe collecting logs that were never at risk, or spend the window arguing in the incident channel while an autoscaler quietly removes the node. The strongest answers are short: stop the clock, record the identity and digest, take memory, take the layer, then argue about everything else with the pod still alive.
- Why record the image digest when you already have the image tag?Tags are mutable - `payments-api:1.4.2` can point to different bytes next week, and in an actively deployed service it often does. The sha256 digest names exactly the bytes the node ran, so months later you can pull that build, diff it against a known-clean one, and say whether the intruder's code was in the image. Without the digest, 'the image was clean' is a claim about a name.
- Why is capturing the pod-IP-to-pod mapping urgent when the flow records are safe?Because the flow records are only interpretable through that mapping. Cluster pod IPs are allocated from a pool and reused within minutes, so a record for 10.42.3.17 attributes to nothing unless you know which pod held that address at that time. The mapping lives in the pod object and node state, which vanish with the pod - capture it and all your network evidence stays attributable.
- Is grabbing the writable layer worth it when you already have process memory?Yes, because they answer different questions. Memory shows what was executing and how it was configured; the writable layer shows what was staged - dropped tooling, archives being prepared, scripts, keys written out. Neither implies the other, and in a container the layer is usually small, so it is a cheap grab compared with imaging a host disk.
saying these in an interview costs you the question
- Starts by pulling the image because it is familiar
- Spends the window exporting logs that were never at risk
- Collects for an hour without pausing the controller
- Records the mutable image tag instead of the digest
- Skips the pod-IP mapping, then cannot attribute flow data