skip to content

Answering For It

What you owe afterwards — evidence kept while remediating under pressure, the notification clock a determination starts, and a retro whose output is a detection. Interviewers probe those trade-offs.

on this pageshow

explore

questions

16

A flagged Kubernetes pod was deleted and replaced before anyone captured it - what is gone for good?

level: juniorimportance: must knowfreq 62%

answer

  1. what only that container held
  2. shipped logs record emissions
  3. the loader never touched disk
  4. no copy means nothing to hash

basics

~20 s

Everything that existed only inside that container: the process memory holding a memory-only loader, its live network and process state, and the writable layer's dropped files. Shipped logs survive, but they record what was emitted, never the code that ran.

solid answer

~50 s

Deleting the pod destroys the only copy of anything that never touched durable storage. A memory-only loader lives in the process address space, so its decoded payload, in-memory configuration, keys and established sockets go with the container, along with whatever it dropped into the writable layer. What survives is second-hand: Kubernetes audit records at the API server, container stdout already shipped off the node, runtime-sensor telemetry, flow records. Those prove that events were emitted - an exec was requested, bytes moved - but they never contain the artefact, so you cannot hash it, analyse it, or match it against anything later. The replacement pod is clean, so you lose forward visibility too, while the access path that got them in is untouched. Pin the image digest and the pod object early; those at least preserve what was supposed to be running.

code

json · 16 lines
json
{
  "kind": "Event",
  "apiVersion": "audit.k8s.io/v1",
  "level": "Metadata",
  "stage": "ResponseComplete",
  "verb": "create",
  "user": { "username": "system:serviceaccount:ci:deploy-bot" },
  "sourceIPs": ["10.42.7.19"],
  "objectRef": {
    "resource": "pods", "subresource": "exec",
    "namespace": "payments", "name": "api-7d9c8-fz2kq"
  },
  "requestURI": ".../pods/api-7d9c8-fz2kq/exec?command=sh&container=api&stdin=true&tty=true",
  "responseStatus": { "code": 101 },
  "stageTimestamp": "2026-03-11T02:41:07Z"
}

go deeper

for a junior

Be ready to name what exists only inside a running container - process memory, live sockets, files dropped into the writable layer - and to say plainly that centrally shipped logs are not a copy of the code.

for a middle

Explain why a memory-only loader leaves nothing on disk to collect later, and which durable records survive the pod - Kubernetes audit, node-shipped stdout, runtime and flow telemetry - and what each one actually proves.

for a senior

Show that you act in the first minute: pin the image digest and the pod object, stop whatever is about to recreate the pod, and only then argue about how long a capture may take.

for a principal

Own the consequence at estate scale. If remediation destroys evidence by default, the organisation cannot answer scope questions after any container intrusion; decide whether that is an accepted risk or a platform change you fund.

## The question behind the question An interviewer asking this is checking one thing: do you understand that **containerised remediation is destructive by default**, and can you name precisely what it destroys? In a Kubernetes estate the normal reaction to a flagged workload — delete the pod, let the controller replace it — is also the fastest possible evidence-destruction procedure, and it happens in seconds, often automatically, usually before a human has looked at anything. ## What lives only inside the container A running container is a process (or a small tree of them) in its own namespaces on a node, with a thin writable layer stacked over the read-only image layers. Three categories of material exist nowhere else: - **Process memory.** The address space of the running process holds the decoded payload of a memory-only loader, its configuration, any keys or tokens it pulled in, decrypted buffers, and the arguments of whatever it is doing right now. A loader that fetches and executes in memory never writes the payload to a filesystem, so this is not one copy of the implant — it is the *only* copy. - **The writable layer.** Anything the process dropped: staged archives, scripts, added binaries, temporary output. It is removed when the container instance goes away. - **Live kernel-side state for those namespaces.** Established sockets and their peers, the process tree and parent relationships, open file descriptors, mounts. Together these answer "what was it talking to, and what started it". Delete the pod and all three are gone with no copy anywhere. This is different from a VM or a laptop, where the disk survives a reboot and can be imaged tomorrow — the container's equivalent of "the disk" is discarded as part of normal operation. ## What survives, and what it actually proves The durable record is second-hand, and each piece proves something narrower than people assume: | What survives | What it proves | What it does not | |---|---|---| | Kubernetes audit events | that an API request was made, by which identity, against which resource and subresource | anything about what happened *inside* a container; for a `pods/exec`, the requested command is in the request URI, but the session content is not recorded | | Container stdout shipped off the node | what the process chose to write to stdout | anything it chose not to write, and anything about its memory | | Runtime-sensor / EDR telemetry | that a rule fired and that certain events were observed | that something malicious happened — a detection firing is a detection firing | | Flow records (NetFlow/IPFIX) | that bytes moved between endpoints, how many, when | what those bytes were — flow records carry no payload | | The image in the registry | what was *supposed* to be running | the runtime-injected implant, which by definition was not in the image | That table is the whole answer to "the SIEM has it": the SIEM has records of emissions and observations. It does not have the artefact. You cannot hash what you never copied, cannot pull strings from it, cannot compare it against a sample from another victim, and cannot hand it to anyone who might identify it. ## The ephemeral-workload twist On a Kubernetes node the deadline is not set by you. A liveness probe restart, an autoscaler scale-down, a rollout, a node drain, or an operator's `delete pod` all end the container, and the ones that are automated do not wait for the incident channel. Meanwhile the replacement pod is clean, so you also lose forward visibility: you cannot watch what the intruder does next in that workload, because the workload they were in no longer exists. That is a second loss, and it is the one people forget — the intruder's access path (a stolen credential, an exposed endpoint, a vulnerable dependency in the image) is untouched by the deletion, so they can come back into a pod you are no longer watching. ## What "deleting the pod cut them off" gets wrong Deleting the pod ends that process. It does not end the intrusion. A projected, bound service-account token stops being accepted once the pod object it is bound to is gone, which helps; but a long-lived Secret-based token, a cloud credential lifted from the node's instance metadata, a key exfiltrated minutes earlier, or the vulnerability that granted execution in the first place are all unaffected. Treating deletion as eradication is the classic wrong answer here: containment stops damage, eradication removes the adversary's paths back, and the two are different phases. ## What to do instead, in the first minute You do not need a forensics programme to avoid the worst of this. Before anything destroys the pod: record the pod's identity and object (name, UID, namespace, node, the spec as it stands), pin the **image digest** rather than the tag, and stop whatever is about to recreate it. Those cost seconds and they preserve the ability to reconstruct the baseline and to interpret every other record you collect later. The expensive part — capturing the running process — is a decision with a cost, and it is the one that has to be argued for while the clock runs.

  • The image the pod ran from is still in the registry - doesn't that give you the malware?
    Only if the malicious code was in the image. A memory-only loader is fetched or injected at runtime, so the image is the clean baseline rather than the implant. Pinning it by sha256 digest is still worth doing immediately: it tells you exactly what was supposed to be running, and diffing against it shows what was added. It answers 'was this a supply-chain compromise', not 'what was the payload'.
  • Does deleting the pod at least cut the intruder off?
    Not reliably. It ends that process, not the access path. A projected bound service-account token stops being accepted once the pod object it is bound to is gone, but a long-lived Secret-based token, a cloud credential lifted from instance metadata, or the vulnerability that granted execution are all unaffected. The controller creates a clean replacement, and the same weakness is reachable again seconds later - in a pod nobody is watching.
  • The node still has the container's stdout log files - is that the same as having the container?
    No. Kubelet keeps pod log files on the node only while the pod object exists, and they contain whatever the process chose to write to stdout - nothing about its memory, its child processes, or its sockets. They are a cheap thing to grab early, not a substitute for the container itself.

Scrapping a crashed car before anyone inspects it. The traffic cameras still show it was on the road at 02:00, but the thing that would tell you why it crashed has been melted down.

saying these in an interview costs you the question

  • Says the SIEM already has everything we need
  • Assumes the container image contains the implant
  • Treats the clean replacement pod as the same evidence
  • Thinks deleting the pod ends the intrusion
  • Believes a fired runtime alert proves what executed

context

open as a page

In a breach determination, what is the difference between personal data being exposed and being accessed?

level: juniorimportance: must knowfreq 57%

basics

~20 s

Exposure means the data became reachable because a control failed. Access means someone actually retrieved it. Notification duties turn on access or its reasonable likelihood, so an investigator looks for a retrieval record, not merely an opportunity.

open as a page

What makes a communications channel genuinely out-of-band during a suspected intrusion?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Out-of-band means the channel shares no dependency with the systems under investigation: different transport, different identity provider, different devices, and a member list not drawn from the compromised directory. A second app behind the same sign-on is still in-band.

open as a page

An intrusion ran through an RDP gateway whose logs only ever wrote to local disk: why does closure raise a collection requirement, not a detection request?

level: middleimportance: must knowfreq 62%

basics

~20 s

No rule can match records that never arrived. The miss was at collection, not logic - the gateway's logs stayed on local disk, so the platform's silence was absence of evidence, not evidence of absence. Ship the source first.

open as a page

A platform SRE wants to drain the Kubernetes node now; you want forty minutes to capture - how do you decide?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Ask what draining buys that a non-destructive control cannot. Cordoning, denying egress and revoking the pod's credentials stop the harm without destroying the container; draining ends it. Take the window only if containment holds and a named person authorises the delay.

open as a page

The forensic picture is incomplete — how do you decide a personal-data breach notification duty has arisen?

level: seniorimportance: must knowfreq 51%

basics

~10 s

Decide on a reasonable degree of certainty that personal data was compromised, not on a finished investigation. Establish that promptly and record it: facts you should have established count as facts you had.

open as a page

How do you stand up an out-of-band incident bridge in a SaaS-first company with no on-premises fallback?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Build a channel that depends on the suspect estate for nothing: not identity, hosting, devices or its member list. Use a separate tenant or personal devices, admit participants one at a time, verify them live, and source contacts outside the compromised directory.

open as a page

After a closed intrusion, what is a collection requirement and how does it differ from a detection request?

level: juniorimportance: should knowfreq 48%

basics

~20 s

A collection requirement asks for telemetry the platform does not receive today - a named source, host set, fields and retention. A detection request asks for logic over records you already collect. One buys visibility; the other buys a rule.

open as a page

A suspect Kubernetes pod may reschedule at any moment - what do you capture, and what can wait?

level: middleimportance: should knowfreq 48%

basics

~20 s

Capture only what dies with the container: the suspect process's memory, the container identity and image digest, the writable layer, live sockets and the pod-IP mapping, and the pod object. The image, audit records and shipped telemetry can wait.

open as a page

A proxy shows 40 GB to a sanctioned file-sync host and the provider's audit trail names the files — what does that prove?

level: middleimportance: should knowfreq 44%

basics

~20 s

Together they prove a named account moved specific file objects to that host, and how many bytes left. Neither record carries contents; what personal data was involved comes from the originals still in the source repository.

open as a page

What does a collaboration suite's audit trail prove when a compromised account joined your incident channel?

level: middleimportance: should knowfreq 52%

basics

~20 s

It proves a session authenticated as that identity performed the recorded actions: joining the channel, downloading a file. It does not prove a person read any message, and no download entry proves nothing was seen.

open as a page

An executive wants your incident briefing to name the attacker group. How do you answer?

level: principalimportance: should knowfreq 38%

basics

~20 s

Give the executive the decisions they actually need rather than the name they asked for, and state what the evidence supports and what it does not. Assessments of who is responsible belong in a separate owned product, not a status briefing.

open as a page

You notified on partial facts and the affected-record count later grew tenfold — how do you handle it?

level: seniorimportance: nice to knowfreq 27%

basics

~10 s

Supplement the original filing with the new figure and the analytic step that produced it. Growth explained by a widened search reads as diligence; growth contradicting an earlier assurance does not.

open as a page

How do you tell 4,000 staff to stop using the company chat when the intruder is reading it?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Send a short instruction that changes behaviour, carries no findings, and can be verified without clicking anything, through a path the adversary does not control. Write it assuming the intruder and the press both read it.

open as a page

Your post-intrusion collection requirement is refused on cost by the platform owner - what does the closure document record?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Record the requirement as raised, the refusal, who decided it and when, the exposure that remains in operational terms, any compensating measure you will actually perform, and a date to revisit. Never quietly delete the requirement.

open as a page

Platform engineering auto-deletes flagged pods, so every intrusion loses its evidence - what do you negotiate?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Not a veto on auto-remediation but a quarantine mode: replace the workload immediately, then hold the original with egress denied and credentials revoked for a hard-capped window, reaped automatically, invoked under standing authority by the incident lead.

open as a page