skip to content

Incident Response Lifecycle

You declare what triage called adversarial, cut the intruder out without losing the evidence that proves it, come back clean, and answer for it afterwards. Interviewers open on ransomware.

on this pageshow

explore

questions

page 1 of 2

A flagged Kubernetes pod was deleted and replaced before anyone captured it - what is gone for good?

level: juniorimportance: must knowfreq 62%

answer

  1. what only that container held
  2. shipped logs record emissions
  3. the loader never touched disk
  4. no copy means nothing to hash

basics

~20 s

Everything that existed only inside that container: the process memory holding a memory-only loader, its live network and process state, and the writable layer's dropped files. Shipped logs survive, but they record what was emitted, never the code that ran.

solid answer

~50 s

Deleting the pod destroys the only copy of anything that never touched durable storage. A memory-only loader lives in the process address space, so its decoded payload, in-memory configuration, keys and established sockets go with the container, along with whatever it dropped into the writable layer. What survives is second-hand: Kubernetes audit records at the API server, container stdout already shipped off the node, runtime-sensor telemetry, flow records. Those prove that events were emitted - an exec was requested, bytes moved - but they never contain the artefact, so you cannot hash it, analyse it, or match it against anything later. The replacement pod is clean, so you lose forward visibility too, while the access path that got them in is untouched. Pin the image digest and the pod object early; those at least preserve what was supposed to be running.

code

json · 16 lines
json
{
  "kind": "Event",
  "apiVersion": "audit.k8s.io/v1",
  "level": "Metadata",
  "stage": "ResponseComplete",
  "verb": "create",
  "user": { "username": "system:serviceaccount:ci:deploy-bot" },
  "sourceIPs": ["10.42.7.19"],
  "objectRef": {
    "resource": "pods", "subresource": "exec",
    "namespace": "payments", "name": "api-7d9c8-fz2kq"
  },
  "requestURI": ".../pods/api-7d9c8-fz2kq/exec?command=sh&container=api&stdin=true&tty=true",
  "responseStatus": { "code": 101 },
  "stageTimestamp": "2026-03-11T02:41:07Z"
}

go deeper

for a junior

Be ready to name what exists only inside a running container - process memory, live sockets, files dropped into the writable layer - and to say plainly that centrally shipped logs are not a copy of the code.

for a middle

Explain why a memory-only loader leaves nothing on disk to collect later, and which durable records survive the pod - Kubernetes audit, node-shipped stdout, runtime and flow telemetry - and what each one actually proves.

for a senior

Show that you act in the first minute: pin the image digest and the pod object, stop whatever is about to recreate the pod, and only then argue about how long a capture may take.

for a principal

Own the consequence at estate scale. If remediation destroys evidence by default, the organisation cannot answer scope questions after any container intrusion; decide whether that is an accepted risk or a platform change you fund.

## The question behind the question An interviewer asking this is checking one thing: do you understand that **containerised remediation is destructive by default**, and can you name precisely what it destroys? In a Kubernetes estate the normal reaction to a flagged workload — delete the pod, let the controller replace it — is also the fastest possible evidence-destruction procedure, and it happens in seconds, often automatically, usually before a human has looked at anything. ## What lives only inside the container A running container is a process (or a small tree of them) in its own namespaces on a node, with a thin writable layer stacked over the read-only image layers. Three categories of material exist nowhere else: - **Process memory.** The address space of the running process holds the decoded payload of a memory-only loader, its configuration, any keys or tokens it pulled in, decrypted buffers, and the arguments of whatever it is doing right now. A loader that fetches and executes in memory never writes the payload to a filesystem, so this is not one copy of the implant — it is the *only* copy. - **The writable layer.** Anything the process dropped: staged archives, scripts, added binaries, temporary output. It is removed when the container instance goes away. - **Live kernel-side state for those namespaces.** Established sockets and their peers, the process tree and parent relationships, open file descriptors, mounts. Together these answer "what was it talking to, and what started it". Delete the pod and all three are gone with no copy anywhere. This is different from a VM or a laptop, where the disk survives a reboot and can be imaged tomorrow — the container's equivalent of "the disk" is discarded as part of normal operation. ## What survives, and what it actually proves The durable record is second-hand, and each piece proves something narrower than people assume: | What survives | What it proves | What it does not | |---|---|---| | Kubernetes audit events | that an API request was made, by which identity, against which resource and subresource | anything about what happened *inside* a container; for a `pods/exec`, the requested command is in the request URI, but the session content is not recorded | | Container stdout shipped off the node | what the process chose to write to stdout | anything it chose not to write, and anything about its memory | | Runtime-sensor / EDR telemetry | that a rule fired and that certain events were observed | that something malicious happened — a detection firing is a detection firing | | Flow records (NetFlow/IPFIX) | that bytes moved between endpoints, how many, when | what those bytes were — flow records carry no payload | | The image in the registry | what was *supposed* to be running | the runtime-injected implant, which by definition was not in the image | That table is the whole answer to "the SIEM has it": the SIEM has records of emissions and observations. It does not have the artefact. You cannot hash what you never copied, cannot pull strings from it, cannot compare it against a sample from another victim, and cannot hand it to anyone who might identify it. ## The ephemeral-workload twist On a Kubernetes node the deadline is not set by you. A liveness probe restart, an autoscaler scale-down, a rollout, a node drain, or an operator's `delete pod` all end the container, and the ones that are automated do not wait for the incident channel. Meanwhile the replacement pod is clean, so you also lose forward visibility: you cannot watch what the intruder does next in that workload, because the workload they were in no longer exists. That is a second loss, and it is the one people forget — the intruder's access path (a stolen credential, an exposed endpoint, a vulnerable dependency in the image) is untouched by the deletion, so they can come back into a pod you are no longer watching. ## What "deleting the pod cut them off" gets wrong Deleting the pod ends that process. It does not end the intrusion. A projected, bound service-account token stops being accepted once the pod object it is bound to is gone, which helps; but a long-lived Secret-based token, a cloud credential lifted from the node's instance metadata, a key exfiltrated minutes earlier, or the vulnerability that granted execution in the first place are all unaffected. Treating deletion as eradication is the classic wrong answer here: containment stops damage, eradication removes the adversary's paths back, and the two are different phases. ## What to do instead, in the first minute You do not need a forensics programme to avoid the worst of this. Before anything destroys the pod: record the pod's identity and object (name, UID, namespace, node, the spec as it stands), pin the **image digest** rather than the tag, and stop whatever is about to recreate it. Those cost seconds and they preserve the ability to reconstruct the baseline and to interpret every other record you collect later. The expensive part — capturing the running process — is a decision with a cost, and it is the one that has to be argued for while the clock runs.

  • The image the pod ran from is still in the registry - doesn't that give you the malware?
    Only if the malicious code was in the image. A memory-only loader is fetched or injected at runtime, so the image is the clean baseline rather than the implant. Pinning it by sha256 digest is still worth doing immediately: it tells you exactly what was supposed to be running, and diffing against it shows what was added. It answers 'was this a supply-chain compromise', not 'what was the payload'.
  • Does deleting the pod at least cut the intruder off?
    Not reliably. It ends that process, not the access path. A projected bound service-account token stops being accepted once the pod object it is bound to is gone, but a long-lived Secret-based token, a cloud credential lifted from instance metadata, or the vulnerability that granted execution are all unaffected. The controller creates a clean replacement, and the same weakness is reachable again seconds later - in a pod nobody is watching.
  • The node still has the container's stdout log files - is that the same as having the container?
    No. Kubelet keeps pod log files on the node only while the pod object exists, and they contain whatever the process chose to write to stdout - nothing about its memory, its child processes, or its sockets. They are a cheap thing to grab early, not a substitute for the container itself.

Scrapping a crashed car before anyone inspects it. The traffic cameras still show it was on the road at 02:00, but the thing that would tell you why it crashed has been melted down.

saying these in an interview costs you the question

  • Says the SIEM already has everything we need
  • Assumes the container image contains the implant
  • Treats the clean replacement pod as the same evidence
  • Thinks deleting the pod ends the intrusion
  • Believes a fired runtime alert proves what executed

context

open as a page

In a breach determination, what is the difference between personal data being exposed and being accessed?

level: juniorimportance: must knowfreq 57%

basics

~20 s

Exposure means the data became reachable because a control failed. Access means someone actually retrieved it. Notification duties turn on access or its reasonable likelihood, so an investigator looks for a retrieval record, not merely an opportunity.

open as a page

What makes a communications channel genuinely out-of-band during a suspected intrusion?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Out-of-band means the channel shares no dependency with the systems under investigation: different transport, different identity provider, different devices, and a member list not drawn from the compromised directory. A second app behind the same sign-on is still in-band.

open as a page

What does an EDR's network-isolate action on a compromised host stop, and what does it not stop?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Network isolation cuts a host's traffic except the EDR agent's own channel, so an intruder loses interactive access from that machine. It kills no running process, removes no persistence, revokes no stolen credential, and touches nothing on any other host.

open as a page

Why is deleting the implant from the one server that alerted not eradication?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Because eradication targets the intruder's whole foothold, not one file. They may hold persistence on other hosts, valid credentials taken from yours, and the same unfixed way in. Deleting an implant removes an artefact, not their access.

open as a page

After a confirmed intrusion, why is the most recent good backup usually the wrong restore point?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Because a restore point after an intrusion is chosen against the intrusion timeline, not against data freshness. Any copy written after the intruder got in can contain their accounts, implants and configuration changes, so restoring it restores them.

open as a page

Why does resetting a compromised user's password not end an intruder's access?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A password change only stops future logins that present the old password. Already-issued sessions, access and refresh tokens, Kerberos tickets, application consent grants, app passwords, API keys and intruder-enrolled MFA devices are separate objects that must each be revoked.

open as a page

During a live intrusion, what does tipping off the intruder mean, and how do they typically react?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Tipping off means taking an action the intruder can observe and infer detection from. The three classic reactions are destroying evidence, going dormant on a foothold you have not found, and accelerating exfiltration or turning destructive.

open as a page

What is 'patient zero' in an intrusion, and why is the alerting host rarely it?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Patient zero is the first asset in your estate the intruder actually controlled. The host that alerted is only where a rule happened to fire, usually later in the chain, once the intruder moved and grew noisier.

open as a page

Why does a ransomware playbook pre-decide the ransom position and who may disconnect?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Because the extortion note arrives with a countdown, and the decisions it forces — pay or refuse, disconnect or keep producing, call the insurer and counsel — belong to executives who are asleep. Pre-deciding removes hours of paralysis.

open as a page

A compromised read-only auditor account could query a regulated customer table — what must you establish before grading that as data exposure?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Entitlement is not access. The account's grants tell you what was reachable; only the warehouse's query-audit records tell you which tables and columns were actually read. Grade the reach now, call it exposure only where query evidence supports it.

open as a page

When an intrusion is declared, which functions beyond the security team join, and what does each decide?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A declared intrusion adds legal counsel, privacy, HR when the subject is an employee, communications, an executive who can authorise disruptive action, and often a forensic firm retained through counsel. Security keeps the facts; they own the decisions.

open as a page

An intrusion ran through an RDP gateway whose logs only ever wrote to local disk: why does closure raise a collection requirement, not a detection request?

level: middleimportance: must knowfreq 62%

basics

~20 s

No rule can match records that never arrived. The miss was at collection, not logic - the gateway's logs stayed on local disk, so the platform's silence was absence of evidence, not evidence of absence. Ship the source first.

open as a page

A CERT's intrusion tip finds no support in your surviving logs — can you close it as unfounded?

level: middleimportance: must knowfreq 66%

basics

~20 s

No. First separate 'we looked and it is not there' from 'we could not have seen it'. A search is only evidence where a source covered that host and window, retained the data, and would have recorded the described activity.

open as a page

Given VPN concentrator logs and CI runner job logs, how do you date an intrusion's spread?

level: middleimportance: must knowfreq 58%

basics

~20 s

Build one entity table — hosts, accounts, credentials — each with earliest and latest observed adversarial activity and the record proving it. Join the two sources on the concentrator-assigned address inside its session window, and mark every cell observed or inferred.

open as a page

An internet-facing file-transfer appliance's access log shows POSTs to an unknown .aspx page in its web root — is that enough to declare an intrusion?

level: middleimportance: must knowfreq 66%

basics

~20 s

Usually yes. A page in the appliance's web root that appears in no vendor manifest, driven from outside with command-shaped parameters, is an artefact no legitimate process explains — enough to declare, though the log proves no code ran.

open as a page

A platform SRE wants to drain the Kubernetes node now; you want forty minutes to capture - how do you decide?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Ask what draining buys that a non-destructive control cannot. Cordoning, denying egress and revoking the pod's credentials stop the harm without destroying the container; draining ends it. Take the window only if containment holds and a named person authorises the delay.

open as a page

The forensic picture is incomplete — how do you decide a personal-data breach notification duty has arisen?

level: seniorimportance: must knowfreq 51%

basics

~10 s

Decide on a reasonable degree of certainty that personal data was compromised, not on a finished investigation. Establish that promptly and record it: facts you should have established count as facts you had.

open as a page

How do you stand up an out-of-band incident bridge in a SaaS-first company with no on-premises fallback?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Build a channel that depends on the suspect estate for nothing: not identity, hosting, devices or its member list. Use a separate tenant or personal devices, admit participants one at a time, verify them live, and source contacts outside the compromised directory.

open as a page

On a 300-server Linux fleet with confirmed persistence on nine hosts, which do you rebuild rather than clean?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Rebuild wherever the adversary reached root, where the package manager or kernel could have been touched, or where telemetry cannot account for what they did. Clean in place only where the mechanism is fully enumerated and the host's activity is fully observed — and check what the rebuild pulls from first.

open as a page

After an estate-wide password reset, which identity artefacts still let an intruder back in?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Anything that authenticates without a user password: consented application grants, service-principal secrets and certificates, API keys, static cloud keys and in-flight role sessions, intruder-registered authenticators, and a federation signing key. Enumerate from the intrusion timeline.

open as a page

How do you stage one simultaneous containment cut across three regions against an intruder holding two footholds?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Enumerate every foothold first, then run one written cut list — an owner and a tested step per row — inside a single window on one clock. Pre-position access that survives the cut, and capture evidence before any destructive step.

open as a page

Executives want to pay an extortion demand to prevent publication — what must the pre-decided ransom position answer?

level: principalimportance: must knowfreq 57%

basics

~20 s

It must separate buying a decryption tool from buying a promise to delete stolen data, and settle in advance who signs, whether sanctions screening and insurer consent allow payment at all, and that engaging a negotiator is not the same as paying.

open as a page

After a closed intrusion, what is a collection requirement and how does it differ from a detection request?

level: juniorimportance: should knowfreq 48%

basics

~20 s

A collection requirement asks for telemetry the platform does not receive today - a named source, host set, fields and retention. A detection request asks for logic over records you already collect. One buys visibility; the other buys a rule.

open as a page

A national CERT says one of your public IPs attacked another company at 02:11 UTC — what do you establish first?

level: juniorimportance: should knowfreq 54%

basics

~20 s

Turn the report into an asset and a time window. Normalise the timestamp, then use NAT, proxy or DHCP records together with the source port to find which internal host held that public address at that exact moment.

open as a page

What does declaring a security incident commit an organisation to that continuing to investigate does not?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Declaring turns an open investigation into a formal intrusion response: evidence must be preserved rather than remediated away, the cyber insurer normally has to be notified, the work moves under legal counsel, and named people gain authority to disconnect systems.

open as a page

A suspect Kubernetes pod may reschedule at any moment - what do you capture, and what can wait?

level: middleimportance: should knowfreq 48%

basics

~20 s

Capture only what dies with the container: the suspect process's memory, the container identity and image digest, the writable layer, live sockets and the pod-IP mapping, and the pod object. The image, audit records and shipped telemetry can wait.

open as a page

A proxy shows 40 GB to a sanctioned file-sync host and the provider's audit trail names the files — what does that prove?

level: middleimportance: should knowfreq 44%

basics

~20 s

Together they prove a named account moved specific file objects to that host, and how many bytes left. Neither record carries contents; what personal data was involved comes from the originals still in the source repository.

open as a page

What does a collaboration suite's audit trail prove when a compromised account joined your incident channel?

level: middleimportance: should knowfreq 52%

basics

~20 s

It proves a session authenticated as that identity performed the recorded actions: joining the channel, downloading a file. It does not prove a person read any message, and no download entry proves nothing was seen.

open as a page

How do you sweep 300 Linux servers for persistence when a third have no EDR agent?

level: middleimportance: should knowfreq 47%

basics

~20 s

Compare every host against its declared configuration and its package database instead of against a sensor. Planted persistence usually surfaces as drift: an added SSH key, an undeclared systemd timer, a new package hook. The gap is unmanaged paths and hosts whose agent stopped reporting.

open as a page

showing 1–30 of 60