skip to content

A platform SRE wants to drain the Kubernetes node now; you want forty minutes to capture - how do you decide?

level: seniorimportance: must knowfreq 55%

answer

  1. drain does two jobs, one is urgent
  2. cordon is not containment
  3. contain without destroying first
  4. what claim does the capture buy
  5. the lead authorises, the log records

basics

~20 s

Ask what draining buys that a non-destructive control cannot. Cordoning, denying egress and revoking the pod's credentials stop the harm without destroying the container; draining ends it. Take the window only if containment holds and a named person authorises the delay.

solid answer

~50 s

Draining does two jobs - stopping harm and removing the workload - and only the first is urgent, so separate them. Cordon marks the node unschedulable and leaves running pods alone; a deny-all egress policy plus revoking the pod's service-account and cloud credentials cuts the operator's reach while the process stays alive. That reframes the ask from 'forty minutes of a live intrusion' to 'forty minutes of a contained, credential-stripped pod'. Then weigh what capture buys - the only copy of the loader, and the basis for a scope claim - against the cost: an operator who may notice the isolation and burn their access, other workloads on the node, customer impact. If data is moving now and you cannot cut it, speed wins and the report says 'consistent with' rather than 'we established'. The incident lead owns the call, and the reasoning goes in the log.

go deeper

for a junior

Know the difference between cordoning a node and draining it, and that draining ends the pod along with everything held only in its memory.

for a middle

Explain the non-destructive levers - deny-all egress, credential revocation, cordon - and why they change the question from how risky the wait is to how contained the pod is.

for a senior

Demonstrate the judgment: name what the capture would let you prove, what the delay costs with an operator who may be watching, and the conditions under which you concede and let them drain.

for a principal

Own the authorisation. Decide who may hold production capacity for evidence, within what budget and for how long, and how the decision is recorded so it survives a review weeks later.

## The binary is usually false "Drain now" versus "give me forty minutes" is presented as a choice between service safety and evidence. It rarely is, because draining is doing two jobs at once — **stopping the harm** and **removing the workload** — and only the first is urgent. The senior move is to separate them: take every control that stops harm without destroying the container, then argue about the remaining risk of the pod merely existing. ## What each lever actually does - **Cordon** marks the node unschedulable. Nothing new lands there. Pods already running, including your suspect, are untouched. It is safe, instant, and it is *not* containment — a candidate who cordons and declares the pod contained has misunderstood the verb. - **Drain** evicts the pods. The suspect container ends; its process memory and writable layer are gone with no copy anywhere. - **Delete the pod** — same destruction, plus the controller creates a clean replacement, so you also lose forward visibility while the access path survives. - **Deny-all egress network policy** cuts the pod's outbound reach. The process stays alive and capturable. It is the highest-value non-destructive control you have. - **Credential revocation** — the pod's service account, any cloud role it assumed, keys it held. This shrinks what the operator can still do with what they already took, which drain does not do at all. - **Egress-gateway or firewall block** of the known peers, for the case where you want the pod's own configuration untouched. Run through that list before conceding the pod. In most cases it turns "forty minutes of an active intrusion" into "forty minutes of a contained, credential-stripped workload on a cordoned node", which is a materially different ask and one a platform engineer can accept. ## What you are buying with the window Be explicit, because "for forensics" is not an argument. The capture buys: the only copy of a memory-only loader and therefore the ability to say *what* ran rather than that *something* ran; the configuration that names the operator's infrastructure, which becomes detections and a hunt across the rest of the estate; and the evidence that supports a scope claim — whether this workload's identity reached other namespaces, other accounts, other data. Without it, scope rests on inference from records of emissions. ## What the wait costs - **An operator who can still act.** A hands-on intruder is not a script. They may notice the isolation and burn what they have, destroy artefacts, or fall back to a second foothold you have not found. Isolation is loud. - **Frozen understanding.** The moment you cut egress, you stop learning what they would have done next. Containment costs visibility as well as buying safety. - **Blast radius.** Other workloads share the node; the namespace may share credentials; the data the pod can still read is data it can still read. - **Service and people.** The node held out of rotation, the platform on-call carrying it, the customer impact if capacity is tight. ## Where the line falls Concede and let them drain when the harm is active and unbounded — data is moving now and you cannot cut it, or the workload is spreading and non-destructive controls are not holding — or when the capture cannot realistically complete before an automated process ends the pod anyway. Take the window when containment demonstrably holds, the cost is bounded and understood, and the claim you would otherwise lose actually matters to the investigation. Middle grounds are real: take a shorter window and accept a partial capture, or grab the cheap perishable metadata and image digest and let the pod go. ## When it goes wrong, say so The realistic outcome is often partial: the pod reschedules mid-capture, and you are left with a truncated image of a process that had no single consistent point in time to begin with. That does not make it worthless — fragments can still support a finding — but it changes the language. A report that says "we established the loader was X" on the strength of a truncated capture will not survive challenge; "consistent with X, on a partial acquisition truncated at 02:58 when the pod rescheduled" will, and it tells the reader exactly how much weight to put on it. Softening the claim is the honest consequence of having chosen speed, and choosing speed can still have been right. ## Who decides, and what is written down This is the incident lead's call, not the analyst's and not the platform team's unilaterally. What goes in the log: the time, the window asked for, what containment is in place, what the alternative was, who agreed on the platform side, and the service impact accepted. Six weeks later someone will ask why a compromised pod was left running for forty minutes; the record is the entire defence, and an undocumented good decision looks identical to a reckless one.

  • The pod rescheduled halfway through your memory capture. What can you still say?
    Less than you hoped, and you have to say so. A capture of a live process has no single point in time even when it completes, and a truncated one may hold only fragments. Report what the fragments support - 'consistent with a memory-resident loader matching X' - state that the acquisition was truncated when the pod rescheduled, and never round that up to an established fact. The limit has to travel with the claim.
  • Doesn't cutting the pod's egress tip the operator off?
    It can, and that is a genuine cost. A hands-on operator whose channel dies may abandon quietly, destroy what they can reach, or fall back to a second foothold you have not found. Weigh it against ongoing exfiltration: if the traffic is the bigger risk, cut it. If the tip-off is, a narrower block that stops bulk data movement while leaving the control channel up can buy quiet time.
  • Who signs off holding the node, and what do you record?
    The incident lead authorises - not the analyst doing the capture, and not the platform team unilaterally. Record the decision time, the window requested, the containment in place, who agreed on the platform side, and the service impact accepted. When someone asks in six weeks why a compromised pod ran for another forty minutes, that record is the entire defence; an undocumented good decision looks identical to a reckless one.

saying these in an interview costs you the question

  • Treats drain and capture as the only two options
  • Cordons the node and calls the pod contained
  • Leaves a live pod uncontained while capturing
  • Insists evidence always outranks availability
  • Makes the call alone and records nothing

context