Platform engineering auto-deletes flagged pods, so every intrusion loses its evidence - what do you negotiate?
answer
- the default has a legitimate owner
- preserve and remediate, not either
- hold, don't destroy - with an expiry
- automate the cheap evidence
- fund it with past unprovable scope
basics
~20 sNot a veto on auto-remediation but a quarantine mode: replace the workload immediately, then hold the original with egress denied and credentials revoked for a hard-capped window, reaped automatically, invoked under standing authority by the incident lead.
solid answer
~50 sAsking a platform team to stop deleting flagged pods will fail, and it should - deletion restores service in seconds and they carry the availability pager. Negotiate the shape instead. Remediation still creates the replacement immediately; the original pod is quarantined rather than destroyed, kept on a cordoned node with a deny-all egress policy and its identity revoked, for a platform-enforced expiry after which it is reaped automatically, with a cap on concurrent holds. Then shrink the human decision: make the cheap evidence automatic on every flagged pod - pod object, runtime metadata, image digest, container filesystem - so the only thing anyone has to argue about at 03:00 is holding a live process. Pre-authorise that within a defined budget so nobody negotiates capacity with a sleeping executive. Fund it with cases, not principles: the intrusions where scope was never provable because the artefact was destroyed.
go deeper
Understand that remediation and evidence pull against each other, and that the team deleting the pod is doing exactly what they were asked to do for availability.
Be able to describe a concrete alternative - replace the workload, hold the original with egress denied, reap it on a timer - rather than arguing for a veto over someone else's automation.
Show that you would cost the proposal in capacity and pager load, automate the cheap evidence so the human decision shrinks, and pre-authorise the call so nothing is negotiated at 03:00.
Own a negotiation with a team that can refuse: fund it with cases where scope was unprovable, accept documented exemptions for workloads that cannot bear a hold, and measure whether artefact recovery actually improved.
## Why the default exists and why "stop doing it" loses Auto-remediation that deletes a flagged pod is a good control. It restores a healthy workload in seconds, it needs no human, and it reduces pages for the team that owns the platform. That team is measured on availability and on toil. A security ask that reads "stop deleting compromised pods" asks them to trade their own metric for someone else's, at 03:00, on the word of a detection that might be a false positive. They can refuse, and if they cannot refuse openly they will route around it later. Treat this as a negotiation with an owner who has a legitimate position, not as a policy you can assert. ## Reframe: preserve *and* remediate, don't choose The proposal that survives contact is one that keeps everything the platform team gets today: 1. **Remediation still happens immediately.** The controller creates the replacement pod; service is restored on the same timeline as now. 2. **The original is quarantined, not destroyed.** It stays on its node, the node is cordoned so nothing new lands there, a deny-all egress policy is applied, and the identity it holds is revoked. The container keeps running, isolated and stripped, so it remains capturable. 3. **A hard, platform-enforced expiry.** After a fixed window the held pod is reaped automatically unless someone re-authorises. The guard rail must be enforced by the system, not by a human promise, because the failure mode is a forgotten hold occupying capacity for a month. 4. **A capacity cap.** A maximum number of concurrent holds, visible on the platform team's own dashboard, with who invoked each and why. Now the conversation is about a bounded, self-releasing resource cost rather than about disabling their control. ## Automate the cheap evidence so the human decision shrinks The decision that fails at 03:00 is the one that requires an expert to be awake and to win an argument. So move everything cheap into the remediation path itself: on deletion, emit the pod object, the container runtime metadata, the image reference **and digest**, the restart history and the container filesystem to an evidence bucket, automatically, for every flagged pod. None of that is contentious, none of it delays remediation meaningfully, and it removes most of the routine loss. What remains — holding a live process so its memory can be captured — is expensive and genuinely needs a decision, so make sure that is the *only* thing anyone has to argue about. ## Pre-authorise, with a budget Standing authority beats escalation. Write it down: the incident lead may quarantine up to N pods or one node's worth of capacity, for up to the enforced window, without further approval; beyond that it escalates to the service owner. Name who may invoke it out of hours. The point is that nobody negotiates capacity with a sleeping executive while an intruder is live, and equally that the hold cannot quietly grow into an outage nobody agreed to. ## Fund it with cases, not with principles "Best practice" is the weakest argument available. Bring the record: the container intrusions in the last year where remediation destroyed the implant, and the specific consequence in each — the report claim that had to be softened from established to consistent-with, the scope question a customer asked that could only be answered with inference, the detection you could not write because you never had the artefact. Engineering leadership funds a capability that closes a named gap; it does not fund forensic tidiness. ## Accept exemptions rather than pretend Some workloads cannot bear it — a capacity-critical path, a regulated environment where a held pod is itself a problem, a node type too scarce to hold. A documented exemption with an agreed alternative (automatic filesystem and metadata capture only, no live hold) is worth more than a universal policy that is quietly ignored. The exemption list is also the honest statement of residual risk, and it belongs where a risk owner signs it. ## Measure whether it worked Two numbers make this reviewable: the proportion of container intrusions where the running artefact was actually recovered, and the time from detection to quarantine. If recovery does not improve, the mechanism exists but is not reachable in practice — usually because invoking it needs a person who is not there. If holds routinely hit the expiry with nobody having captured anything, the capture capability is the gap, not the policy. Either way, report it against the cases you used to fund it. ## The failure modes to name in an interview Demanding a veto over another team's automation; designing a hold with no expiry; making security the only party who can release it; assuming access to the cluster is the same as authority to change how it behaves; and building a beautiful mechanism whose invocation depends on a specialist being awake.
- The platform team offers a node disk snapshot instead of holding the pod. Is that good enough?Better than nothing, and weaker than it sounds. A disk snapshot may preserve the container filesystem, but the process memory that holds a memory-only loader is not on disk, and untangling which container owned what on a busy shared node costs hours. Take it as a fallback for the routine case, and keep asking for the live hold in the cases where the running process is the evidence.
- How do you stop an evidence hold being abused or simply forgotten?Give it an expiry enforced by the platform rather than by a human promise: the held pod is reaped when the window ends unless someone re-authorises. Cap concurrent holds, surface them on the platform team's own dashboard, and record who invoked each and why. A control that only the security team can release is one the owning team will eventually route around.
- Leadership asks why this deserves engineering effort. What is the argument?Frame it as questions you currently cannot answer. In each recent container intrusion the implant was destroyed by remediation, so scope rested on inference and the report said 'consistent with' where it should have said 'established'. The ask is not forensic tidiness - it is the ability to state what an intruder did and did not reach, which is what customers, insurers and the board actually ask for.
saying these in an interview costs you the question
- Demands the platform team disable auto-remediation
- Assumes security can mandate changes to another team's system
- Designs an evidence hold with no expiry or capacity cap
- Leaves the 03:00 decision to a heroic manual capture
- Argues from best practice instead of past unprovable cases