skip to content

A containment playbook isolated a compromised CI runner but its artifact-quarantine step had failed silently for a month — how do you handle the case now?

level: seniorimportance: must knowfreq 50%

answer

  1. green run, failed task
  2. continue-on-error on a containment action
  3. the target's audit trail is the evidence
  4. the failed step names the open path
  5. believed-contained versus actually-contained

basics

~20 s

The run log is a claim, not evidence: rebuild what landed from each target's audit trail. The artifact stayed published, so the path was never closed. Quarantine it manually, re-scope to its consumers, re-date containment.

solid answer

~50 s

The run finished green only because the quarantine task was set to continue on error, so containment was half-applied: the runner was isolated, the artifact was not. I reconstruct reality from the targets — the EDR host record for the isolation, the artifact repository's audit log for the quarantine that never happened — rather than from the playbook status. The failed step defines an open re-entry path, so scope now covers every build and workstation that consumed that artifact while it stayed published, which is how the intruder came back after we called it contained. I quarantine it manually, verify at the repository, and record two timestamps: when we believed we were contained and when we actually were. Then I check every other case run since the connector died a month ago, and change the playbook so a failed containment task blocks the case and pages rather than passing through.

code

json · 18 lines
json
{
  "playbook": "contain-build-compromise",
  "case": "INC-4412",
  "status": "completed",
  "tasks": [
    { "n": 2, "name": "isolate-ci-runner",
      "connector": "edr",  "status": "completed",
      "response": { "http": 200 } },
    { "n": 3, "name": "disable-build-service-account",
      "connector": "idp",  "status": "completed",
      "response": { "http": 200 } },
    { "n": 4, "name": "quarantine-published-artifact",
      "connector": "artifact-repo", "status": "completed",
      "response": { "http": 401, "error": "invalid_token" },
      "on_error": "continue" }
  ]
  ...
}

go deeper

for a junior

Know that a playbook can finish green with a failed step inside it, and that you confirm containment in the target system's console or audit log rather than in the automation.

for a middle

Explain how continue-on-error produces a half-applied containment, and how you diff the playbook's intended actions against each target's audit trail to find what actually landed.

for a senior

Demonstrate the case judgment: the failed step defines an open re-entry path, scope follows that path, containment is re-dated from verified evidence, and every case since the connector died gets reviewed.

for a principal

Own the standard that stops it recurring — containment tasks that halt and page rather than continue, in-playbook verification at the target, and a scheduled exercise budget so a broken response path is found in hours.

## Why the run is green when the containment is not Response platforms let each task declare what happens on failure. For enrichment tasks — a reputation lookup, an asset-owner lookup — continuing past a failure is correct, because the case is still workable without that context. When the same setting is applied to a *containment* task, the platform faithfully does what it was told: it records the failure inside the task, moves to the next one, and finishes the run with an overall status of completed. Nobody reads task-level detail on a run that says completed. In the record above, task 4's connector received HTTP 401 with an `invalid_token` error — the artifact repository rejecting the call as unauthenticated, because the grant the connector holds expired or was revoked a month ago. The isolation landed. The account disable landed. The artifact stayed downloadable. ## Half-applied containment is a specific and dangerous state Containment is not one action; it is a set of actions that together close every path the adversary has. Applying a subset produces something worse than applying none, because it *feels* finished: - the incident lead reports containment complete; - the case moves into eradication and recovery; - everyone stops watching the path that is still open. Here, the still-open path is a poisoned artifact in a repository that downstream jobs and developers keep pulling. The adversary does not need the isolated runner back; every consumer of that artifact is a new foothold, which is exactly how re-entry happened after the case was declared contained. ## Reconstructing what actually landed Work from the targets, in this order: 1. **List the actions the playbook was supposed to perform**, from the playbook definition, not from the run. 2. **For each one, find the target system's own audit entry**: the EDR's host record and its containment timestamp; the identity provider's audit entry for the disabled account; the artifact repository's audit log for a quarantine event that will not be there. 3. **Diff the two lists.** Everything in the first without a match in the second is an action you have not performed, whatever the run log says. 4. **Perform the missing actions by hand now**, and verify each at the target before writing it down. This diff is also the honest answer to "which steps ran?" — the question the playbook's engineer will be asked, and the one the execution record alone cannot answer credibly, because it is the artefact under suspicion. ## Re-scoping Two different windows open here, and both need work. **The case window.** The artifact was published at some point and never quarantined. Scope is therefore every build job, image, and workstation that pulled it between publication and the manual quarantine — obtained from the repository's download or pull logs, not guessed. Anything that consumed it is a candidate host and must be triaged on its own merits. **The estate window.** The connector has been dead for a month. Every case in that month whose playbook included an artifact-repository action is suspect in exactly the same way. That review is unglamorous and it is not optional; a silent connector failure is never a single-case problem. ## The number that has to be written down Record the moment containment was *believed* complete (the green run) and the moment it was *actually* complete (the verified manual quarantine). The gap is a measured duration, and it belongs in the case as a fact, not a footnote: it bounds the adversary's opportunity, it explains the re-entry, and it is the honest core of the correction the incident lead has to issue to everyone who acted on the earlier statement. Eradication here also happens late by definition, so re-entry watching is not optional: keep detections on the artifact's hash, the repository account, and the runner fleet running well past the point the case would normally close. ## Fixing the class, not the instance Repairing the connector's credential fixes today. The class of failure needs three changes: - **A failed containment task must never pass through.** Set containment tasks to halt, mark the case blocked, and page the incident lead. A response platform that cannot perform containment should be loud about it, at exactly the moment it matters. - **Verify at the target inside the playbook.** After each containment action, read the target's state back and fail the task if it does not reflect the change. An acknowledgement is not a result. - **Exercise the path when there is no incident.** A playbook only runs during cases, so the absence of errors proves nothing about the connectors — it may only mean the estate was quiet. A scheduled exercise that quarantines and releases a dummy artifact, against the real repository with the real credential, turns a month of silent breakage into a few hours. ## What the incident lead says The correction is short and factual: what was believed, what is now known, the window between them, and what is being re-scoped as a result. It is far cheaper to say it in the case than to have it discovered by the adversary's second visit — which is what happened here.

  • How would you have caught the dead connector before the incident?
    By exercising it when there is no case. Playbooks only run during incidents, so a month without errors may just mean a quiet month — silence is not health. A scheduled end-to-end exercise against a decoy object, using the real credential and the real API, plus alerting on any connector authentication failure, turns a month of silent breakage into hours.
  • Where does the corrected containment timestamp come from?
    From the target system's own audit entry for the action that actually landed — the repository's quarantine event, the EDR's containment state change. Never from the playbook run, which records only what was sent. If the target offers no audit entry, the timestamp is the moment you personally observed the resulting state.
  • Is continue-on-error ever the right setting in a response playbook?
    Yes, for best-effort context: enrichment lookups, ticket updates, chat notifications. The case is still workable without them. It is never right for a containment or eradication action, where continuing converts a visible failure into a false claim that the estate is safe.
  • The runner is ephemeral and was already destroyed. Does that make the isolation moot?
    It makes the isolation irrelevant, not the case. Scope moves to what the compromised job produced and touched — the published artifact, the credentials the job held, and the downstream consumers. It also means your evidence window is short, so pull the job logs and build metadata before the platform's retention removes them.

saying these in an interview costs you the question

  • Treats the playbook's completed status as proof of containment
  • Re-runs the whole playbook without checking which actions actually landed
  • Dates containment from the run rather than the verified action
  • Fixes the connector and closes the case without re-scoping
  • Treats isolating the runner as having eradicated the intrusion
  • Reviews only this case, not the month the connector was dead

context