skip to content

An atomic test created a scheduled task on a production laptop and left it there — what did the tool not do for you?

level: seniorimportance: should knowfreq 42%

answer

  1. green in a lab, residue on a real host
  2. cleanup is a separate step
  3. no snapshot to roll back
  4. track, clean, verify, log
  5. tool ran it; you own the consequences

basics

~20 s

It did not clean up after itself. Atomic Red Team runs the technique but only reverts it if you explicitly run the cleanup command. A test that passed in a lab leaves real artefacts — a scheduled task, a registry key, a dropped file — on a production host, plus no chaining, no decisions and no evidence capture. Those are all yours.

solid answer

~60 s

The tool executed the technique; owning the consequences is on you. Atomic Red Team's cleanup is a separate `cleanup_command` that runs only if you invoke it — run the executor and skip cleanup and the scheduled task persists on the host exactly as it would after a real attack. On a lab image you reset the snapshot and never notice; on a production developer laptop there is no snapshot, so the residue stays and can trip your own detections later or confuse the next responder. This is the general shape of "what each tool class does not do for you": Atomic Red Team does not clean up unless told, does not chain techniques, does not decide anything, and does not capture the evidence of what it did. An operator using Cobalt Strike or Sliver has to plan cleanup and record their own actions too. The discipline is to track every artefact you create, run the cleanup, verify it is gone, and keep your own log so the exercise is reversible and defensible.

code

yaml · 11 lines
yaml
- name: Scheduled Task Startup Script
  supported_platforms:
    - windows
  executor:
    name: command_prompt
    elevation_required: true
    command: >
      schtasks /create /tn "AtomicTask" /tr calc.exe /sc onlogon
    cleanup_command: >
      schtasks /delete /tn "AtomicTask" /f
  # run the executor and skip cleanup_command -> the task stays on the host

go deeper

for a junior

Know that Atomic Red Team has a separate cleanup command and it only runs if you invoke it — a passed test can leave the artefact behind.

for a middle

Explain why lab snapshots hide the residue problem and why cleanup, chaining and decisions are not things the library does for you.

for a senior

Show the production discipline: track artefacts, run and verify cleanup, keep your own timeline, and prefer reversible tests on real endpoints so an exercise never degrades the estate.

for a principal

Own the guardrails for running emulation against production — reversibility requirements, evidence-capture standards, and when a test must move to a disposable host — so the program is safe and defensible at scale.

## The failure in the question An atomic for `T1053.005` (Scheduled Task) runs `schtasks /create ...` and succeeds. In a lab you were testing against a **VM snapshot**, so you rolled the machine back afterwards and the task vanished with everything else. On a **production developer laptop** there is no snapshot to roll back to. Unless you ran the atomic's **cleanup_command** (`schtasks /delete ...`), the task is still there — a persistence artefact you planted and forgot. **Cleanup in Atomic Red Team is not automatic.** Each atomic can define a `cleanup_command`, but the executor and the cleanup are separate steps. Running the test does not run the cleanup; you (or `Invoke-AtomicRedTeam` with the cleanup flag) must do it explicitly. This is by design — sometimes you want the artefact to remain so a detection can find it — but it means residue is the default, not the exception. ## The general lesson: what the tool class does NOT do The scheduled task is one instance of a broader truth about emulation tooling — every class leaves work to a human: - **Cleanup / reversibility.** Atomic Red Team gives you a cleanup command but will not run it for you. Operator C2 (Cobalt Strike, Sliver) gives you nothing automatic — every file dropped, service created, or token stolen is yours to reverse. Emulation on production hosts must be reversible or you are degrading the estate you are supposed to be testing. - **Chaining.** A library runs one technique; it will not walk from persistence to credential access to exfiltration. You sequence that yourself, or move to a framework. - **Decision-making.** The tool does what you told it. It does not read the host and decide the next move; a planner follows rules and an operator supplies judgment. - **Evidence capture.** This is the one people miss most. The tool does not, by default, keep a defensible record of *what it did and when* in a form the blue team can reconcile against their telemetry. The operator's own timeline — the framework's task/beacon history, timestamps, the exact commands — is the ground truth that lets you later prove what was actually sent. If you do not capture it, a missed detection is unattributable: you cannot tell a real telemetry gap from "the test never actually ran the way you think." ## Why this matters on production hosts specifically The budgeted environment here is developer laptops, mixed macOS and Windows — real endpoints, not lab images. On real hosts residue has real costs: - it can **fire your own detections** days later and burn an analyst's time chasing a self-inflicted alert; - it can be **mistaken for a real compromise** by the next responder who was not on the exercise; - a **persistence artefact left behind** (a scheduled task, a launch agent on macOS, a run key) is exactly the kind of thing a later, real intruder could find and reuse; - if you never cleaned up and never logged, you cannot even prove the artefact is yours. ## The disciplined operator's checklist - **Track every artefact you create** as you create it — task names, files, keys, accounts. - **Run cleanup and verify it**, do not assume the command worked; confirm the task is gone. - **Keep your own timeline** of actions and timestamps so the exercise is reconcilable and defensible. - **Prefer reversible tests on production**, and where a test cannot be cleanly reverted, run it on a disposable host instead. ## The interview signal The weak answer treats a green test result as success. The strong answer separates *the technique ran* from *the exercise is complete and safe*: the tool proved a behaviour executed, but cleanup, sequencing, judgment and evidence are the human's job, and skipping them on a production host leaves persistence behind and makes the result unattributable.

  • Why does the residue problem barely show up in a lab but bite on a production laptop?
    In a lab you test against a VM snapshot and roll it back afterwards, so any artefact — the task, a dropped file, a registry key — disappears with the revert and you never run the cleanup command consciously. A production developer laptop has no snapshot, so the same skipped cleanup leaves a real, persistent artefact on a machine people use. The lab hides the discipline gap that production exposes.
  • Why is keeping your own operator log part of not-leaving-residue?
    Because an uncaptured action is both irreversible and unattributable. If you did not record that you created a task at a given time, you cannot reliably clean it up later, and the blue team cannot tell your artefact from a real one — nor tell a genuine detection gap from a test that misfired. The operator's timeline is the ground truth that makes the exercise both cleanable and defensible.

saying these in an interview costs you the question

  • Assuming cleanup runs automatically after the test
  • Treating a green lab result as a completed exercise
  • Running irreversible tests on production endpoints
  • Not tracking artefacts created during the run
  • Leaving persistence behind that a real intruder could reuse

context