skip to content

Adversary Emulation

Running a known technique against your own estate on purpose, then reading what the defence did: telemetry, rule, triage, response. Interviewers use it to tell asserted coverage from proven coverage.

on this pageshow

explore

questions

page 1 of 2

In a purple-team session, why re-run the technique after the detection engineer edits the rule?

level: juniorimportance: must knowfreq 62%

answer

  1. an edit is a hypothesis, not a result
  2. the edit changes matching, not history
  3. delivery, parsing, schedule, suppression, routing
  4. green once, by one person, one way

basics

~10 s

Because an edited rule is only a hypothesis until the behaviour it targets is performed again while it is live. Re-execution proves the whole path, from event delivery to a routed alert, actually works.

solid answer

~50 s

The loop is execute, observe, tune, retest. Editing a rule changes what will be matched from now on; it does not change what was already recorded, and it proves nothing about the alerting path. Between a control-plane audit record existing and an analyst seeing an alert there are several independent stages: delivery to the platform, parsing into the fields the rule names, matching when the rule next runs, suppression or deduplication, and routing to a queue. A hand-written test event exercises the syntax and the parser only. Re-executing the technique exercises all of it at once, with a real record produced by real behaviour. That is why the retest is the evidence step of the session, and why a session that ends with an edit and no re-execution has produced a claim rather than a result.

go deeper

for a junior

Be ready to name the four steps of the loop in order and say plainly what the re-execution adds: it is the only step that proves the alert actually arrives, not just that the query looks right.

for a middle

Expect to be asked what sits between the record and the alert - delivery, parsing, schedule and lookback, suppression, routing - and which of those a synthetic test event never touches.

for a senior

Demonstrate that you separate a telemetry gap from a detection gap on the first pass, and that you record a green as evidence about one variant rather than as coverage.

for a principal

Own the framing that a session's product is evidence, not shipped rules, and that the value of both teams sitting together is the length of the feedback loop, not the volume of findings.

## The loop A purple-team session runs a short cycle with both teams in the room: the operator **executes** one technique against a real estate, the detection engineer **observes** what the telemetry actually recorded, the engineer **tunes** the query or rule, and the operator **re-executes** the same technique so the edited logic meets live behaviour. The cycle repeats until the rule fires on a fresh execution, or until the team concludes the behaviour is not detectable with the telemetry present, which is itself a legitimate result. ## What each pass answers The first execution answers a question that is not about detection at all: *is the behaviour recorded?* If the operator mints a second access key for an existing cloud identity and then attaches an inline policy to it (ATT&CK `T1098.001`, additional cloud credentials), the control plane either writes audit records naming the caller, the target identity and the API action, or it does not. If no record exists, no rule can be written, and what you have found is a **telemetry gap**, not a detection gap. Those route to different owners and different fixes. Every later pass answers a different question: *does the candidate logic match a real occurrence of the behaviour, delivered through the real pipeline, on the real schedule?* ## Why an edit is not evidence An edit changes matching, not history. Two things follow. First, the records from the previous execution are already written; re-running a saved search over them tells you the query would have selected those rows. That is worth doing - it is the fast inner check - but it validates a query, not a detection. A detection is a rule plus a schedule plus a suppression policy plus a destination. Second, the failure modes that bite in production live outside the query. A rule can be perfectly correct as a search and still never alert because a field it names is populated on only some events, because its scheduled lookback is shorter than delivery lag so late-arriving records fall outside every run, because a broad suppression swallows it, or because it routes to a queue nobody works. Only a live re-execution walks that whole path end to end. ## What the first green does and does not prove A green retest proves one thing precisely: *this behaviour, performed this way, by this person, with this tooling, produced an alert.* That is a real result and it is not nothing. It is also not coverage. The same technique performed from a console rather than a command-line client, with the two API calls in the opposite order, or with a delay between them, is a different execution and an untested one. The exit criterion for the loop is not the first green; it is a green on a variant the engineer did not watch being built. ## What to record when it goes green Keep a session log that lets someone else reconstruct the claim months later: the exact execution time, the identity and account used, the technique and variant, the rule version that fired, the observed alert time, and the delta between them. Note explicitly that it is one variant by one operator. That single sentence is what stops a session result being read as a coverage claim. ## Why this is asked in interviews Candidates who have only read about purple teaming describe it as a report exchange: red hands over findings, blue writes rules later. The whole value of running it in one sitting is the tightness of the feedback - the operator is still there, the estate is still in the state that produced the record, and a broken assumption is discovered in minutes rather than in the next quarter's exercise. Someone who cannot say why the re-execution matters has not sat in one.

  • The engineer re-runs the saved search over the last hour instead of asking for a re-execution. Is that the same thing?
    No. It checks the query against records already stored, which is a useful fast inner check, but it exercises none of the live path: the schedule, the lookback window, suppression and routing all go untested. It also re-uses the same single execution, so it cannot reveal that the logic is leaning on something specific to that one run.
  • What do you write down when the retest goes green?
    The execution time, the identity and account, the technique and the exact variant performed, the rule version that fired, the alert time and the delay between the two. Add one sentence stating that this is one variant by one operator, so the result is never later read as a coverage claim.
  • The first execution produced no audit record at all. What have you found?
    A telemetry gap, not a detection gap. No rule can be written over a record that does not exist, so the finding is about collection - a source not enabled, an account not covered, an event class not captured - and it belongs to whoever owns that pipeline rather than to the detection engineer.

Editing the rule is writing down a prediction. Re-executing the technique is running the experiment. A session that stops after the edit has only ever made predictions.

saying these in an interview costs you the question

  • Says the rule is proven because it matched a hand-written test event
  • Treats the first green retest as the end of the loop
  • Thinks editing a rule changes what was already recorded
  • Calls a backfill search over old records a retest
  • Describes purple teaming as red handing a report to blue

context

open as a page

Why doesn't replaying saved log records prove a detection still works after a sensor upgrade?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Replaying saved records tests only the stages after collection: parsing and rule logic. It cannot show that the upgraded sensor still emits that event with the same fields, or that the forwarder still ships it. Only re-executing the real technique exercises the whole path.

open as a page

In a red-team engagement, what is a deconfliction contact and what question do they exist to answer?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A named, reachable person on each side of the exercise whose job is to answer one question fast: is this specific activity ours? They attribute or disown observed behaviour. They do not authorise it and they do not order a stand-down.

open as a page

A vulnerability scan, a pentest and a red team engagement all test security — what does each actually prove?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A scan proves a weakness is present. A pentest proves an operator could exploit it and how far the chain reaches. A red team proves whether your defenders detect and stop a realistic path to an objective.

open as a page

Why derive an emulation technique set from a threat profile rather than a popularity list?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A popularity list tests the techniques other organisations happen to report; a threat profile tests the ones the adversary interested in you actually uses. Only the second supports a claim about the intrusion you are likely to face.

open as a page

In an adversary emulation run, does a prevented technique count as detected?

level: juniorimportance: must knowfreq 70%

basics

~10 s

No. A prevention proves the control convicted the behaviour, not that the SOC saw it. Record prevented, detected-only and undetected as three separate outcomes, because each one needs a different fix.

open as a page

What is Atomic Red Team, and what does running a single atomic test actually do?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Atomic Red Team is an open-source library of small, per-technique tests mapped to MITRE ATT&CK. Running one atomic executes a short scripted action for a single technique, so you can check whether your telemetry recorded it and your detections fired.

open as a page

A purple-team technique you executed produced no alert — at which stages could the miss have occurred?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A miss can sit at four stages: the record was never generated, no rule matched the record, no alert reached the queue, or nobody acted on the alert. The absence of an alert names none of them.

open as a page

An emulation proved an exec into a privileged pod went unseen — do you detect it or prevent it?

level: middleimportance: must knowfreq 58%

basics

~20 s

Prevent, primarily. A pod that runs privileged and mounts the node filesystem has almost no legitimate use, so admission control can make the technique impossible, while a detection only tells you it already happened. Detect the narrow exception paths the policy has to allow.

open as a page

When do you choose atomic single-technique tests over one full-chain emulation campaign?

level: middleimportance: must knowfreq 57%

basics

~20 s

Choose atomic tests when the question is which behaviours you can see, because each result is isolated and diagnosable. Choose a chained campaign when the question is whether people and process turn a realistic sequence into a verdict.

open as a page

When do you reach for Atomic Red Team, CALDERA, or an operator C2 like Cobalt Strike or Sliver?

level: middleimportance: must knowfreq 50%

basics

~20 s

Use Atomic Red Team to validate detection of one technique at a time; use CALDERA to chain techniques automatically through agents and a planner; use Cobalt Strike or Sliver when you need a live operator, an interactive C2 channel, and human decision-making the automated tools cannot supply.

open as a page

How do you prove an emulated technique's miss was absent auditd telemetry, not an unmatched rule?

level: middleimportance: must knowfreq 55%

basics

~20 s

Search the host's own audit log over the known window, against a positive-control host that does produce the record. No record on the host is a collection gap; a record present clears collection and moves the test to the rule.

open as a page

A rule tuned live in a purple-team session goes green - how do you prove it detects the technique, not the operator?

level: seniorimportance: must knowfreq 51%

basics

~20 s

Separate what is invariant about the technique from what is incidental to the person: user agent, source address, parameter spelling and call ordering are the operator. Then have a second person execute it differently, and re-open if it misses.

open as a page

A caller says the 02:00 credential-theft alert on your domain controller is red-team activity. How do you verify it?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Treat the claim as unverified until a callback on a pre-recorded number produces a per-activity attribution: this host, this command, this time, this source. Keep containment and evidence collection running while you verify, and reopen anything the operators cannot account for.

open as a page

A purple-team test proved a technique went undetected: what are the routing options for that gap?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Three: detect it, meaning a rule that fires next time; prevent it, meaning a configuration change so the technique stops working; or accept it, recorded with a reason and a named person. Leaving it in the report is not one of them.

open as a page

Cloud control-plane audit events arrive minutes after the API call - how does that shape a live tune-and-retest session?

level: middleimportance: should knowfreq 44%

basics

~20 s

Delivery lag sets the floor tempo: you cannot retest faster than records arrive. Measure the lag at the start, check the raw record before declaring a miss, and pipeline executions instead of one edit per pass.

open as a page

Which platform changes should trigger re-executing a detection test out of cadence?

level: middleimportance: should knowfreq 46%

basics

~20 s

Any change to something a detection depends on: the source's audit configuration or firmware, the agent or forwarder version, the network path to the collector, and the platform's ingest or normalisation schema. Re-run the specific tests that depend on the changed component, not the whole suite.

open as a page

What must a signed red-team authorisation letter contain for a SOC analyst to act on it at 02:00?

level: middleimportance: should knowfreq 55%

basics

~20 s

Two halves: the legal core — a signatory with authority, in-scope and explicitly out-of-scope systems, a dated window, prohibited actions — and fields a defender can check an alert against: operator source addresses, test markers, out-of-band deconfliction numbers.

open as a page

What is an assumed-breach start, and which detections can it never test compared with earned initial access?

level: middleimportance: should knowfreq 58%

basics

~20 s

An assumed-breach engagement begins with the operator already holding a foothold or valid credentials, so everything up to initial access is skipped. The edge and phishing detections are never exercised, and their silence proves nothing about them.

open as a page

A threat report says the crew phishes OAuth consent for mail-read scope. What turns that into an executable test?

level: middleimportance: should knowfreq 44%

basics

~20 s

A named behaviour is a class, not a test. You must add the procedure: which application, which permission scope, which consent path, which target identity, plus the observable you expect and the criterion that decides pass or fail.

open as a page

Your operator's log-clear was blocked and no Windows Security 1102 exists — what may you conclude?

level: middleimportance: should knowfreq 52%

basics

~10 s

Only that the clear never completed. Windows Security 1102 is written when the audit log is cleared, so its absence is indistinguishable from nobody trying. The prevention event is your only positive evidence.

open as a page

Your emulation gap register has forty open rows and no closures in two quarters — what do you do?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Diagnose why rows do not close before adding any more. Most stalls are one of three: no real owner, an owner who never agreed, or no stated closure test. Then stop intake, re-route what the owner cannot deliver, and explicitly accept the tail.

open as a page

With capacity to re-execute 20 of 200 passing attack techniques a quarter, how do you choose?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Rank by consequence of a single success, fragility of the telemetry path, and exposure to change since the last pass. Give a small permanent core the estate cannot survive missing, spend the rest on recently changed and single-source detections, and rotate the tail so nothing is untested indefinitely.

open as a page

An unannounced red team met its objective and no alert ever fired — what does that report actually prove?

level: seniorimportance: should knowfreq 46%

basics

~20 s

It proves one path to the objective existed and went unnoticed by the rules and people on duty that week. It does not say which stage failed, and it measures nothing the operator never attempted.

open as a page

Tamper protection blocked your emulation technique — how do you learn whether a detection existed behind it?

level: seniorimportance: should knowfreq 46%

basics

~10 s

The block truncated the behaviour, so nothing downstream could fire. Check what telemetry still arrived, then re-run in a scoped, time-boxed detect-only window or against a variant the control does not convict.

open as a page

An atomic test created a scheduled task on a production laptop and left it there — what did the tool not do for you?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It did not clean up after itself. Atomic Red Team runs the technique but only reverts it if you explicitly run the cleanup command. A test that passed in a lab leaves real artefacts — a scheduled task, a registry key, a dropped file — on a production host, plus no chaining, no decisions and no evidence capture. Those are all yours.

open as a page

Your Cobalt Strike beacon triggered an EDR alert on its default profile — can you claim the technique is detected?

level: seniorimportance: should knowfreq 48%

basics

~20 s

No. If the alert fired on Cobalt Strike's default artefacts — its stock malleable profile, default named pipe names or default certificate — you detected the tool's defaults, not the technique. Change the profile and the same behaviour would sail through. You proved the SOC catches an out-of-the-box beacon, not the tradecraft.

open as a page

An emulated technique's auditd record exists and the rule matches it on replay, yet no alert reached the queue — where did it die?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Between the rule matching and a human seeing it: the rule was disabled or in test mode, a scheduled search ran before the data was indexed, suppression swallowed the firing, or it was routed below the severity anyone reads.

open as a page

Mid-session, your new key-creation rule also matches the CI pipeline's own key rotation - do you keep tuning in the room?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

No. First confirm from the audit records that the extra matches really are the pipeline, not a second unknown actor. Then stop tuning: an owner outside the room is involved, so the gap leaves the session still open.

open as a page

A re-executed attack technique that alerted last quarter now produces nothing. What do you do?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Confirm the technique really executed, then treat it as a regression rather than a new gap: find what changed, scope every other detection sharing that dependency, and account for the blind window between the change and this discovery. It is a change-management finding, not a coverage backlog item.

open as a page

showing 1–30 of 36