skip to content

Tune and Retest Loop

Both teams in one room: execute, watch the console, adjust the rule, execute again within the hour. Interviewers ask what that tempo costs and when it overfits a rule to one operator.

on this pageshow

explore

questions

4

In a purple-team session, why re-run the technique after the detection engineer edits the rule?

level: juniorimportance: must knowfreq 62%

answer

  1. an edit is a hypothesis, not a result
  2. the edit changes matching, not history
  3. delivery, parsing, schedule, suppression, routing
  4. green once, by one person, one way

basics

~10 s

Because an edited rule is only a hypothesis until the behaviour it targets is performed again while it is live. Re-execution proves the whole path, from event delivery to a routed alert, actually works.

solid answer

~50 s

The loop is execute, observe, tune, retest. Editing a rule changes what will be matched from now on; it does not change what was already recorded, and it proves nothing about the alerting path. Between a control-plane audit record existing and an analyst seeing an alert there are several independent stages: delivery to the platform, parsing into the fields the rule names, matching when the rule next runs, suppression or deduplication, and routing to a queue. A hand-written test event exercises the syntax and the parser only. Re-executing the technique exercises all of it at once, with a real record produced by real behaviour. That is why the retest is the evidence step of the session, and why a session that ends with an edit and no re-execution has produced a claim rather than a result.

go deeper

for a junior

Be ready to name the four steps of the loop in order and say plainly what the re-execution adds: it is the only step that proves the alert actually arrives, not just that the query looks right.

for a middle

Expect to be asked what sits between the record and the alert - delivery, parsing, schedule and lookback, suppression, routing - and which of those a synthetic test event never touches.

for a senior

Demonstrate that you separate a telemetry gap from a detection gap on the first pass, and that you record a green as evidence about one variant rather than as coverage.

for a principal

Own the framing that a session's product is evidence, not shipped rules, and that the value of both teams sitting together is the length of the feedback loop, not the volume of findings.

## The loop A purple-team session runs a short cycle with both teams in the room: the operator **executes** one technique against a real estate, the detection engineer **observes** what the telemetry actually recorded, the engineer **tunes** the query or rule, and the operator **re-executes** the same technique so the edited logic meets live behaviour. The cycle repeats until the rule fires on a fresh execution, or until the team concludes the behaviour is not detectable with the telemetry present, which is itself a legitimate result. ## What each pass answers The first execution answers a question that is not about detection at all: *is the behaviour recorded?* If the operator mints a second access key for an existing cloud identity and then attaches an inline policy to it (ATT&CK `T1098.001`, additional cloud credentials), the control plane either writes audit records naming the caller, the target identity and the API action, or it does not. If no record exists, no rule can be written, and what you have found is a **telemetry gap**, not a detection gap. Those route to different owners and different fixes. Every later pass answers a different question: *does the candidate logic match a real occurrence of the behaviour, delivered through the real pipeline, on the real schedule?* ## Why an edit is not evidence An edit changes matching, not history. Two things follow. First, the records from the previous execution are already written; re-running a saved search over them tells you the query would have selected those rows. That is worth doing - it is the fast inner check - but it validates a query, not a detection. A detection is a rule plus a schedule plus a suppression policy plus a destination. Second, the failure modes that bite in production live outside the query. A rule can be perfectly correct as a search and still never alert because a field it names is populated on only some events, because its scheduled lookback is shorter than delivery lag so late-arriving records fall outside every run, because a broad suppression swallows it, or because it routes to a queue nobody works. Only a live re-execution walks that whole path end to end. ## What the first green does and does not prove A green retest proves one thing precisely: *this behaviour, performed this way, by this person, with this tooling, produced an alert.* That is a real result and it is not nothing. It is also not coverage. The same technique performed from a console rather than a command-line client, with the two API calls in the opposite order, or with a delay between them, is a different execution and an untested one. The exit criterion for the loop is not the first green; it is a green on a variant the engineer did not watch being built. ## What to record when it goes green Keep a session log that lets someone else reconstruct the claim months later: the exact execution time, the identity and account used, the technique and variant, the rule version that fired, the observed alert time, and the delta between them. Note explicitly that it is one variant by one operator. That single sentence is what stops a session result being read as a coverage claim. ## Why this is asked in interviews Candidates who have only read about purple teaming describe it as a report exchange: red hands over findings, blue writes rules later. The whole value of running it in one sitting is the tightness of the feedback - the operator is still there, the estate is still in the state that produced the record, and a broken assumption is discovered in minutes rather than in the next quarter's exercise. Someone who cannot say why the re-execution matters has not sat in one.

  • The engineer re-runs the saved search over the last hour instead of asking for a re-execution. Is that the same thing?
    No. It checks the query against records already stored, which is a useful fast inner check, but it exercises none of the live path: the schedule, the lookback window, suppression and routing all go untested. It also re-uses the same single execution, so it cannot reveal that the logic is leaning on something specific to that one run.
  • What do you write down when the retest goes green?
    The execution time, the identity and account, the technique and the exact variant performed, the rule version that fired, the alert time and the delay between the two. Add one sentence stating that this is one variant by one operator, so the result is never later read as a coverage claim.
  • The first execution produced no audit record at all. What have you found?
    A telemetry gap, not a detection gap. No rule can be written over a record that does not exist, so the finding is about collection - a source not enabled, an account not covered, an event class not captured - and it belongs to whoever owns that pipeline rather than to the detection engineer.

Editing the rule is writing down a prediction. Re-executing the technique is running the experiment. A session that stops after the edit has only ever made predictions.

saying these in an interview costs you the question

  • Says the rule is proven because it matched a hand-written test event
  • Treats the first green retest as the end of the loop
  • Thinks editing a rule changes what was already recorded
  • Calls a backfill search over old records a retest
  • Describes purple teaming as red handing a report to blue

context

open as a page

A rule tuned live in a purple-team session goes green - how do you prove it detects the technique, not the operator?

level: seniorimportance: must knowfreq 51%

basics

~20 s

Separate what is invariant about the technique from what is incidental to the person: user agent, source address, parameter spelling and call ordering are the operator. Then have a second person execute it differently, and re-open if it misses.

open as a page

Cloud control-plane audit events arrive minutes after the API call - how does that shape a live tune-and-retest session?

level: middleimportance: should knowfreq 44%

basics

~20 s

Delivery lag sets the floor tempo: you cannot retest faster than records arrive. Measure the lag at the start, check the raw record before declaring a miss, and pipeline executions instead of one edit per pass.

open as a page

Mid-session, your new key-creation rule also matches the CI pipeline's own key rotation - do you keep tuning in the room?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

No. First confirm from the audit records that the extra matches really are the pipeline, not a second unknown actor. Then stop tuning: an owner outside the room is involved, so the gap leaves the session still open.

open as a page