A rule tuned live in a purple-team session goes green - how do you prove it detects the technique, not the operator?
answer
- a sample of one execution
- which fields change for free next time
- user agent and source address are the person
- ask a second operator to do it differently
- the first green is not the exit criterion
basics
~20 sSeparate what is invariant about the technique from what is incidental to the person: user agent, source address, parameter spelling and call ordering are the operator. Then have a second person execute it differently, and re-open if it misses.
solid answer
~50 sA green in the room proves that this behaviour, performed this way, by this person, with this tooling, produced an alert. Tuning under an audience pushes the engineer toward whatever condition makes the screen go green soonest, and the fastest conditions are almost always incidental: the client's user agent string, the operator's source address, the exact ordering and spacing of the two API calls, the spelling of a parameter or policy name. Those are indicator-grade artefacts and they are attacker-controlled. Anchor instead on what the technique cannot avoid - the API action itself, the relationship between the caller and the target identity, that the target already holds credentials, the rarity of that identity performing it. Then test it: a second operator, a different client, the calls in the opposite order with a delay. The loop exits when a variant nobody in the room built fires, not on the first green.
code
json · 14 lines{
"eventTime": "2026-03-04T14:02:11Z",
"eventSource": "iam.amazonaws.com",
"eventName": "CreateAccessKey",
"awsRegion": "us-east-1",
"sourceIPAddress": "198.51.100.24",
"userAgent": "aws-cli/2.15.30 Python/3.11.8 Darwin/23.4.0",
"userIdentity": {
"type": "IAMUser",
"arn": "arn:aws:iam::111122223333:user/rt-operator"
},
"requestParameters": { "userName": "svc-backup" },
...
}go deeper
Know that a rule can fire on one demonstration and still miss the same technique performed slightly differently, and that things like a client version string are easy for anyone to change.
Be able to sort the fields of an audit record into incidental to this operator and invariant to the technique, and explain why the first group makes fragile rule conditions.
Show that you design the second test rather than accepting the first green - a different operator, a different client, reversed ordering - and that you re-open the finding when the variant misses.
Own how session results are worded downstream, so a single-variant green never travels through the organisation as a coverage claim about a technique.
## The failure this question is about A rule tuned in a live session is fitted to a sample of one. The engineer has exactly one execution in front of them, the operator is waiting, and every condition added narrows the rule toward that single record. The result is a rule that goes green in the room and encodes the operator's tradecraft rather than the technique. ## Invariant versus incidental Take a cloud persistence step: mint a second access key for an existing identity, then attach an inline policy to it (ATT&CK `T1098.001`, additional cloud credentials). The audit records carry two very different kinds of field. **Incidental to this operator** - changes for free, next time: - the `userAgent` string, which names the client and its version - `sourceIPAddress` - the ordering of the two calls and the seconds between them - the target identity's name, the policy name, the exact policy document formatting - the region the console happened to be pointed at **Invariant to the technique** - hard to avoid while still achieving the objective: - the API actions themselves, and the fact that both occur against the same target identity - the caller creating credentials for an identity that is not the caller - the target already holding an active key, so this is an additional credential rather than a first one - the rarity of this principal performing credential creation at all, measured against the account's own history - the absence of the change from any change record or deployment window A rule leaning on the first list is an indicator rule wearing a behaviour rule's clothes. This is the indicator-versus-behaviour distinction with money on it: an artefact you observed is cheap for an adversary to change, a behaviour required by the objective is not. ## How you actually test it, in the room Ask for a variant that changes the field you most suspect the rule is leaning on. If you narrowed on the client, have the operator repeat the technique from the console or an SDK. If you narrowed on ordering, ask for the policy attach first, or a fifteen-minute gap between the calls. Best of all, have a *different person* execute it - somebody who was not watching the rule being written and will not unconsciously reproduce the shape it expects. A useful in-room habit: before the retest, the engineer states out loud which condition they expect the variant to break. If the rule survives a variant the engineer predicted would kill it, that is worth more than three greens on the original. ## The exit criterion The loop does not exit on the first green. It exits when a variant that nobody in the room designed the rule against fires - and if the second variant misses, the correct move is to **re-open the finding**, not to argue that the first result stands. The miss is the more informative result: the difference between the variant that fired and the one that did not is the specification for the next edit. This matters because of what happens to session output afterwards. A green from a session gets written into a report and read as coverage. If the only evidence behind it is one operator's execution, the organisation now believes it detects a technique it detects one performance of. ## Documenting honestly Write the variants down: which were executed, by whom, with which client, in which order, and which fired. State the ones you did *not* test. A finding that says *detected for CLI-driven executions by a single operator; console and delayed-ordering variants untested* is far more useful six months later than a green tick, and it tells the next person exactly which execution to try first. ## What interviewers listen for They want to hear you volunteer that a same-sitting green is weak evidence, name specific fields as incidental rather than gesturing at overfitting in the abstract, and describe a concrete second test. Candidates who describe the loop as *tune until it fires* have described the defect.
- The operator has time for one more execution. Which variant do you ask for?The one that breaks the condition you most suspect is load-bearing. If the rule narrowed on the client, ask for the same two actions from the console or an SDK; if it narrowed on ordering or timing, ask for them reversed with a long gap. Say out loud beforehand which condition you expect the variant to break.
- The second variant does not fire. What is the status of the finding?Re-opened. The earlier green was evidence about one execution path and it still stands as that; the miss is simply the more informative result. Record both variants, keep the gap open, and treat the difference between them as the specification for the next edit rather than as a reason to defend the rule.
- Is keying a rule on the tool's user agent ever acceptable?As enrichment or a low-confidence supplementary signal, sometimes. As the load-bearing condition, no: a user agent is client-supplied, trivially changed, and describes an artefact rather than the behaviour required by the objective. It is the sort of condition that survives an exercise and dies against the first real intruder.
- How do you write the result up so it is not later read as coverage?Name the variants executed, the operator and client used, and explicitly list what was not tested. Detected for CLI-driven executions by one operator, console and delayed variants untested is honest and actionable; a green tick against the technique name is neither.
Tuning against one execution is fitting a curve through a single point: it passes perfectly through the sample and tells you nothing about the next one.
saying these in an interview costs you the question
- Treats the first green retest as coverage of the technique
- Anchors the rule on user agent or source address
- Confuses an observed artefact with the behaviour itself
- Narrows conditions until it fires, then stops
- Refuses to re-open the finding when a variant misses