What does a purple-team test where your detection fired prove about a quiet quarter?
answer
- supplying a known positive
- tests the live chain, not a fixture
- record the stage, not pass or fail
- only the technique you executed
- generalises to detection, not to 03:00
basics
~20 sIt proves the chain from telemetry to rule to analyst worked for the one technique executed, on the hosts in scope, at that hour. It is a positive control for the pipeline, not evidence no adversary was present.
solid answer
~50 sA purple-team result is a positive control: you inject a known-malicious behaviour so a pipeline whose normal output is silence produces something observable. What it licenses you to claim is narrow and should be stated that way - this technique, executed this way, on these hosts, at this hour, was detected. Report the stage it reached, because 'the rule fired' and 'an analyst escalated within twelve minutes' are different claims, and so is 'the telemetry was there but nothing matched it'. What it cannot do is stand in for real-adversary evidence: it says nothing about techniques you did not execute, about a variant that changes the observable the rule keys on, or about the hours your two-person team is not staffed. In a report for a quarter with no intrusion it is the strongest positive finding you have, and it rebuts 'the sensors are blind' only for the slice you tested.
go deeper
Know that an exercise deliberately runs attacker behaviour so defenders can see whether anything notices, and that a detection firing during it proves the rule matched that behaviour and nothing more.
Explain why silence-producing systems need a positive control, and break a result into telemetry, detection, analyst and response so you can say which stage the estate actually reached.
Show judgment about scope: which techniques were worth executing against this estate, what a variant would do to the rule that caught it, and how you keep the untested list in front of the reader instead of quoting a coverage percentage.
Set the standard for what exercise results are allowed to claim in reporting to executives, and make sure the untested surface is stated with the same prominence as the passes.
## Why an exercise result exists at all A detection programme has a measurement problem that most engineering systems do not: **its normal output is nothing, and nothing is also what failure looks like.** You cannot compute how often it missed, because misses leave no record. The only way to get an observable out of it is to **supply a known positive** - run the adversary behaviour yourself and watch what the estate does. That is a positive control, in the same sense a lab runs a known-positive sample to prove the assay still works. This is why a purple-team result is qualitatively different from replaying a saved input through a rule. Replaying a captured record tests the rule's logic against a fixture. Executing the behaviour tests **collection, rule, analyst and response together, in the live configuration**, including all the things nobody wrote down: whether the agent was actually installed on that host, whether the log source was still forwarding, whether the alert routed to a queue anyone reads. ## Report the stage, not a pass A useful result is recorded as the furthest stage it reached: - **Telemetry present** - the records describing the behaviour arrived (for example a process-creation record carrying the command line and the parent image). - **Detection fired** - a rule matched those records and raised an alert. - **Analyst acted** - a human worked the alert and reached the right verdict, within a measurable time. - **Response landed** - the containment action actually executed on the target. Collapsing all four into 'pass' or 'fail' destroys the diagnosis. 'The data was there and no rule matched' is a detection-engineering gap. 'The rule fired and the alert sat unworked for two days' is a staffing or routing gap. Those need different money and different fixes, and in a budget conversation they are different arguments. ## The limits you must state - **Technique scope.** A detection for one technique identifier - say scheduled-task persistence, `T1053.005` - was exercised. Nothing was established about the techniques you did not run. Coverage claims made from a handful of executions are the most common overreach in this area. - **Variant fragility.** Many rules key on a specific observable rather than the behaviour: a binary name, an argument string, a parent-child process pair. An operator who changes that observable can defeat the rule while performing the same technique. A pass tells you the rule caught **that** implementation. - **Time of day.** In a small team with no overnight coverage, a test run at 11:00 exercises a staffed analyst stage. The same behaviour at 03:00 reaches a queue nobody is reading until morning. The detection result generalises; the response time does not. - **Artificial conditions.** If the exercise ran from a host chosen because it definitely had the endpoint agent, or from an account exempted from blocking so the chain could continue, the result is optimistic relative to a real intrusion. - **No adversary adaptation.** A red-team operator working to a scenario does not iterate against your alerts the way an intruder who notices a response will. ## What it can and cannot substitute for It **can** substitute for the claim 'we would have seen it' in the narrow tested slice, and it is the only evidence available in a quarter with no real intrusion that is generated by the defence rather than by luck. It **cannot** substitute for evidence about adversary presence or absence, and it must never be reported as though it were. The honest sentence in a report is: 'we executed N techniques against production; M were detected end to end; here is the list, and here is what remains untested.' ## Reporting language that survives challenge Name the technique, the hosts, the date and hour, the stage reached, and the time between execution and analyst verdict. Keep the untested list visible in the same table, because a table of only passes invites the reader to generalise, and the first challenge you will get is exactly that generalisation.
- The rule fired during the exercise but nobody worked the alert. Is that a pass?No, and it is the most useful result of the exercise. Record it as 'detection fired, analyst stage not reached'. The telemetry and rule are sound and the gap is routing or staffing, which is a different fix and a different budget line from writing a new detection. Reporting it as a pass hides exactly the failure the exercise was run to find.
- The exercise ran at 11:00 on a Tuesday. What does it not tell you about 03:00?Nothing about the human stages. Collection and rule matching are continuous, so a detection that fired at 11:00 will fire at 03:00, but with no overnight staffing the alert waits until morning. So the exercise supports a claim about detection coverage and not about response time outside staffed hours, and I would report those as two separate figures.
- How many techniques do you need to execute before you can claim coverage?Coverage is not a number you reach; it is a list you publish. I report which techniques were executed and detected, which were executed and missed, and which were never tested, prioritised by what is plausible against this estate. A percentage over a framework's full matrix invites a false sense of completeness and I avoid quoting one.
It is the known-positive control on a lab assay. A positive result proves the test can still detect the sample you spiked it with; it says nothing about the samples you never ran.
saying these in an interview costs you the question
- Presents an exercise result as proof the estate was clean
- Calls it a pass when the rule fired but no analyst acted
- Generalises one detected technique to broad coverage
- Ignores that the operator ran from an exempted account
- Treats exercise response times as overnight response times
- Says an exercise measures the false-negative rate