What does a green T1218 cell overstate when only two of its signed-binary proxy sub-techniques were tested?
answer
- the parent is not the children
- twelve squares wearing one colour
- same idea, different binaries and parents
- roll up on the worst child
- the variant you ran is the claim
basics
~10 sIt overstates the unit of the claim. Sub-techniques use different binaries, command lines and parent processes, so catching rundll32 and regsvr32 says nothing about the ten siblings nobody executed.
solid answer
~50 sA technique-level cell aggregates over children that share a name and very little else. `T1218` covers proxy execution through a range of signed system binaries, and a rule keyed to `rundll32.exe` command-line shapes (`T1218.011`) tells you nothing about `mshta` or `odbcconf` - different image, different parent, sometimes a different log source entirely. So a green cell built on two tested children is claiming coverage of ten it never touched. The second overstatement is inside the child: executing one procedure variant proves that variant was caught, and the same sub-technique invoked from a scheduled task or a COM hijack arrives with a different parent process and may not match the rule that fired in the test. The honest unit is sub-technique per platform, with the tested variant named. If you must roll up for a rollup audience, derive the parent's colour from its **worst** child, never its best, and carry the tested-of-total count with it.
code
json · 11 lines{
"technique": "T1218",
"description": "proxy execution via signed system binaries",
"status": "covered",
"platforms_claimed": ["windows"],
"sub_techniques_in_scope": 12,
"sub_techniques_validated": ["T1218.011", "T1218.010"],
"variant_executed": "shell parent, export from user-writable path",
"rules": ["SOC-0421", "SOC-0433"],
"last_validated": "2026-02-14"
}go deeper
Know that techniques have sub-techniques and that testing one does not cover the others. Being able to say a rule for one signed binary does not carry over to another one is enough at this level.
Explain the mechanics: different image, parent process and command-line grammar per child, so the detection logic does not transfer. Then go one level deeper and note that even a tested child is only proven for the variant that was executed.
Show the fix rather than the complaint - measure at sub-technique per platform, roll up on the worst child, carry the tested-of-total count, and record the executed variant so a challenged cell can be defended or re-run.
Be ready to decide the granularity the organisation reports at, knowing a finer grid is more honest and less readable, and to hold the line that a rollup must inherit its weakest child's colour when a leader asks for one number.
## Why one square over a dozen children is the classic overstatement A coverage heatmap has to be drawn at some level of granularity, and the level almost everyone draws it at - the technique - is the level at which the underlying behaviours have least in common. `T1218` is the standard illustration: proxy execution through a signed, already-trusted system binary. Its children include `T1218.011` (`rundll32.exe`), `T1218.010` (`regsvr32.exe`) and `T1218.005` (`mshta.exe`), among others. They share an idea, not a detection. Different image names, different command-line grammars, different typical parents, and in some cases a different observation surface altogether. If detection engineering wrote two rules keyed to the `rundll32` and `regsvr32` cases, executed both, and coloured the parent cell green, the cell is now asserting something about every other child on the strength of evidence about two. Nothing about the artefact records that. ### The second overstatement lives inside the child Even a validated sub-technique is only validated for the *procedure variant* that was executed. Suppose the test launched `rundll32.exe` with an unusual export from a user-writable directory, with a shell as the parent process, and the rule fired on the parent-child pair. The same sub-technique invoked from a scheduled task, from a service, or through a hijacked COM registration arrives with a completely different parent, and a rule whose logic leans on that parent will not fire. Nobody enumerated that persistence variant, so nobody tested it, so the green square silently covers it. This is the difference between a technique (an idea), a sub-technique (a narrower idea) and a procedure (the concrete way somebody did it). Detections are written against procedures and claim coverage of ideas. ### The third axis: platform One cell is usually drawn once for the whole estate, but the same behaviour class has different maturity per platform. A technique can be genuinely well covered on the managed Windows endpoints, thin on the Linux server fleet where the audit rules were never extended, and impossible on the container platform. Rendering that as one colour picks a winner, and the winner is always the platform that was measured most. ### What to do about it **Draw the grid at the level you actually test.** Sub-technique per platform is the honest cell. If the grid is too large to read that way, keep sub-technique as the record and render the technique-level view from it, rather than the other way round. **Roll up on the worst child, not the best.** A parent coloured by its strongest child is a marketing artefact. Coloured by its weakest, it is a work list. The rollup should also carry a count - two of twelve children tested - so the reader knows the denominator. **Record the variant.** Alongside a validated child, record what was actually executed: the parent process, the invocation shape, the host class. That single field is what lets a later reader tell a broad claim from a narrow one, and it is what makes the cell re-testable when someone challenges it. **Name untested siblings explicitly.** An enumerated but untested child is a work item. A child nobody ever wrote down is a blind spot that will not appear in any review. ### Reading the artefact critically When you are handed a coverage matrix - your own or a supplier's - the questions that collapse the overstatement fastest are: at what level was this drawn and at what level was it measured; how many children does each green parent have and how many were exercised; which platforms is the colour asserting over; and, for each validated cell, what exactly was run. Coverage claims almost never survive the fourth question in their original colour, and interviewers ask this precisely because it separates people who produce the artefact from people who audit it. ### The counting trap underneath all of it Counting rules is not counting coverage. Ten rules against one sub-technique and none against its eleven siblings is a large rule count and a narrow claim; teams that report rule counts as a coverage metric are measuring their own output rather than the adversary behaviour they can see. The unit of a coverage claim is always a behaviour that was executed and observed, never an artefact that was authored.
- Is a technique-level view ever the right thing to publish?Yes, for an audience that cannot read a two-hundred-row grid. The conditions are that the record underneath is kept at sub-technique level, the parent's colour is derived from its weakest child, and the tested-of-total count travels with the cell. A rollup that hides its denominator is where the overstatement enters.
- One rule catches the behaviour only when the parent process is an Office application. How should that cell read?As a conditional claim, not a green one. Record the condition as part of the assertion - detected when spawned by Office, untested otherwise - because the same sub-technique launched by a scheduled task or a service will not match. If the grid has no room for the condition, the honest colour is the one for partial, with the tested variant named behind it.
- Why is 'we have forty rules for this technique' a poor answer to a coverage question?Because it counts your own output rather than adversary behaviour that was observed. Forty rules can all cover one sub-technique and one invocation shape. The unit of a coverage claim is a behaviour executed and seen, so the answer an interviewer wants is which children were exercised, on which platforms, and when.
saying these in an interview costs you the question
- Assuming sub-techniques of one technique share telemetry
- Rolling a parent cell up to its best child
- Treating one procedure variant as the whole sub-technique
- Counting rules written instead of behaviours observed
- Colouring one cell for an estate with several platforms