skip to content

A threat report says the crew phishes OAuth consent for mail-read scope. What turns that into an executable test?

level: middleimportance: should knowfreq 44%

answer

  1. a class of behaviour is not a test
  2. the procedure is the missing half
  3. consent path and scope are choices
  4. name the observable before you run
  5. success and detection are separate outcomes

basics

~20 s

A named behaviour is a class, not a test. You must add the procedure: which application, which permission scope, which consent path, which target identity, plus the observable you expect and the criterion that decides pass or fail.

solid answer

~50 s

A technique identifier tells you what class of behaviour to reproduce; an operator cannot execute a class. The test case has to specify the procedure: an application registration under your control, the exact delegated scope requested, whether the grant goes through user consent or admin consent, which test identity grants it, and what the token then does - for example a paced read of a seeded mailbox rather than bulk collection. It also has to state the **expected observable** (the identity provider's consent and permission-grant audit entries, the tenant's own audit log for the API reads) and the **success criterion**, and those are two different things: succeeding at the behaviour and being detected are independent outcomes you must record separately. Where the report is silent - and it usually is on consent path and scope - you choose a variant, write the assumption down, and flag it, because detections often key on one implementation.

code

text · 12 lines
text
THREAT PROFILE - CREW "HOLLOW-KESTREL" (financial services SaaS tenants)

Initial access: consent phishing. Targets receive a link to an
attacker-registered multi-tenant application requesting delegated
mail-read scope. No credential is captured and no endpoint payload
is deployed.

Collection: mail is read through the granted token via the vendor API
over several weeks; volume is low and paced to resemble a mail client.

Persistence: the grant itself. A password reset does not remove it.
...

go deeper

for a junior

Know that a named technique is a category of behaviour, and that somebody still has to decide the concrete steps, the target and the tooling before anything can be executed or judged.

for a middle

Be able to list what a test case must specify beyond the technique name - procedure parameters, expected observable, success criterion, abort condition - and explain why the expected observable is written before the run rather than after.

for a senior

Demonstrate that you scope claims to the fidelity you actually achieved, plan variants to measure how narrow a detection is, and record the assumptions you invented where the intelligence was silent.

for a principal

Own the standard: what an emulation report is permitted to assert given the fidelity it bought, and how much realism the organisation will fund before the residual gap is accepted rather than closed.

## Technique, procedure, test case Three things get confused here, and interviewers separate them deliberately. - A **technique** is a class of adversary behaviour with a stable identifier, useful for labelling and reporting. - A **procedure** is one concrete way that behaviour was carried out: the tooling, the parameters, the sequence, the artefacts it leaves. - A **test case** is a procedure plus everything needed to run it and judge it: preconditions, execution steps, the observable you expect, the criterion for success, and the criterion for stopping. A plan that lists technique identifiers is a scope document. An operator handed only that will improvise a procedure, and the improvised one is what your results are actually about. ## What the report gives you and what it does not A profile paragraph typically gives the shape of the behaviour and the outcome. It rarely gives: - **Which consent path.** A user consenting for themselves and an administrator consenting on behalf of the tenant are different acts, leave different audit entries, and are gated by different tenant settings. If the estate blocks user consent outright, one variant is unexecutable and that fact is itself a finding about a preventive control. - **Which scope.** Read-only mail access and broader mailbox permissions are different grants and different subsequent behaviour. - **What the token is used for and how fast.** Slow paced reads and a bulk pull look nothing alike in volume-based signals. - **What the application looks like.** Display name, publisher state, redirect target: detections frequently key on exactly these attributes rather than on the behaviour. The honest response to each gap is the same: pick the variant that is plausible for your estate, record the assumption in the test case, and mark it as an open question for the intelligence side rather than silently deciding. ## Fidelity decides what a result means Emulation trades fidelity against safety and effort. If your implementation differs from the reported one in the attribute a rule keys on, a no-detection result is ambiguous: it may mean the estate is blind to the behaviour, or it may mean your application did not look like the one the rule was written for. That distinction is the difference between a detection-engineering finding and a note that your tooling was unrepresentative. The useful discipline is to state, per test case, which attributes are faithful and which are stand-ins. Faithful: the grant really happens, the token is really issued, the API really returns mail. Stand-in: the application is yours, the mailbox is seeded, the volume is bounded, nothing leaves the tenant. The claim you write afterwards is scoped to the faithful part. ## Expected observable and success criterion Every test case names, in advance, **where the evidence should appear**: the identity provider's audit trail for the consent event and the resulting delegated-permission grant, and the SaaS tenant's audit records for the mail reads made with that token. Naming it in advance is what makes the exercise diagnostic - if nothing appears in the log at all, that is a telemetry finding; if the record is there and no rule fired, that is a detection finding; if a rule fired and no analyst acted, that is a process finding. Writing the observable afterwards lets you rationalise whatever happened. Success criteria are separate from detection outcomes. "Did the operator obtain mail?" and "did anything fire?" are independent, and all four combinations are informative - including the uncomfortable one where the behaviour succeeded and an alert fired but nobody worked it. ## Variants are part of the design One technique usually deserves more than one test case, because a rule that keys on a single implementation will pass one variant and miss the next. Two or three procedures for the same behaviour - different scope, different consent path, a different application profile - turn a binary result into a description of how narrow the detection is. ## What the artefact looks like A good test case reads like an experiment: preconditions, steps precise enough that a different operator reproduces it, the expected observable and where to look for it, success criteria, abort conditions, and the assumptions you had to invent because the report was silent. That last section is what lets the next cycle improve the fidelity instead of repeating the same guess.

  • The report does not say whether consent was granted by a user or by an administrator. How do you proceed?
    Pick the variant your tenant actually permits, write the assumption into the test case, and raise the ambiguity with the intelligence side. The two paths produce different audit entries and are gated by different settings, so the choice changes both what you can execute and what a result means. If user consent is blocked outright, that is a preventive-control finding worth reporting on its own.
  • Why does procedure fidelity change what a no-detection result proves?
    Detections often key on implementation attributes - an application's display name, the requested scope string, the client identifier - rather than on the behaviour itself. If your stand-in differs there, a miss may say your tooling was unrepresentative rather than that the estate is blind. State per test case which attributes were faithful, and scope the claim to those.
  • How many test cases should one technique get?
    More than one whenever variants are cheap. A single procedure yields a binary result; two or three - different scope, different consent path, a different application profile - tell you how narrow the detection is. That is usually the more actionable finding, because it predicts which real-world variation would slip past.

saying these in an interview costs you the question

  • Treats the technique identifier as the test case
  • Runs a generic tool and calls it emulation of that crew
  • Never states which log should carry the evidence
  • Assumes any implementation of the behaviour proves the same thing
  • Confuses the operator succeeding with the behaviour being undetected

context