In an EDR bake-off, how would you falsify a vendor's claim that behavioural AI convicts fileless attacks?
answer
- Test the output, not the model
- Turn each claim into a prediction that could fail
- Your image, your policy, your envelope
- Count files through, not alerts raised
- Ask what it will not convict
basics
~20 sYou cannot inspect the model, so test its output. Execute the behaviour yourself in the trial tenant under the policy you would really deploy, and measure what got through and the damage before the action — not alert counts.
solid answer
~50 sTreat the claim as a hypothesis and design the test that would disprove it. Since the model is not inspectable, judge only the observable output: run the behaviours yourself in the vendor's proof-of-concept tenant, on your own build image, with the policy you would really deploy rather than their demo profile. Include the cases the claim is weakest at — a signed, allow-listed tool driven maliciously; a chain halted halfway; a host that has lost its cloud connection. Then measure four things separately, because a failure at any one looks the same from the console: did telemetry arrive, did the engine convict, did it act, and how much damage preceded the action. Report damage in files, not alerts. Finally, ask each vendor in writing to name three behaviours the engine will not convict, and what rollback does not restore.
go deeper
Understand that a vendor demonstration is designed to succeed, and that an evaluation only means something when your own team executes the test on your own configuration.
Be ready to convert a capability claim into a concrete test and to separate the four ways a miss can happen: no telemetry, no conviction, no action, or action too late.
Show you can design the adversarial cases that would break the claim — a signed allow-listed tool driven maliciously, a halted chain, a sensor cut off from the cloud — and measure damage rather than alerts.
Own the procurement judgment: weight autonomous action for the team you actually have, demand written failure conditions, and commit to re-testing after purchase because a behavioural claim decays silently.
## The problem you are actually solving As the manager running the bake-off you are being asked to sign for a control whose internals you are not allowed to see. You will not get the model, the feature set, or the thresholds. That is not unusual and it is not automatically bad faith, but it does mean the only honest evaluation is behavioural: **you cannot test the mechanism, so you test the claim's consequences.** The framing that keeps this rigorous is falsification. Do not ask "can it catch things" — every product can catch something, and a vendor-run demo is built to prove exactly that. Ask instead: *if this claim were false, what would I see?* Then go and look for that. ## Turn each marketing sentence into a testable prediction - "Behavioural AI convicts fileless attacks" predicts: with no file on disk to score, the engine still convicts and still acts. Test: run in-memory execution techniques from an interpreter on your own image and observe whether anything beyond telemetry happens. - "Convicts living-off-the-land abuse" predicts: a signed, allow-listed administrative tool driven maliciously is convicted on conduct. Test: drive exactly such a tool through a full impact chain — enumerate, remove local restore points, bulk rewrite — against scratch data. - "One-click rollback" predicts a recovery boundary. Test each edge separately: local volume, mapped share, files written after the last snapshot, and a group left in detect-only. Run the halted case too. Stop the chain after the second step and see whether anything at all is convicted, because a product that only convicts complete attacks tells you nothing until the damage is done. ## Control the conditions, or you have measured the vendor's lab Three conditions decide whether the result transfers to your estate. - **Your image and your policy.** Deploy the configuration you intend to run, not the profile the sales engineer sets up. If you will run some server groups in detect-only during rollout, test that state, because it is the state the estate will really be in. - **Your operating envelope.** Sever the sensor's cloud connectivity mid-test. Many behavioural claims lean on cloud-side scoring; you need to know what the local engine still does when the link is down, because an adversary can arrange for it to be down. - **Your people.** With one analyst and no overnight cover, a product whose value arrives as console enrichment for a human to interpret is worth much less than one whose autonomous action is trustworthy. Weight the criteria to match the team you have, not the team the reference architecture assumes. ## Measure four things, never one When something is missed, the console shows the same emptiness regardless of cause. Separate them deliberately: 1. **Did the telemetry arrive?** Was the activity recorded at all, whatever the verdict? 2. **Did the engine convict?** Was a verdict reached from that telemetry? 3. **Did it act?** Kill, quarantine, isolate, roll back — or only report? 4. **How much happened first?** Files written, minutes elapsed, hosts reached. A product that records everything and convicts nothing is a very different purchase from one that convicts and is configured not to act. Only the second is fixable with policy. ## What not to count **Alert volume is not a score.** A noisier product will win a naive count and lose the estate, because one analyst cannot work an inflated queue. **Detection counts on a public test corpus are not your estate**, and neither are the specific named attacks the vendor's own material references. **A demo you did not drive is evidence about the demo.** ## Ask for the failure conditions in writing The most informative question in a bake-off is not about strengths. Ask each vendor to name three behaviours their engine will not convict, and to state precisely what rollback does not restore. A vendor with a real engineering culture answers this readily and specifically; the answer tells you where to look next and becomes the basis of what you are willing to promise your own business. A vendor who cannot produce it has given you the most useful finding of the evaluation. ## Close the loop after purchase The bake-off result decays. The behaviour that was convicted at trial may not be convicted after a sensor upgrade or a policy change, and unlike a rule you can replay against a stored input, the only proof a behavioural control still works is **executing the behaviour again**. Schedule that re-execution as a standing exercise; it is the same test, and it is the only evidence that the thing you bought is still the thing you have.
- Two products both convicted the chain. How do you pick between them?Compare what happened before the action: files written before mitigation, whether the whole process tree was killed or only one process, whether the recovery boundary was where the vendor said, and how much analyst time reaching the verdict took. With one analyst and no night cover, the earlier conviction and the more trustworthy autonomous action wins, even if the other console is richer.
- The vendor refuses to allow adversary simulation in their trial tenant. What do you conclude?Not necessarily bad faith — some trials are shared infrastructure — but it means the trial cannot produce the evidence you need. Either negotiate an isolated tenant with a scoped and written rules-of-engagement, or record that the capability claim went untested and weight it at zero rather than accepting it on assertion.
- A year after purchase, how do you know the behavioural engine still convicts what it convicted at trial?Re-execute the behaviour. A behavioural control's silence is indistinguishable from a quiet estate, so its false-negative rate is not computable from its own output and it cannot be proven by replaying a stored input. Running the technique again on a scheduled basis is the only evidence, and it also catches the case where a policy or exclusion change quietly removed the coverage.
saying these in an interview costs you the question
- Accepts a vendor-driven demo as evaluation evidence
- Scores products by number of alerts raised
- Tests only the attacks the vendor's material names
- Treats an uninspectable model as unfalsifiable and gives up
- Never tests the sensor with cloud connectivity removed
- Assumes the trial result stays true after upgrades