skip to content

Negative Test Criteria

An abuse case earns its keep when it becomes an acceptance criterion and then an automated test that fails if the abuse succeeds. Interviewers ask which abuse cases resist automation entirely.

on this pageshow

questions

3

How do you turn an abuse case into an acceptance criterion and then an automated negative test?

level: middleimportance: must knowfreq 62%

answer

  1. invert the story into a must-not
  2. attach an observable, not a procedure
  3. fixture sits in the attacker's seat
  4. refusal, unchanged state, recorded evidence
  5. races assert invariants, never timing

basics

~10 s

Restate the attacker's goal as a must-not statement with an observable outcome, then write a test that runs from the adversary's position and asserts the refusal, the unchanged state, and the recorded evidence.

solid answer

~50 s

There are three artifacts and you move between them deliberately. The abuse case is prose: an attacker goal against a feature. The acceptance criterion is that goal inverted into a must-not with something measurable attached. The negative test is the executable version: it authenticates as the adversary's principal, performs the attempt, and asserts three things — the specific refusal the design chose, that the targeted state did not change, and that the denial left evidence. Take a gift-card balance API. The abuse case is "redeem the same card twice at once". The criterion becomes: for a card with one unit of balance, any number of concurrent redemptions settles exactly one, the rest are refused, and the ledger holds exactly one debit. The test seeds the card, fires ten concurrent redemptions from that customer's session, and asserts on the invariant — one success, zero balance, one debit — never on timing.

go deeper

for a junior

Be ready to say what a negative test is: a check that a disallowed action stays disallowed, run on purpose from the attacker's side rather than the happy path. Know that it belongs in the same suite as functional tests.

for a middle

Expect to walk the three artifacts out loud — narrative, must-not criterion with an observable, executable test — and to name the assertions you would write for a concrete case such as double-spending a stored balance.

for a senior

Show the discipline: adversary-position fixtures, invariants instead of timing for races, explicit numeric bounds for resource abuse, and assertions on unchanged state and audit evidence, not just on the refusal.

for a principal

Own the argument that these criteria are part of the definition of built, and be able to say which threats justify the translation cost and which are better served by a design change that removes the abuse case entirely.

## Three artifacts, not one An **abuse case** is a short narrative of what an adversary wants to achieve against a feature. It is prose and it is not checkable: "an attacker drains a gift card" names no position, no input and no observable. An **acceptance criterion** is that narrative inverted into a *must-not* statement with a measurable outcome attached, written in the same place as the feature's ordinary criteria so it is part of what "built" means. A **negative test** is the executable form of that criterion: a deterministic check that a specific disallowed outcome does not occur. Getting from the first to the third is the whole skill this question probes. ## The translation, step by step 1. **Name the adversary's position.** Anonymous caller, authenticated low-privilege tenant, insider with legitimate access, a compromised third party. This decides what credentials the test fixture holds, and it is the step most often skipped — a test run with an admin token proves nothing about a customer-level attack. 2. **Name the asset and the disallowed outcome.** Money moved, a record read, a service made unavailable, an audit entry missing. 3. **Invert into a must-not with an invariant.** "The attacker can X" becomes "the system must refuse X, and Y must remain true afterwards." The invariant is what makes it testable; the refusal alone is often too weak. 4. **Choose the observables.** The refusal semantics the design actually chose (a specific status and error code), the state that must be unchanged, and the evidence that must exist. 5. **Make it deterministic.** "Try it fast" is not a test. A race becomes N concurrent requests plus a counting invariant. A resource-exhaustion case becomes explicit numeric bounds. ## Worked example — money, authenticated customer Gift-card balance API. Abuse case: an authenticated customer fires two redemptions of one card simultaneously, hoping both settle. | Artifact | Content | | --- | --- | | Abuse case | "Redeem the same card twice at once and get double value." | | Criterion | "For a card with one unit of balance, any number of concurrent redemptions results in exactly one successful settlement; the others are refused with a defined error; the ledger contains exactly one debit." | | Test | Seed the card. Fire ten concurrent redemptions as that customer. Assert exactly one success, final balance zero, exactly one ledger debit, refusals carry the expected error. Repeat the whole test several times. | Note what the assertion is *about*: the invariant (one debit), not the timing. A concurrency negative test that asserts "the second call was slower" is a flake generator. ## Worked example — availability, authenticated claimant An insurance claims intake accepts attachments. Abuse case: a claimant uploads a 900 MB file, or a small archive that expands to enormous size when unpacked, and the intake service stops answering anyone. The criterion has to carry numbers, because "the service should survive" is not an assertion: reject above the stated size limit *before* the body is buffered; refuse an archive nested deeper than the stated depth or exceeding the stated expansion ratio; return the refusal within the stated wall-clock bound. The test then asserts the refusal, and — the part people forget — that a normal request still succeeds immediately afterwards, which is the actual availability property. ## What the test asserts beyond "it failed" A refusal on its own is weak evidence. Add the two companions: **state unchanged** (no debit, no record opened, no row written) and **evidence present** (a denial event exists with enough detail to investigate). For a document e-signature service, the abuse case "open an envelope addressed to someone else by walking the signing link" produces both an assertion that the unauthorised holder is refused *and* an assertion that the refusal itself was audited — because the asset there is non-repudiation and audit truth, not just confidentiality. ## Why this is a threat-model deliverable A threat that produced a paragraph in a document regresses silently the first time someone refactors the guard. A threat that produced a criterion and a test regresses loudly, on every release, in every environment the suite runs against. That durability is the argument for spending the extra half hour on the translation, and it is why negative tests are treated as an output of the modeling session rather than a favour the QA team might do later. ## Failure modes to name in an interview Writing "tester attempts double redemption" as the criterion (a procedure, not an outcome); running the attempt with elevated credentials; asserting only "not 200"; asserting on error-message text; and assuming the positive test passing implies the abuse fails, which it never does.

  • "An attacker drains the gift-card balance" — why is that not yet an acceptance criterion?
    It names no adversary position, no input and no observable, so nobody can tell whether the system passes it. A criterion has to say who is attempting it, what they send, what the system must refuse, and what must still be true afterwards — for example exactly one debit on the ledger. Until those exist there is nothing a test can assert on, and two engineers will implement two different checks.
  • How do you write a negative test for an availability abuse case such as an oversized or deeply nested upload?
    Put numbers in the criterion: a maximum accepted size, a maximum archive depth or expansion ratio, and a wall-clock bound for the refusal. Assert that the oversized input is rejected before the service buffers or expands it, and assert that a normal request still succeeds straight afterwards. That last assertion is what actually tests availability; the refusal alone only tests validation.
  • What makes a derived negative test a durable deliverable rather than a one-off check?
    It asserts an invariant of the design, not a symptom of one build, so it stays meaningful after refactors. A test like "a member cannot read another member's points ledger" is written once and then run against every environment the service is deployed to, which is what catches the configuration drift that silently re-opens a threat you already closed in code.

A user story says who the door must let through; the abuse case says who must be turned away. The negative test is the turnstile test for the second door, and it also checks the camera recorded the attempt.

saying these in an interview costs you the question

  • Writes "tester tries to hack the endpoint" as the criterion
  • Asserts only that the request did not succeed
  • Runs the attempt with admin credentials, not the attacker's role
  • Asserts on timing instead of a counting invariant
  • Assumes a passing positive test implies the abuse fails
  • Leaves the criterion in the threat model instead of beside the feature

context

open as a page

How do you keep a negative test from passing without proving the control works?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Pair every negative assertion with a positive control in the same run, assert the exact refusal the design chose plus unchanged state and a denial record, and confirm the test fails when the guard is disabled.

open as a page

How do you handle the steps of an abuse case that no automated test can cover?

level: principalimportance: nice to knowfreq 28%

basics

~10 s

Split the narrative into steps, automate the technical invariants, and keep each human-dependent step visible in the model with a stated reason and a named substitute exercise instead of quietly deleting it.

open as a page