skip to content

Non-Functional Testing

Testing what a functional check never touches: speed under load, behaviour when a dependency dies, other platforms and locales, and refusing abuse. Most candidates only ever exercise the happy path.

on this pageshow

questions

page 1 of 2

How do you derive abuse and misuse cases from a documented happy-path flow?

level: juniorimportance: must knowfreq 64%

answer

  1. Success is only half the specification
  2. Walk the steps, not the screens
  3. Ask what an actor does instead
  4. Omit, repeat, reorder, substitute, exceed
  5. Each case names its expected refusal

basics

~20 s

Walk each step of the happy path and ask what an actor could do instead: skip it, repeat it, reorder it, supply another actor's identifier, or exceed a limit. Each answer becomes a case the system must refuse.

solid answer

~50 s

Write the happy path as numbered steps, recording for each step who performs it, what state it requires and what it changes. Then apply a fixed set of operators to every step: omit it, repeat it, reorder it against its neighbours, substitute the identifier with one from another actor or account, substitute the actor with an unauthenticated or lower-privileged one, exceed the expected quantity, act outside the allowed time window, and malform the payload. Each combination is a candidate case. Prune to the ones a stated rule says must be refused, plus the ones nobody can answer — those are specification gaps found early. Finish each surviving case with a precise expected refusal: what the caller sees, that no state changed, and whether the attempt is recorded. "It should not work" is not an expected result.

code

pseudocode · 13 lines
pseudocode
step = {id: 5, actor: "planner", action: "publish_draft", requires: ["draft_locked"], changes: ["timetable_version"]}

candidates = []
candidates.add(omit_precondition(step, "draft_locked"))
candidates.add(repeat(step, times=2))
candidates.add(reorder(step, before="run_clash_check"))
candidates.add(substitute_identifier(step, scope="other_school"))
candidates.add(substitute_actor(step, role="teacher"))
candidates.add(substitute_actor(step, role="anonymous"))
candidates.add(shift_time(step, to="after_term_start"))

for candidate in candidates:
    candidate.expected = {outcome: "refused", state_change: "none", audited: true}

go deeper

for a junior

Be ready to take a short happy path an interviewer describes and produce misuse cases out loud. Naming the operators — omit, repeat, reorder, substitute, exceed, malform — is what gets you past this question.

for a middle

Explain the mechanics: how a step's preconditions and state changes tell you which operator applies, and how you turn a candidate into a case with a checkable expected refusal rather than a vague "should fail".

for a senior

Show judgement about pruning and about where each case runs. An interviewer expects you to say which candidates are worth keeping, and to insist that a case the interface makes impossible still runs at the service boundary.

for a principal

Own the practice, not the list. Be ready to say how derivation gets scheduled, who reviews the undefined outcomes, and how you keep a growing catalogue of refusal cases from outgrowing the team that maintains it.

### A happy path is a sequence, not a picture A happy path is the ordered list of steps that carries an actor from an intent to a completed outcome with every precondition met and every input well formed. It is the flow the specification was written around, the flow the demo shows, and usually the only flow the suite proves. An **abuse case** is a use of the same flow by an actor whose intent is hostile; a **misuse case** is a use by an ordinary actor whose intent is innocent but whose path was never envisaged. Both produce the same artefact for a tester: a case whose expected outcome is a **refusal**, and whose oracle - the thing that says the observed behaviour is wrong - is that the refusal did not happen, or happened incompletely. The derivation is mechanical enough to do in a room with a whiteboard, which is why it is asked at screening depth. You write the happy path as numbered steps. For each step you record who performs it, what state it requires, what it changes, and what it returns. Then you apply a fixed set of operators to that step and read off the cases. ### The operators - **Omit** the step. Jump straight to step five. Does the system enforce that steps one to four happened, or does it trust the caller to have walked the sequence? - **Repeat** the step. Send the same completed action twice, or twice concurrently. Is the second one refused, ignored, or applied again? - **Reorder** the steps. Perform the confirm before the reserve. This is where the most expensive defects on this leaf usually live, because ordering is an assumption nobody wrote down. - **Substitute the identifier.** Perform the step with an identifier that belongs to a different actor, a different account or a different scope than the one the session was issued for. - **Substitute the actor.** Perform the step as an unauthenticated caller, as a lower-privileged role, and as an equally-privileged actor from a different scope. - **Exceed the quantity.** Ask for ten thousand where the flow expects three; upload past the stated limit; request a page far beyond the end. - **Stretch the timing.** Perform the step after the window closed, after the actor's access was removed, or while another actor performs the same step on the same record. - **Malform the payload.** Omit a required field, add an unexpected one, send the wrong type, send an empty collection where one item was assumed. Each operator applied to each step yields a candidate. You then prune: keep the candidates where a stated rule says the system must refuse, plus the ones where nobody can tell you what should happen - those are the valuable ones, because an undefined refusal is a specification gap discovered before it is a defect. ### A worked derivation Take a school timetable planner. The happy path for publishing a term timetable is: (1) a planner opens the draft term, (2) assigns each class to a room and a slot, (3) runs the clash check, (4) locks the draft, (5) publishes it to teachers and students. Applying the operators to step 5 alone gives: publish without having run the clash check (omit); publish twice within the same second (repeat); publish a draft that was never locked (reorder); publish a draft belonging to a different school (substitute the identifier); publish as a teacher rather than a planner (substitute the actor); publish after the term start date has passed (timing). A team of four working through five steps this way generates on the order of thirty candidates in an hour and prunes to perhaps eighteen worth keeping. The reorder case is the instructive one here: the planner had an **ordering assumption** - that a draft reaching publish must already have passed the clash check, because the interface only offers the publish control after the check completes. The service itself never re-checked. A request assembled out of order published a timetable with two classes in the same room, and the interface-level assumption was invisible until someone wrote the case that violated it. ### What the derived case must say A derived case is not finished when it names the misuse. It has to state the expected refusal precisely enough to be checkable: the outcome the caller sees, the fact that no state changed, and whether the attempt should leave a record. "It should not work" is not an expected result - it passes when the endpoint has been renamed and the request now fails for an unrelated reason. Two distinctions keep the catalogue honest. First, **must-refuse versus cannot-happen**: if a case is impossible only because the interface hides the control, it is a must-refuse case at the service boundary and must be exercised there, not through the interface that hides it. Second, **refuse versus tolerate**: a repeated submission may legitimately be absorbed rather than rejected, and the case must say which, because a test asserting rejection against a system that correctly absorbs it is a false alarm that trains people to ignore the suite.

  • What is the difference between an abuse case and a misuse case, and does it change the test you write?
    An abuse case assumes a hostile actor deliberately working against the system; a misuse case assumes an ordinary actor taking a path nobody envisaged, such as opening two windows and submitting the same form twice. The distinction matters for prioritisation and for who reviews the list, but the artefact is identical: a case whose expected outcome is a refusal or a safe absorption, with the state change asserted either way.
  • You derive a candidate case and nobody on the team can say what the system should do. What do you do with it?
    Treat it as the most valuable candidate in the batch. An undefined outcome is a specification gap, and it is far cheaper to close in a five-minute conversation than after a release. Take it to whoever owns the rule, get a decision recorded, and then write the case against the decision. Do not guess an expected result and encode your guess as the oracle — that bakes an unreviewed opinion into the suite.
  • How do you stop this technique generating an unbounded pile of cases?
    The operators produce candidates, not tests. Prune by keeping the cases a stated rule requires the system to refuse, the ones with no stated answer, and the ones whose failure would be expensive or irreversible. Drop the ones already covered by an equivalent case at the same boundary. A five-step flow should yield a couple of dozen candidates and roughly two-thirds that many kept cases, not hundreds.

A building inspector does not only check that the front door opens; they try every window, the fire exit from the wrong side, and the door left propped while the alarm is armed.

saying these in an interview costs you the question

  • Treating "it should not work" as a complete expected result
  • Deriving cases only from screens instead of from steps
  • Assuming the interface prevents an out-of-order request
  • Listing misuse cases but never asserting state was unchanged
  • Only varying the input, never the actor or the ordering

context

open as a page

What is volume testing, and how does it differ from a load test that raises request rate?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Volume testing keeps the request rate fixed and grows the stored dataset instead: more rows, longer collections, larger files. It exposes defects that appear only at size, such as unbounded queries, memory that scales with result size, and jobs that outgrow their window.

open as a page

What is the difference between an internationalization defect and a translation defect?

level: juniorimportance: must knowfreq 46%

basics

~20 s

An internationalization defect is in the product code: a string that was never extracted, a hardcoded date format, a container that cannot hold a longer word. A translation defect is in the text itself, wrong or missing wording.

open as a page

In a performance run, what are ramp-up, steady state and warm-up, and which window do you measure?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Ramp-up is the stretch while offered load climbs to the target level; steady state is where that load is held constant; warm-up is the early portion where caches, pools and compiled code settle. Report the steady state only.

open as a page

What is a support matrix in compatibility testing, and what does one cell of it commit you to?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A support matrix is the published grid of platform combinations — operating system, version, engine, device class — where a product promises to work. Each listed cell is a commitment: a defect that reproduces there is one you owe a fix.

open as a page

Before you fail a dependency in a resilience test, what must you write down first, and why?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Write the expected degraded behaviour first: for each affected feature, what the caller gets, what the interface shows, which side effects stay allowed, and what must never happen. Without that written oracle, anything short of a crash looks like a pass.

open as a page

How do you build an authorization matrix as a test artefact, and what do its negative cells assert?

level: middleimportance: must knowfreq 57%

basics

~20 s

Rows are actor classes, columns are operations, and each cell says allow or deny. A parameterized suite iterates the whole table, and every deny cell asserts a refusal, an unchanged state, no information leak and the required record of the attempt.

open as a page

How do you tell queueing delay from service time when a performance run recorded only end-to-end duration?

level: middleimportance: must knowfreq 58%

basics

~20 s

Service time is what one request costs with nothing competing for the resource; queueing delay is what waiting adds. Measure at a concurrency of one, then watch duration grow as concurrency rises — the growth is wait, not work.

open as a page

In a performance run, what does a gap between offered and achieved throughput tell you?

level: middleimportance: must knowfreq 62%

basics

~20 s

Offered throughput is the demand the run applied; achieved throughput is the work that actually completed and was counted. A gap means requests were refused, expired or never finished, so the completed-work figure describes only part of the intended load.

open as a page

What is the difference between an open and a closed workload model in a load test?

level: middleimportance: must knowfreq 54%

basics

~20 s

In an open model the generator sends requests at a chosen arrival rate regardless of how the system responds. In a closed model a fixed population of simulated users each waits for a response, then thinks, then sends again, so slowness throttles the load.

open as a page

Past a system's intended load limit, what must "degrades gracefully" become before a performance run can assert it?

level: middleimportance: must knowfreq 62%

basics

~20 s

"Degrades gracefully" must become outcome rules written before the run: which requests are still completed and inside what response-time bound, which may be refused, how fast a refusal must arrive, and which results fail outright.

open as a page

Before a performance run starts, what must its pass rule fix for the run's verdict to mean anything?

level: middleimportance: must knowfreq 70%

basics

~20 s

Fix five things before the run starts: which measurement, at which point in its distribution, over which window, under which workload, and read from which vantage point. Anything left open is chosen afterwards to suit the numbers.

open as a page

Why can a performance run's per-interval 95th-percentile times not be averaged into one figure for the whole run?

level: middleimportance: must knowfreq 62%

basics

~20 s

Percentiles are ranks over a set of samples, not quantities that add or average. The mean of per-interval figures weights a quiet minute like a busy one and throws away each interval's shape. Merge raw samples or bucket counts instead.

open as a page

What has to be held constant for two performance runs to be comparable at all?

level: middleimportance: must knowfreq 62%

basics

~20 s

Comparable performance runs hold the same build, the same dataset size and shape, the same environment and its neighbours, the same warm state and the same applied workload. When one of those differs, a change in the numbers cannot be attributed.

open as a page

How do you decide how many hours a sustained load run must hold?

level: middleimportance: must knowfreq 58%

basics

~10 s

Derive it from the build-up you expect to surface: the hold must be long enough for that accumulation to clear measurement noise and cover a readable fraction of the headroom. Calendar convenience decides nothing.

open as a page

After a surge in demand subsides, what must be true before you call the system recovered?

level: middleimportance: must knowfreq 55%

basics

~20 s

Recovery needs more than response times falling. The backlog must drain back to its pre-surge depth, resources must be released rather than settling higher, accepted work must still complete, and that state must hold across a defined observation window.

open as a page

How do you parameterise a sudden surge in arrival rate so a performance run's result is attributable to it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Fix three numbers before the run: the settled baseline arrival rate, the multiple applied to it, and the transition time over which the rate rises. Vary one per run, and hold the peak long enough to read a result.

open as a page

In a performance run, why is a reply that arrives with a success status not necessarily a success?

level: juniorimportance: should knowfreq 55%

basics

~20 s

A success status only says something answered the request. The body may be a rendered error page, a truncated document, or an empty result where a data record was expected, so a run reading only the status overstates how much work really succeeded.

open as a page

In a load run held steady for hours, why assert on the growth trend rather than the peak?

level: juniorimportance: should knowfreq 47%

basics

~20 s

A peak says how high a figure got; accumulation is a direction. Measure retained memory, open handles, cached entries and disk use as a slope across the hold, then compare that slope with a period expected to be level.

open as a page

Why does uniformly generated seed data understate how a system behaves at volume?

level: middleimportance: should knowfreq 54%

basics

~20 s

Generated rows tend to be uniform: equal group sizes, same-width text, no gaps, one format. Real data is skewed, with hot keys, long tails and rare variants. Uniform seeds hide the worst case, so the run measures an average production never has.

open as a page

What is pseudo-localization, and which defects does it catch before any translation exists?

level: middleimportance: should knowfreq 33%

basics

~20 s

Pseudo-localization replaces every extracted string with a mechanically accented, padded, bracketed version of itself and runs the product against that fake locale. Anything still in plain source text was never extracted, and padded text exposes truncation.

open as a page

How do you confirm a fixed pool of connections or workers, not the resource behind it, caps a performance run's throughput?

level: middleimportance: should knowfreq 50%

basics

~20 s

A fixed pool caps throughput near its slot count divided by mean service time, so throughput flattens while duration rises with offered concurrency. Change the slot count: if the ceiling moves in proportion, the pool was the cap.

open as a page

How does an overload run assert that refusing work is a correct outcome rather than a defect?

level: middleimportance: should knowfreq 44%

basics

~20 s

By asserting three things about every rejection: it comes back quickly rather than after a full wait, it reaches the caller as an explicit refusal it can act on, and it never counts as a success.

open as a page

A performance run passed on average response time although many requests were far slower. How do you write the pass rule so that cannot happen?

level: middleimportance: should knowfreq 52%

basics

~20 s

Attach the bound to a stated point in the distribution rather than a central measure, pair it with an outer cap on the worst reported interval, and say how many intervals may breach before the run fails.

open as a page

In a performance run mixing several operation types, how can one slow operation vanish from the aggregate 95th percentile?

level: middleimportance: should knowfreq 55%

basics

~20 s

An aggregate percentile ranks every request together, so weight follows request share, not importance. An operation carrying under 5% of requests can be entirely above the 95% cut and never move it. Report a distribution per operation type.

open as a page

Before you call a difference between two performance runs a regression, what must you measure first?

level: middleimportance: should knowfreq 52%

basics

~10 s

Measure the run-to-run spread first: repeat one unchanged configuration several times and record how much the compared figure varies on its own. A difference smaller than that spread is not evidence of anything.

open as a page

When does an emulated device environment give a false pass that real hardware would catch?

level: middleimportance: should knowfreq 57%

basics

~20 s

Emulated environments borrow the host's processor, memory, storage speed and network, and they run a clean reference build of the platform. So they miss slow or thermally throttled hardware, vendor-modified system builds, real sensors and cameras, memory-pressure eviction, and real network transitions.

open as a page

How do you report a defect that reproduces on only one platform in your support matrix?

level: middleimportance: should knowfreq 53%

basics

~20 s

State the exact platform coordinates where it fails and the cells where the same steps passed, isolate one axis at a time so the report names the differing variable, give a reproduction rate, and name the failing cell's tier.

open as a page

A failover test passes. What does verifying failback to the primary catch that the failover leg does not?

level: middleimportance: should knowfreq 46%

basics

~20 s

Failback catches everything the switch itself created: writes accepted on the standby that must reach the primary, a replication path that has to reverse, stale reads right after the return, clients still pinned to the old node, and configuration edited during the incident.

open as a page

How do you design tests that catch tenant crossing and privilege escalation in a multi-tenant system?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Seed two fully populated accounts with distinguishable data, then replay every recorded happy-path request with a session from the other account and with lower-privileged roles. Each must be refused, change no state, and leak nothing about the other account.

open as a page

showing 1–30 of 60