skip to content

Non-Functional Testing

Testing what a functional check never touches: speed under load, behaviour when a dependency dies, other platforms and locales, and refusing abuse. Most candidates only ever exercise the happy path.

on this pageshow

questions

page 2 of 2

How do you design a volume run that reveals super-linear cost growth before production reaches that size?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Run the same fixed workload against several dataset sizes — today's, four times, sixteen times — and compare the shape of the curve rather than one pass or fail. Anything whose cost rises faster than the data is the finding.

open as a page

A monthly utility billing run double-charged accounts on the night clocks shifted back. How would you test its timezone and calendar handling?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Drive the run from an injected clock, not the machine clock, and enumerate boundary dates deliberately: the repeated local hour, the missing one, month ends, fractional offsets. Assert one charge per account per period.

open as a page

Client-side timings exceed the system's own recorded durations for the same requests in a performance run. Which path segments explain the gap?

level: seniorimportance: should knowfreq 43%

basics

~20 s

A client's clock and the system's own clock cover different intervals. A gap that does not change with load is fixed transport cost outside the handler. A gap that widens with load is waiting in a segment nobody times.

open as a page

How do you attribute a performance run's slow response tail to one step on the request path when each step's own average looks acceptable?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Per-step averages hide a step that is slow occasionally or runs many times per request. Compare each step's share of total time in the slowest requests against median ones, then confirm by neutralising the suspect.

open as a page

In a performance run, how do you count requests that expired, were reset, or were never sent?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Count them as outcomes, not as missing data. Each stays inside the issued total, lands in a named failure category, and the report states what time was recorded for it. Deleting them describes only the requests the system managed to serve.

open as a page

Why can a load generator's reported p99 be far better than the latency real users would have seen?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Because a generator that waits for each response stops sending while the system is stalled, so the requests that would have been slowest are never issued and never measured. This is coordinated omission: the missing samples are exactly the bad ones.

open as a page

How do you tell whether the load generator, not the system under test, is the bottleneck in a run?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compare achieved load against offered load and instrument the generator hosts: processor use, memory pressure, port and descriptor exhaustion, network saturation. If the achieved rate falls short while the system under test sits idle, the rig is the limit.

open as a page

Which results fail an overload run outright, whatever its response times show?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Three results fail regardless of timing: work applied in part and left inconsistent, a request accepted and lost with nobody told, and a queue or in-flight set that grows with no ceiling. These are damage, not slowness.

open as a page

Under overload, two systems both show a falling success rate. Which observations tell shedding apart from collapse?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Shedding holds completed work near the sustainable level, answers rejections in milliseconds, keeps survivors near normal speed, and holds backlog and memory flat. Collapse shows completed work falling, everyone waiting, and a backlog that keeps growing.

open as a page

A performance run finished with no pass rule agreed beforehand. What can its report honestly conclude?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Description only: what was applied, where timings were taken, and where the numbers sat. Not a pass or a fail - a bound chosen once results are visible is fitted to them. Its real output is the next run's rule.

open as a page

Why can one latency pass rule not judge a long steady hold, a sudden step in arrival rate and a deliberate overload run?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Each shape asks a different question. A long hold must still meet the bound in its final intervals; an abrupt increase in arrival rate is judged on recovery time; a run past the intended rate is judged on refusal, not latency.

open as a page

A performance run's response times form two distinct clusters. What does one 95th-percentile figure get wrong?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It implies one population with one centre. With two clusters the figure lands in whichever cluster holds the rank, describes only that one, and moves with the proportion between them rather than with either cluster's own speed.

open as a page

A request crosses four services, each with a known 99th-percentile latency. Why is the end-to-end 99th percentile not their sum?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Summing per-hop 99th percentiles prices a request in which all four hops were simultaneously in their own slowest one percent, which is rare — so the sum usually overstates. Correlated hops can make it understate. Measure whole-path time per request instead.

open as a page

How do you set a regression threshold for performance runs from measured noise rather than a round number?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Derive it from the measured run-to-run spread: set the alarm level a multiple above the spread of repeats, per metric, then check that size against what change actually matters. A round percentage chosen by habit is either noise or nothing.

open as a page

Response times climb across a long steady-load run while per-request resource use stays flat - what do you check?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Check whether the climb advances with work done or with elapsed time. Replot it against cumulative completed work, run a low-rate control, probe a path the run never writes to, and restart the process keeping its data.

open as a page

When extra capacity arrives minutes after a surge, how do you keep a performance run from hiding that delay?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Start at the size the system would reach under baseline demand, never pre-raised to peak, and timestamp two series: offered demand and capacity actually serving. The interval between them is the exposure window, and what happened inside it is the result.

open as a page

How do you prove across a demand surge that no unit of work was lost or done twice?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Tag every submitted unit with a unique identifier before demand rises, then reconcile after the observation window: submitted equals completed plus refused plus pending. Separately compare distinct completed identifiers against total completions, because totals alone hide loss and repetition together.

open as a page

How do you verify that saved user data survives both an upgrade and a rollback to the previous release?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Keep saved data written by every supported release, upgrade each one and assert the content afterwards, then reverse it: install the previous release over the upgraded data and check the older build reads it without discarding what it cannot interpret.

open as a page

How do you test that no data is corrupted when a service is killed abruptly mid-write?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Kill deliberately at the boundaries between irreversible steps, restart, then judge by invariants over the stored data: acknowledged operations minus applied ones must be empty, nothing applied twice, no orphaned rows, and cross-store totals reconcile.

open as a page

How would you build a workload model from real production traffic and keep it representative over time?

level: principalimportance: should knowfreq 42%

basics

~20 s

Derive the transaction mix, arrival shape and data profile from observed traffic rather than opinion, state the reference period, and keep the model a versioned, owned artefact that is re-derived on a schedule instead of ageing into fiction.

open as a page

When is adopting a new performance reference run legitimate rather than moving the goalposts?

level: principalimportance: should knowfreq 38%

basics

~20 s

It is legitimate when the shift in level has an identified, intentional cause, was reproduced across repeats, and still meets the obligation the team owes. It is a moved goalpost when the reason is that meeting the old figure became inconvenient.

open as a page

How do you decide which platform versions to drop from a support matrix, and what do you owe users when you drop one?

level: principalimportance: should knowfreq 43%

basics

~20 s

Weigh your own usage telemetry and the revenue behind it against the cell's cost, honour stated commitments and the vendor's own support dates, then drop with announced lead time, a pinned last-supported release, and a check that the tail migrated.

open as a page

How do you test that a revoked, expired or tampered credential is actually refused?

level: middleimportance: nice to knowfreq 27%

basics

~20 s

Treat the three as separate cases. Expiry needs a controllable clock rather than a sleep, revocation needs a stated bound on how long access may survive withdrawal, and tampering means altering one field and asserting refusal rather than silent acceptance.

open as a page

How do you test locale-sensitive sorting and matching in a customer or account list?

level: middleimportance: nice to knowfreq 17%

basics

~20 s

Build a golden ordered fixture per locale containing accents, case variation, a locale-specific digraph and punctuation, and assert the full order. Then assert every layer uses the same comparator, and that composed and decomposed spellings compare equal.

open as a page

How do you verify a one-off data backfill will finish inside its window before running it in production?

level: seniorimportance: nice to knowfreq 21%

basics

~20 s

Rehearse it against a volume-representative copy, measure the rate chunk by chunk instead of only the total, and extrapolate with the non-linear effects stated. Then set a go/no-go checkpoint during the real run and practise the abort.

open as a page

A slowdown reproduces only when requests overlap. How do you design the follow-up performance run that confirms it?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Vary overlap alone. Run the same arrival rate with few busy clients and then with many idle ones: if duration differs, overlap is the cause. Repeat with requests spread across distinct records to locate what is shared.

open as a page

In a performance run, how much of each reply can you afford to verify in flight?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Only checks cheap enough to fit inside the applying side's own budget: a length floor, a required field, a marker. Deep parsing per reply steals the capacity that generates demand, so sample it instead and report what share was verified.

open as a page

When a 99th percentile is computed from bucketed latency counts, what error do the bucket boundaries introduce?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

The read-out locates the bucket holding the ranked request, not its time, so the figure carries that bucket's width as uncertainty. Interpolating inside assumes a spread the slow end does not have, and an unbounded top bucket reports nothing at all.

open as a page

How do you decide whether a long sustained run should hit scheduled and rotational work or avoid it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Decide from what the hold must prove. Arrange the run to hit periodic work - scheduled jobs, credential rotation, cache expiry, index maintenance - when that interaction is the risk; otherwise disable it and record that the trend excludes it.

open as a page

How do you establish that a second, larger request peak after an applied surge was self-inflicted rather than real demand?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Compare what the load generator offered against what the service received: the offered rate is known and unchanged, so any excess was manufactured inside the system. Repeats of identifiers already submitted, arriving synchronised at a fixed delay, confirm it.

open as a page

showing 31–60 of 60