Non-Functional Testing
Testing what a functional check never touches: speed under load, behaviour when a dependency dies, other platforms and locales, and refusing abuse. Most candidates only ever exercise the happy path.
on this pageshowhide
explore
- Performance Workloads40 questions
- Load Modeling & Measurement5 questions
- Run Shapes13 questions
- Reading Results22 questions
- Resilience & Recovery3 questions
- Platform Compatibility5 questions
- Locale Readiness4 questions
- Abuse & Negative Cases4 questions
- Data Volume Growth4 questions
- AI & Data Scientistrole
- AI Engineerrole
- Backend Developerrole
- Data Engineerrole
- Frontend Developerrole
- Full Stack Developerrole
- Game Developerrole
- Java Backend Developerrole
- Java SDETrole
- Kotlin Backend Developerrole
- MLOps Engineerrole
- Machine Learning Engineerrole
- QA Engineerrole
- Software Architectrole
- iOS Developerrole
questions
page 2 of 2How do you design a volume run that reveals super-linear cost growth before production reaches that size?
basics
~20 sRun the same fixed workload against several dataset sizes — today's, four times, sixteen times — and compare the shape of the curve rather than one pass or fail. Anything whose cost rises faster than the data is the finding.
A monthly utility billing run double-charged accounts on the night clocks shifted back. How would you test its timezone and calendar handling?
basics
~10 sDrive the run from an injected clock, not the machine clock, and enumerate boundary dates deliberately: the repeated local hour, the missing one, month ends, fractional offsets. Assert one charge per account per period.
Client-side timings exceed the system's own recorded durations for the same requests in a performance run. Which path segments explain the gap?
basics
~20 sA client's clock and the system's own clock cover different intervals. A gap that does not change with load is fixed transport cost outside the handler. A gap that widens with load is waiting in a segment nobody times.
How do you attribute a performance run's slow response tail to one step on the request path when each step's own average looks acceptable?
basics
~20 sPer-step averages hide a step that is slow occasionally or runs many times per request. Compare each step's share of total time in the slowest requests against median ones, then confirm by neutralising the suspect.
In a performance run, how do you count requests that expired, were reset, or were never sent?
basics
~20 sCount them as outcomes, not as missing data. Each stays inside the issued total, lands in a named failure category, and the report states what time was recorded for it. Deleting them describes only the requests the system managed to serve.
Why can a load generator's reported p99 be far better than the latency real users would have seen?
basics
~20 sBecause a generator that waits for each response stops sending while the system is stalled, so the requests that would have been slowest are never issued and never measured. This is coordinated omission: the missing samples are exactly the bad ones.
How do you tell whether the load generator, not the system under test, is the bottleneck in a run?
basics
~20 sCompare achieved load against offered load and instrument the generator hosts: processor use, memory pressure, port and descriptor exhaustion, network saturation. If the achieved rate falls short while the system under test sits idle, the rig is the limit.
Which results fail an overload run outright, whatever its response times show?
basics
~20 sThree results fail regardless of timing: work applied in part and left inconsistent, a request accepted and lost with nobody told, and a queue or in-flight set that grows with no ceiling. These are damage, not slowness.
Under overload, two systems both show a falling success rate. Which observations tell shedding apart from collapse?
basics
~20 sShedding holds completed work near the sustainable level, answers rejections in milliseconds, keeps survivors near normal speed, and holds backlog and memory flat. Collapse shows completed work falling, everyone waiting, and a backlog that keeps growing.
A performance run finished with no pass rule agreed beforehand. What can its report honestly conclude?
basics
~20 sDescription only: what was applied, where timings were taken, and where the numbers sat. Not a pass or a fail - a bound chosen once results are visible is fitted to them. Its real output is the next run's rule.
Why can one latency pass rule not judge a long steady hold, a sudden step in arrival rate and a deliberate overload run?
basics
~20 sEach shape asks a different question. A long hold must still meet the bound in its final intervals; an abrupt increase in arrival rate is judged on recovery time; a run past the intended rate is judged on refusal, not latency.
A performance run's response times form two distinct clusters. What does one 95th-percentile figure get wrong?
basics
~20 sIt implies one population with one centre. With two clusters the figure lands in whichever cluster holds the rank, describes only that one, and moves with the proportion between them rather than with either cluster's own speed.
A request crosses four services, each with a known 99th-percentile latency. Why is the end-to-end 99th percentile not their sum?
basics
~20 sSumming per-hop 99th percentiles prices a request in which all four hops were simultaneously in their own slowest one percent, which is rare — so the sum usually overstates. Correlated hops can make it understate. Measure whole-path time per request instead.
How do you set a regression threshold for performance runs from measured noise rather than a round number?
basics
~20 sDerive it from the measured run-to-run spread: set the alarm level a multiple above the spread of repeats, per metric, then check that size against what change actually matters. A round percentage chosen by habit is either noise or nothing.
Response times climb across a long steady-load run while per-request resource use stays flat - what do you check?
basics
~20 sCheck whether the climb advances with work done or with elapsed time. Replot it against cumulative completed work, run a low-rate control, probe a path the run never writes to, and restart the process keeping its data.
When extra capacity arrives minutes after a surge, how do you keep a performance run from hiding that delay?
basics
~20 sStart at the size the system would reach under baseline demand, never pre-raised to peak, and timestamp two series: offered demand and capacity actually serving. The interval between them is the exposure window, and what happened inside it is the result.
How do you prove across a demand surge that no unit of work was lost or done twice?
basics
~20 sTag every submitted unit with a unique identifier before demand rises, then reconcile after the observation window: submitted equals completed plus refused plus pending. Separately compare distinct completed identifiers against total completions, because totals alone hide loss and repetition together.
How do you verify that saved user data survives both an upgrade and a rollback to the previous release?
basics
~20 sKeep saved data written by every supported release, upgrade each one and assert the content afterwards, then reverse it: install the previous release over the upgraded data and check the older build reads it without discarding what it cannot interpret.
How do you test that no data is corrupted when a service is killed abruptly mid-write?
basics
~20 sKill deliberately at the boundaries between irreversible steps, restart, then judge by invariants over the stored data: acknowledged operations minus applied ones must be empty, nothing applied twice, no orphaned rows, and cross-store totals reconcile.
How would you build a workload model from real production traffic and keep it representative over time?
basics
~20 sDerive the transaction mix, arrival shape and data profile from observed traffic rather than opinion, state the reference period, and keep the model a versioned, owned artefact that is re-derived on a schedule instead of ageing into fiction.
When is adopting a new performance reference run legitimate rather than moving the goalposts?
basics
~20 sIt is legitimate when the shift in level has an identified, intentional cause, was reproduced across repeats, and still meets the obligation the team owes. It is a moved goalpost when the reason is that meeting the old figure became inconvenient.
How do you decide which platform versions to drop from a support matrix, and what do you owe users when you drop one?
basics
~20 sWeigh your own usage telemetry and the revenue behind it against the cell's cost, honour stated commitments and the vendor's own support dates, then drop with announced lead time, a pinned last-supported release, and a check that the tail migrated.
How do you test that a revoked, expired or tampered credential is actually refused?
basics
~20 sTreat the three as separate cases. Expiry needs a controllable clock rather than a sleep, revocation needs a stated bound on how long access may survive withdrawal, and tampering means altering one field and asserting refusal rather than silent acceptance.
How do you test locale-sensitive sorting and matching in a customer or account list?
basics
~20 sBuild a golden ordered fixture per locale containing accents, case variation, a locale-specific digraph and punctuation, and assert the full order. Then assert every layer uses the same comparator, and that composed and decomposed spellings compare equal.
How do you verify a one-off data backfill will finish inside its window before running it in production?
basics
~20 sRehearse it against a volume-representative copy, measure the rate chunk by chunk instead of only the total, and extrapolate with the non-linear effects stated. Then set a go/no-go checkpoint during the real run and practise the abort.
A slowdown reproduces only when requests overlap. How do you design the follow-up performance run that confirms it?
basics
~20 sVary overlap alone. Run the same arrival rate with few busy clients and then with many idle ones: if duration differs, overlap is the cause. Repeat with requests spread across distinct records to locate what is shared.
In a performance run, how much of each reply can you afford to verify in flight?
basics
~20 sOnly checks cheap enough to fit inside the applying side's own budget: a length floor, a required field, a marker. Deep parsing per reply steals the capacity that generates demand, so sample it instead and report what share was verified.
When a 99th percentile is computed from bucketed latency counts, what error do the bucket boundaries introduce?
basics
~20 sThe read-out locates the bucket holding the ranked request, not its time, so the figure carries that bucket's width as uncertainty. Interpolating inside assumes a spread the slow end does not have, and an unbounded top bucket reports nothing at all.
How do you decide whether a long sustained run should hit scheduled and rotational work or avoid it?
basics
~20 sDecide from what the hold must prove. Arrange the run to hit periodic work - scheduled jobs, credential rotation, cache expiry, index maintenance - when that interaction is the risk; otherwise disable it and record that the trend excludes it.
How do you establish that a second, larger request peak after an applied surge was self-inflicted rather than real demand?
basics
~20 sCompare what the load generator offered against what the service received: the offered rate is known and unchanged, so any excess was manufactured inside the system. Repeats of identifiers already submitted, arriving synchronised at a fixed delay, confirm it.
showing 31–60 of 60