skip to content

Testing & Quality Disciplines

The vocabulary and practice of software quality: test levels and the pyramid, test doubles, test structure and independence, case-design techniques, coverage and its limits, TDD and BDD, and static analysis. Interviewers cover this because every role writes tests, and how you talk about them reveals how you work.

on this pageshow

explore

questions

651 · 13 sections

What are the three phases of the Arrange-Act-Assert test structure?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Arrange builds the fixture, inputs and collaborators the test needs. Act invokes the one behaviour under test and captures its result. Assert compares that outcome with the expected one. Given-When-Then names the same three phases.

open as a page

What is acceptance testing, and what decides whether a change passes it?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Acceptance testing judges a change against the stated acceptance criteria of the requirement it implements, not against the shape of the code. It passes when every criterion is demonstrably met, and those criteria are agreed before the work starts.

open as a page

What is an approval test, and how does it differ from a test with hand-written assertions?

level: juniorimportance: must knowfreq 46%
basics
~20 s

An approval test runs the code, writes the output to a received artefact, and compares it against a previously approved artefact. Nobody types the expected value: a human reviews the output once and approves it, and the test then guards that exact output.

open as a page

In boundary value analysis, which values do you test for a field that accepts 1 through 60?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Test the edges, not the middle: 0 and 1 low, 60 and 61 high. Defects hide where a comparison flips from accept to reject, so each limit gets a value on it and one just past.

open as a page

In a consumer-driven contract test, what does the recorded contract assert about the provider, and what does it deliberately not check?

level: juniorimportance: must knowfreq 42%
basics
~20 s

A consumer-driven contract records the requests one consumer sends and the response parts it actually reads, then asserts the provider can still produce them. It checks the shape of the boundary, not whether the provider's answers are correct.

open as a page

What does "emergent design" mean in test-driven development, and where does the structure come from?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Emergent design means the structure of the code is an output of the cycle, not a plan made before it. Each restructuring step reshapes working code, so the design arrives gradually instead of being guessed up front.

open as a page

In TDD, what is the difference between outside-in and inside-out development?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Outside-in TDD starts at an outer boundary and invents collaborator roles as they are needed, standing them in until they are built. Inside-out TDD starts with domain behaviour built from real objects and wires the outer layers in last.

open as a page

Why must a test-driven test be watched failing before you write the code that passes it?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Watching the test fail proves the assertion can actually detect the behavior's absence. A test that has never been red may be mis-wired, never executed, or trivially true, and would report green forever while guarding nothing.

open as a page

Why run the test suite after each small refactoring move instead of once at the end?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Because a green run after every move tells you exactly which move changed behaviour. Run once at the end and a red suite only says something in the last hour was wrong, so you debug instead of undoing.

open as a page

What is test-first development, and how does it differ from writing tests after the code?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Test-first means writing an executable test for behaviour that does not exist yet, then writing code to satisfy it. Test-after writes tests against finished code, so the cases are shaped by what the code already does.

open as a page

In double-loop TDD, what are the outer and inner loops, and which turns green last?

level: juniorimportance: must knowfreq 58%
basics
~20 s

The outer loop is a failing acceptance scenario written in the user's language; the inner loop is the fast unit cycle beneath it. The inner loop goes green many times over; the outer scenario goes green last, when the behaviour is complete.

open as a page

What is a step definition in a scenario suite, and what binds it to a line of a scenario?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A step definition is a function holding the automation code for one scenario step. The runner matches the step's plain-language text against a pattern registered with that function, converts any captured values into arguments, and calls it.

open as a page

What does a Background section in a feature file do to the scenarios below it?

level: juniorimportance: must knowfreq 66%
basics
~10 s

A Background holds context steps that run again before every scenario in that file, so shared setup is written once. Every scenario inherits those steps, whether or not it needs them.

open as a page

In a behaviour scenario, what do the Given, When and Then steps each describe?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Given states the context that already holds before the behaviour under test. When names the single event that triggers it. Then states the observable outcome that must follow. All three describe behaviour, not interface mechanics.

open as a page

What makes a scenario report generated from the last run 'living documentation'?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The document is generated from scenarios that actually executed, so it describes only behaviour the suite exercised. When the system changes and the scenario does not, the scenario fails and the page visibly breaks instead of quietly going stale.

open as a page

What is a static-analysis baseline, and why create one when adopting an analyser on an existing codebase?

level: juniorimportance: must knowfreq 60%
basics
~20 s

A static-analysis baseline is a recorded snapshot of the findings an analyser already reports on existing code. The gate ignores those and fails only on findings outside the snapshot, so an old codebase can adopt analysis without a mass cleanup first.

open as a page

Why run the same static analysis rule in the editor, at pre-commit, and in the build rather than only in the build?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Each placement buys something different: the editor gives feedback in seconds while you type, the pre-commit check keeps a whole commit clean, and the build is the only placement that runs for everyone and can actually block a merge.

open as a page

What is the difference between a formatting lint rule and a correctness lint rule?

level: juniorimportance: must knowfreq 74%
basics
~20 s

A formatting rule governs how code looks - spacing, indentation, ordering - and never changes behaviour. A correctness rule flags a shape that is likely a bug: a discarded return value, an unreachable branch, a resource never released.

open as a page

In security static analysis, what are a source, a sink and a sanitizer in taint tracking?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A source is where untrusted input enters the program, a sink is a sensitive operation that must not receive it, and a sanitizer neutralises the data. A finding is a source-to-sink path with no sanitizer on it.

open as a page

What does a static type checker prove before a program runs, and which bugs does it never catch?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A static type checker proves, without running the program, that every operation receives a value whose declared or inferred type it accepts. It rules out type-mismatch errors on all code paths, but says nothing about whether the logic is right.

open as a page

What are entry criteria for a test cycle, and what do they prevent?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Entry criteria are the conditions that must hold before a test cycle starts: a deployed build with a known change list, a healthy environment, seeded data, and agreed acceptance criteria. They prevent a cycle from burning its budget on setup failures.

open as a page

What is the difference between verification and validation in a software lifecycle?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Verification asks whether the product was built according to its specification. Validation asks whether that specification described the right thing to build. One checks conformance to a written statement, the other checks fitness for a real need.

open as a page

In a test cycle status report, why is a blocked case counted separately from a failed one?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A failed case ran and gave the wrong result, so it is evidence about the product. A blocked case never ran, because something stopped it. Merging the two hides untested scope and blames the product for an environment problem.

open as a page

What is the difference between confirmation testing and regression testing after a defect fix?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Confirmation testing re-runs the case that exposed the defect, to prove the fix works. Regression testing re-runs other cases around the change, to prove the fix broke nothing that previously passed. Both are needed after a fix.

open as a page

What is risk-based test selection, and how do likelihood and impact rank what gets tested first?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Risk-based test selection ranks each area by how likely it is to fail and how badly a failure would hurt, then spends the deepest testing on the highest-ranked areas and the least on the lowest.

open as a page

What must a defect report contain for someone else to reproduce and fix it?

level: juniorimportance: must knowfreq 74%
basics
~10 s

An unambiguous one-line title, the exact build identifier and environment, numbered preconditions and steps, the expected result and the actual result written separately, and evidence such as logs, screenshots or traces.

open as a page

What is defect density, and why does the denominator you choose change the answer?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Defect density is a count of confirmed defects divided by a size unit — a thousand lines of code, a functional size unit, a module, a feature. Change the unit and the same product scores differently.

open as a page

What is the difference between a defect's origin phase and its escape phase?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Origin phase is where the defect was introduced - the requirement, design, code, data or configuration work that created it. Escape phase is the last check that should have caught it and did not. Every defect has both.

open as a page

In a tracked defect's explanation, how do the symptom, the condition that produced it, and the weakness that let it reach a user differ?

level: juniorimportance: must knowfreq 64%
basics
~20 s

The symptom is what was observed. The condition that produced it is the specific state or input that made the code fail that time. The weakness is the missing guard or assumption that allowed the whole family of failures.

open as a page

What is the difference between a defect's severity and its priority?

level: juniorimportance: must knowfreq 84%
basics
~20 s

Severity measures a defect's technical impact — how badly the product's behaviour breaks. Priority measures business urgency — how soon it should be fixed relative to other work. Different people set them, and the two values move independently.

open as a page

How do you derive abuse and misuse cases from a documented happy-path flow?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Walk each step of the happy path and ask what an actor could do instead: skip it, repeat it, reorder it, supply another actor's identifier, or exceed a limit. Each answer becomes a case the system must refuse.

open as a page

What is volume testing, and how does it differ from a load test that raises request rate?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Volume testing keeps the request rate fixed and grows the stored dataset instead: more rows, longer collections, larger files. It exposes defects that appear only at size, such as unbounded queries, memory that scales with result size, and jobs that outgrow their window.

open as a page

What is the difference between an internationalization defect and a translation defect?

level: juniorimportance: must knowfreq 46%
basics
~20 s

An internationalization defect is in the product code: a string that was never extracted, a hardcoded date format, a container that cannot hold a longer word. A translation defect is in the text itself, wrong or missing wording.

open as a page

In a performance run, what are ramp-up, steady state and warm-up, and which window do you measure?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Ramp-up is the stretch while offered load climbs to the target level; steady state is where that load is held constant; warm-up is the early portion where caches, pools and compiled code settle. Report the steady state only.

open as a page

What is a support matrix in compatibility testing, and what does one cell of it commit you to?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A support matrix is the published grid of platform combinations — operating system, version, engine, device class — where a product promises to work. Each listed cell is a commitment: a defect that reproduces there is one you owe a fix.

open as a page

Why narrow a defect to the shortest reliable reproduction before you report it?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A short reproduction proves which conditions actually cause the failure. Every step you can delete without losing the failure is noise that slows diagnosis, invites a cannot-reproduce close, and hides the real trigger from whoever has to fix it.

open as a page

What is a bug bash and who takes part in one?

level: juniorimportance: must knowfreq 56%
basics
~20 s

A bug bash is a short, time-boxed event where many people — developers, support, product and testers — hunt defects in one shared build at once, each covering an assigned area and filing into one channel.

open as a page

What is a test charter in exploratory testing, and what does its three-part shape name?

level: juniorimportance: must knowfreq 74%
basics
~20 s

A test charter is a one- or two-sentence mission for a single exploratory session. The common shape names three things: what to explore, what to explore it with, and what information the session should discover.

open as a page

What is a test oracle, and what makes one fallible?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A test oracle is whatever you use to decide that an observed behaviour is wrong: a written specification, a standard, a comparable product, an earlier version, or your own expectation. Each is an imperfect model, so it can be wrong too.

open as a page

What is a timeboxed test session in session-based test management?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A session is one uninterrupted block of testing, commonly 60, 90 or 120 minutes, spent on a single charter and written up as one reviewable record. The session, not the test case, is the unit of work you plan and count.

open as a page

Machine-drafted cases triple a regression suite's case count in a week. What has that growth not proven?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Case count measures how much was written, not how much can be detected. Drafted cases usually re-walk journeys existing cases already walk and check outcomes already checked, so the set of failures the suite can catch may not grow.

open as a page

When a suite repairs its own broken locators and still passes, which product defects does that green run hide?

level: juniorimportance: must knowfreq 55%
basics
~20 s

A repair hides every change that broke the original anchor: a control that moved, was relabelled, lost its announced name, or was replaced by a different one. The suite reports the flow works while the interface changed.

open as a page

What should a service-interface test assert about a response instead of comparing the whole payload?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Assert the status code, the few fields the case is actually about, and the response shape the contract promises — and on failure paths the error envelope's code. Whole-body equality breaks whenever an unrelated field changes.

open as a page

How does an automated case observe an outbound callback that the system sends to a configured destination?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Point the system's configured destination at a receiver the run stands up, then read the recorded request back from it. The receiver must be reachable from where the system runs, not only from the machine running the suite.

open as a page

Which property should an automated test assert on when the effect it checks lands only after the call returns?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Assert on a settled invariant, a property that stays true once the flow has finished, such as a terminal status or a final total. Do not assert on a counter or an intermediate status that is true for only a moment.

open as a page

What is the difference between an object mother and a test-data builder?

level: juniorimportance: must knowfreq 66%
basics
~20 s

An object mother exposes named ready-made instances, so a test asks for a case by name. A test-data builder starts from valid defaults and lets a test override only the fields that case depends on.

open as a page

Why do test suites run against an in-memory database stand-in, and what does that hide?

level: juniorimportance: must knowfreq 68%
basics
~20 s

An in-memory database stand-in starts in milliseconds, needs no external service and resets cleanly between cases, so a suite runs fast anywhere. It hides every behaviour where the production engine differs: dialect, type coercion, collation, constraint enforcement and locking.

open as a page

How do hand-built fixtures, generated records and production extracts differ as test-data sources?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Hand-built fixtures state only the few records a case needs, so they read clearly and carry no privacy risk. Generated records buy volume and variety cheaply. Production-derived extracts show real shapes and skew, but must be cut down and masked first.

open as a page

Why must a test's teardown still run after an assertion fails partway through the test?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A failing assertion aborts the test body, so cleanup written after it never executes and the records it created leak into later tests. Register teardown as an after-each hook or a resource-scoped block so it runs on both exits.

open as a page

Why should a test that generates random data record and print the seed it used?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Random data makes a failure depend on values nobody chose. Recording the seed and printing it in the failure output lets anyone re-run the identical values, so the failure can be reproduced and fixed instead of guessed at.

open as a page

What is fuzzing, and what counts as a failure when there is no expected output to compare against?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Fuzzing drives a program with large volumes of generated, malformed or mutated input. Because no per-input expected result exists, the oracle is the program's own bad behaviour: crashes, hangs, assertion failures, and reports from an instrumented build.

open as a page

What is model-based testing, and where do the executable test cases come from?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Model-based testing builds an explicit model of intended behaviour, usually a state machine, and derives cases from it automatically. A generator walks paths through the model; each walk becomes a test, and the model predicts the expected result.

open as a page

What is a round-trip property in property-based testing, and what bug can it miss?

level: juniorimportance: must knowfreq 63%
basics
~20 s

A round-trip property asserts that transforming a value and then reversing the transformation returns the original value, for every generated input. It misses paired defects: when both directions are wrong in mirror-image ways, the trip still returns the input unchanged.

open as a page

How do dumb mutation, grammar-aware generation and coverage-guided fuzzing differ, and when is each right?

level: middleimportance: must knowfreq 55%
basics
~20 s

Dumb mutation flips bytes in existing inputs and knows nothing about the format. Grammar-aware generation builds well-formed inputs from a description of the format. Coverage-guided fuzzing uses an instrumented build to keep inputs that reach new code and mutate those further.

open as a page

How do you state a property for a function that has no inverse to round-trip against?

level: middleimportance: must knowfreq 56%
basics
~10 s

Assert a rule the output must satisfy rather than a specific value: an invariant or postcondition, idempotence, agreement with a simpler reference implementation, or a metamorphic relation between two runs on related inputs.

open as a page

In what order should you read a code change under review, and why does the order matter?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Read the intent first — what problem the change claims to solve — then the contract it exposes, then the edge cases and error paths, then the tests. That order settles whether the change is right before you argue about how it is written.

open as a page

Why does a defect found in production usually cost more to fix than the same defect found during development?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A production defect costs more because the code change is only a fraction of the work. Someone must notice it, reproduce it and re-learn code written weeks ago, and the team then pays for data repair, support and an unplanned release.

open as a page

What are the driver and navigator roles in pair programming, and why rotate them?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The driver types and handles the mechanics. The navigator stays a step ahead — unhandled cases, naming, the remaining plan — and keeps talking. Rotating the keyboard on a timer keeps both people active and able to explain the change.

open as a page

What are the common quality ownership models, from an independent test group to whole-team quality?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Four shapes recur: an independent test group that verifies finished work, testers embedded in each delivery team, whole-team quality where every engineer tests and no dedicated tester exists, and a small coaching group that builds testing skill in others.

open as a page

How do you measure a test suite's flake rate, and why is its trend more useful than its level?

level: middleimportance: must knowfreq 60%
basics
~20 s

Flake rate is the share of runs or cases giving a different verdict on the same unchanged revision. Measure it by recording every attempt, not just the final one, and track the trend: an acceptable level depends on suite size.

open as a page

A feature's answer text comes from a generative step. In what distinct ways can that step fail to deliver an answer?

level: juniorimportance: must knowfreq 56%
basics
~20 s

Four delivery failures matter: nothing comes back inside the deadline, the call is refused for volume, the provider is unavailable, or the input exceeds what the step accepts. Each strands the user differently, so each earns its own test case.

open as a page

For a feature whose answer text comes from a generative step, what makes an acceptance criterion adjudicable?

level: middleimportance: must knowfreq 64%
basics
~20 s

An adjudicable acceptance criterion names one observable property of the response, says where to look for it, and fixes the decision rule in advance, so two testers reading the same response record the same pass or fail.

open as a page

Before writing test cases for an AI-powered feature, why agree how often each acceptance criterion may miss?

level: middleimportance: must knowfreq 57%
basics
~20 s

A criterion with no stated allowance is read as absolute, so the first response that misses it forces an unplanned argument. Agreeing the allowance up front makes a rare miss an expected, recorded outcome rather than a crisis.

open as a page

When a test of a feature that answers from a generative step fails and will not reproduce, what must that test run itself have recorded?

level: middleimportance: must knowfreq 62%
basics
~20 s

Write the evidence at failure time, because a repeat may never reproduce it: the exact input, the material and instruction text given to the generative step, the model identifier and generation settings, the raw output, and a correlation id.

open as a page

Which layers can own a wrong answer from an AI-powered product feature, and how do you decide which one it is?

level: middleimportance: must knowfreq 62%
basics
~20 s

Five layers can own it: the ordinary product code around the generative step, the instruction text sent to it, the material supplied with it, the configuration, and the generative model's own capability ceiling. Attribute only after evidence separates them.

open as a page