What is model-based testing, and where do the executable test cases come from?
answer
- The cases are not written by hand
- One artefact holds the intended behaviour
- States, actions, guards
- A generator walks it and emits paths
- It predicts the expected result too
basics
~20 sModel-based testing builds an explicit model of intended behaviour, usually a state machine, and derives cases from it automatically. A generator walks paths through the model; each walk becomes a test, and the model predicts the expected result.
solid answer
~50 sIn model-based testing you write the model instead of the cases. The model is an explicit, machine-readable description of intended behaviour — most often a state machine: named states, the actions that move between them, and guards saying when an action is legal. A generator traverses that model and emits paths, and an adapter layer translates each abstract action into a real call against the system and each abstract state into something observable. Because the model says where every action should lead, it also supplies the expected result, so it doubles as the oracle. The pay-off is that new cases come from editing one artefact rather than writing dozens of scripts, and the generator explores orderings nobody would have thought to script. The costs are the abstraction decision, the adapter code, and keeping the model honest as the product changes.
code
pseudocode · 16 linesmodel BillingRun:
states = [Draft, ReadingsLoaded, Estimated, Validated, Invoiced, Disputed, Settled]
initial = Draft
action loadReadings: from Draft -> ReadingsLoaded
action estimate: from ReadingsLoaded -> Estimated when missingReadings > 0
action validate: from Estimated -> Validated
action issueInvoice: from Validated -> Invoiced
action raiseDispute: from Invoiced -> Disputed
action settle: from Invoiced -> Settled
for path in generate(model = BillingRun, criterion = ALL_TRANSITIONS):
system = adapter.startFreshBillingRun()
for step in path:
adapter.perform(step.action, system)
assert adapter.observe(system) == step.expectedStatego deeper
Be ready to define the technique in two sentences and name the three pieces: a model of intended behaviour, a generator that walks it, and an adapter that connects abstract actions to the real system. Say plainly that the model also supplies the expected result.
Expect to explain the mechanics: what a state, an action and a guard are, how a walk becomes a runnable case, and why the adapter usually dwarfs the model in size. Be able to sketch a small state machine for a stateful workflow on request.
Show judgement about where the technique earns its keep. Talk about stateful workflows whose faults live in orderings, about the abstraction level you would pick, and about triaging a mismatch across three possible culprits: the system, the model and the adapter.
Own the framing that a model is a shared, long-lived asset with a carrying cost. Be ready to say which systems in an estate justify one, who maintains it, and how you would keep it from becoming a second specification that nobody trusts.
## The shape of the technique A hand-written suite stores its knowledge of intended behaviour inside the cases themselves: every case is a small private restatement of what the system should do, and that knowledge is duplicated across hundreds of files. Model-based testing pulls the knowledge out into one artefact — the **model** — and lets a program produce the cases from it. Three parts make up a working setup. **1. The model.** An explicit, machine-readable description of *intended* behaviour. The most common form is a finite state machine: a set of named states, a set of actions (events) that move the system between them, and **guards** — boolean conditions that say when an action is legal. Other forms exist: decision rules over inputs, pre- and postcondition contracts on operations, or a grammar over legal action sequences. The essential property is not the notation but that a program can traverse it: at any point it must be able to ask which actions are legal here and where each one leads. **2. The generator.** A program that walks the model and emits sequences of abstract actions. It is bounded by a selection criterion — visit every state, exercise every transition, stay under a path length — and usually by a budget, because most non-trivial models contain loops and therefore infinitely many walks. **3. The adapter.** Abstract action names mean nothing to the running system. The adapter maps each abstract action onto real work — a call, a form submission, a queued message — and maps observable output back onto model concepts so predictions can be compared with reality. It is routinely the largest body of code in the setup and the part newcomers underestimate. ## A worked example Take a utility billing run owned by a 4-person team. The states are Draft, ReadingsLoaded, Estimated, Validated, Invoiced, Disputed, Adjusted, Settled and Cancelled — nine of them. The actions are load readings, estimate missing readings, validate, issue invoice, raise dispute, apply adjustment, settle and cancel, with guards: issue-invoice is legal only from Validated, apply-adjustment only from Disputed, and cancel is legal from every state before Settled. That is 17 legal transitions. Nobody writes 17 scripts. The team writes the nine states, the eight actions and the guards — around sixty lines — plus an adapter, and the generator produces paths such as Draft → load → ReadingsLoaded → estimate → Estimated → validate → Validated → issue → Invoiced → dispute → Disputed → adjust → Adjusted → settle → Settled. Asking for every transition plus every guard-false rejection yields a set of 214 short paths. The real pay-off arrives at the next change. When the product adds a recall-invoice action, the team edits the model once — one action, two transitions, one guard — and re-derives the entire set, including orderings no one would have scripted by hand, such as dispute then recall then adjust. ## Offline generation versus driving the system live **Offline**: generate the paths first, store them, and run them like ordinary cases. They can be reviewed, diffed and re-run; a failure is reproducible because the path is a file. **On-the-fly**: the generator drives the system live, choosing the next action from the model and consuming the observed result before choosing the one after it. This copes with legitimate non-determinism, where an action may legally lead to one of several states, and it supports very long random walks. The price is reproducibility: unless the seed and the executed action trace are recorded, a failure a hundred thousand steps deep is hard to get back. ## What it buys and what it costs It buys a single place to change, sequences a human would not think to write, a machine-checkable statement of intent that doubles as documentation, and scale — the number of cases grows with the size of the model rather than with author hours. It costs an abstraction decision (model too fine and it becomes a second implementation with its own bugs; too coarse and it predicts nothing worth checking), the adapter code, triage effort when a mismatch could be a fault in the system, in the model or in the adapter, and continuous upkeep as behaviour changes. ## What it is not It is not exhaustive verification of a design: proving that every reachable configuration of a specification satisfies a property is a separate discipline with separate tooling. It is not a route to shipping code — the model here exists to produce cases and predictions. And it does not retire hand-written tests: one awkward boundary case is far cheaper as a single hand-written case than as an amendment to a shared model. Interviewers push hardest on the last point of the model itself: whatever the model does not describe is precisely what the generated suite cannot see.
- What does the adapter layer do, and why is it usually the bulk of the code?The adapter binds abstract model vocabulary to the real system: it turns an action name into an actual call or interaction, and turns observable output back into the model's notion of state so a prediction can be compared. It is large because real systems need setup, teardown, waiting for asynchronous work to settle, and a projection from rich internal state onto the handful of concepts the model names. Most early failures in a model-based effort are adapter faults, not system faults.
- What is the difference between generating a suite offline and driving the system on the fly?Offline generation produces the paths first and stores them, so they can be reviewed, diffed and re-run exactly; a failure is reproducible because the path is a file. Driving on the fly lets the generator pick each next action after seeing the observed result, which copes with legitimate non-determinism and supports very long walks, but a deep failure is only reproducible if the seed and the executed action trace were recorded.
- Does this technique replace hand-written test cases?No. It pays off where behaviour is stateful and the interesting faults live in orderings, because there the case count grows faster than anyone can script. A single awkward boundary value, a regression pinned to one reported defect, or a case that documents one specific business rule is cheaper and clearer written by hand than expressed as an amendment to a shared model. Most teams run both.
It is the difference between writing out every legal itinerary through a rail network by hand and drawing the network once, then letting a program list the journeys.
saying these in an interview costs you the question
- Calls the model a diagram rather than executable input to a generator
- Thinks the model is discovered from the system rather than stating intent
- Claims generated paths amount to exhaustive testing
- Forgets the model must also predict the expected result
- Assumes no adapter code is needed to reach the real system
- Believes writing the model is the whole cost of the technique