skip to content

When is the upkeep of a behaviour-model test suite no longer worth it for a small team?

level: principalimportance: should knowfreq 31%

answer

  1. Build cost is visible, carry cost is not
  2. The suite stays green about vanished behaviour
  3. Watch edits, triage mix, quarantine, ownership
  4. Defects per maintenance hour, not case count
  5. Freeze, promote the valuable paths, delete the rest

basics

~20 s

Stop when the model has drifted into a stale second specification, when most generated failures are model or adapter noise rather than defects, or when one person alone can edit it. Judge by defects found per maintenance hour.

solid answer

~50 s

Treat a behaviour model as a long-lived asset with a fixed build cost and a permanent carrying cost, and review it like any other investment. It pays where behaviour is stateful, long-lived and order-sensitive, and where each product change costs a small model edit rather than a rewrite. Watch four signals of decay: model edits growing faster than product changes, a rising share of failures triaged as model or adapter faults rather than real defects, generated paths quarantined instead of fixed, and a bus factor of one on the model itself. Measure the thing that matters — genuine defects found per maintenance hour — and compare it against the hand-written suite it displaced. Retirement is not failure: freeze the model, promote the highest-value generated paths into ordinary maintained cases, and delete the generator so nobody inherits an artefact they cannot read.

go deeper

for a junior

Understand that a model is maintained forever, not written once, and that behaviour changing in the product means the model must change in the same piece of work. Notice when a generated failure reflects a deliberate change rather than a defect.

for a middle

Be able to describe drift concretely and its cost: stale predictions producing failures that consume triage and return nothing. Know that the fix for a failing generated path is a model edit, never a skip, and be able to make that edit yourself.

for a senior

Show that you track the signals — edit ratio, triage mix, quarantine, run time — and act on them before trust is lost. Be ready to describe coarsening an abstraction that had grown into a second implementation and what it bought.

for a principal

Own the investment case and the exit. Say which systems in an estate justify a model, who carries it and how authorship is spread, what evidence would make you stop, and how you retire it so the valuable generated paths survive and the pipeline does not linger half-maintained.

## The investment framing A behaviour model is not a test; it is an asset with three cost lines. **Build**: the abstraction decision, the model, the adapter, the harness. **Carry**: every product change that touches modelled behaviour also costs a model edit, a review and often an adapter change. **Triage**: every failure needs sorting into system fault, model fault or adapter fault before anyone can act. The build cost is visible and gets estimated; the carry and triage costs are invisible at the decision point and are what actually kill the effort. So the question is never "is model-based testing good?" but "does this system, with this team, at this rate of change, return more than the carry?" ## What drift actually is Drift is the model and the system diverging while the suite continues to run. It is insidious because a drifted model does not go silent — it keeps producing green runs about behaviour that no longer exists, and red runs about behaviour that changed deliberately. Each such red costs triage and returns nothing, and the team learns to distrust the signal. The usual mechanism is asymmetric pressure. Shipping is urgent, the model is not, so a rule changes in the product on Tuesday and the model gets it on some later Tuesday, or never. Once the gap is wide enough that a model edit is a small project rather than a small change, the technique has already failed; what remains is deciding when to say so. ## The four signals **Edit ratio.** Track model edits against product changes touching modelled behaviour. A healthy model absorbs a change in one small, reviewable edit. When a routine change needs restructuring the state machine, the abstraction was wrong — usually too fine, so the model has become a second implementation with its own defects. **Triage mix.** Tag every failure by culprit. Early on, roughly one genuine defect for every two or three model or adapter faults is normal and should improve. When the ratio moves the other way over a quarter, the suite is generating work rather than findings. **Quarantine.** Generated paths being skipped rather than reconciled is the clearest terminal symptom. A quarantined generated path cannot be fixed in isolation, because it was derived — skipping it means the model is wrong and nobody has time to say how. **Ownership.** A 4-person team is the interesting stress case. If the model's original author owns every edit, three of the four are blocked whenever behaviour changes, and the model outlives its author's tenure by weeks. Models die of a bus factor of one more often than of any technical cause. ## Measuring rather than arguing The metric worth defending is **genuine defects found per maintenance hour**, compared against what the displaced hand-written cases were returning, with two adjustments. Count the defects the technique is uniquely able to find — order-dependent and long-sequence faults — rather than the ones any smoke case would have caught. And count the model as documentation value if, and only if, someone outside the team actually reads it; most teams claim this benefit and cannot name a reader. Run time belongs in the same conversation. When the billing model grew from 9 states to 34 states and 88 transitions, the nightly generated run reached 71 minutes, which changed who was willing to wait for it and therefore how quickly drift got noticed. A suite people do not wait for stops being a feedback loop and becomes an archive. ## Keeping it alive when it is worth keeping Coarsen deliberately: fewer, more meaningful states, with detail pushed into state invariants rather than new states. Make the model part of the definition of done for any change to modelled behaviour, reviewed in the same change rather than followed up. Spread authorship — pair on the first few edits until every engineer has made one. And bound the generated set with an explicit time budget so growth is a visible decision rather than a slow slide. ## Retiring it well Retirement is a legitimate outcome, not an admission of failure, and doing it badly is worse than not starting. Freeze the model rather than deleting it silently. Take the generated paths that have historically caught real defects — usually a small fraction — and promote them to ordinary hand-maintained cases with the intent written into their names, so the accumulated knowledge survives. Then delete the generator and the adapter, because a half-maintained generation pipeline is exactly the kind of artefact a future engineer will neither trust nor dare remove. ## What an interviewer is listening for That you treat this as a portfolio decision rather than a technology preference: which systems in an estate justify a model at all, who carries it, what evidence would make you stop, and what the exit looks like. Confident advocacy with no exit criterion is the answer that fails; so is dismissal of the technique on the strength of one team's bad adapter.

  • What single change most often rescues a model that has become expensive to maintain?
    Coarsening the abstraction. Expensive models are usually too fine — they have grown a state for every distinguishable condition and become a second implementation, so every product change ripples through them. Collapsing states that behave identically under every action, and expressing the lost detail as invariants asserted on arrival, cuts both the edit cost and the generated path count without weakening the oracle much.
  • How would you decide whether a new system in your estate deserves a behaviour model at all?
    Ask four things. Is the behaviour genuinely stateful, with faults that live in orderings rather than in single calls? Will the system live long enough to repay the build cost? Is the intended behaviour written down well enough to model from, and is someone available to review the model against it? And is there more than one engineer who will maintain it? Two or more nos, and hand-written cases are the better buy.
  • What do you keep when you retire a model-based suite?
    The paths that earned their place. Look at which generated sequences historically caught real defects, promote those into ordinary maintained cases with names that record the intent, and check them in. Then remove the generator, the adapter and the model from the build, keeping the model in history as documentation. Leaving a partly maintained pipeline behind is worse than removing it, because nobody later will trust it or dare delete it.

saying these in an interview costs you the question

  • Treats the model as free once it has been written
  • Keeps a model only its original author can edit
  • Measures success by the number of generated cases
  • Quarantines generated paths instead of reconciling the model
  • Abandons the technique after one bad adapter experience
  • Assumes a larger, finer model is always a better model

context