skip to content

How do you contain state explosion in a behaviour model that is too large to test?

level: seniorimportance: nice to knowfreq 22%

answer

  1. The count multiplies, it does not add
  2. Ask whether the arrows differ or just the data
  3. Shared arrows drawn once above the group
  4. Independent concerns modelled side by side
  5. Bound the coverage where you cannot shrink the model

basics

~10 s

Stop encoding data in states. Move data conditions into guards, fold related states under a superstate, split independent concerns into separate parallel models, and cover only the risky regions at a higher switch level.

solid answer

~50 s

State explosion happens when independent conditions get multiplied into the state space: three lifecycle stages times two payment flags times two channels is twelve states instead of three. Four containment moves, roughly in order of payoff. **Guards** — push a data condition off the state space onto the arrows, so one state fans out on data rather than duplicating. **Superstates** — group states that share outgoing arrows and draw the common arrow once at the group level, which collapses a fan of repeated transitions. **Orthogonal models** — when two concerns vary independently, model them as two small machines rather than one machine over their cross-product; a nine-state and a four-state model beat one thirty-six-state model. **Scoped coverage** — keep 0-switch across the whole model as the floor and buy 1-switch only on the region where history actually matters. Then delete unreachable states, which are pure maintenance cost.

code

pseudocode · 13 lines
pseudocode
# BEFORE: data folded into the state space -> 12 states
DraftWeb, DraftStore, DraftTransfer,
ReservedWeb, ReservedStore, ReservedTransfer, ...

# AFTER: 4 states, channel carried as data on a guard
STATE Reserved
  ON pick WHEN channel = TRANSFER -> Picked / skipLabel()
  ON pick WHEN channel <> TRANSFER -> Picked / printLabel()

# AFTER: shared arrows lifted to a superstate
SUPERSTATE Open { Draft, Reserved, Picked }
  ON cancel -> Cancelled / releaseHold()
  ON expire -> Expired   / releaseHold()

go deeper

for a junior

Recognise the symptom: state names that read like combinations of conditions, and a diagram nobody can read. Know that a data condition usually belongs on an arrow rather than in a new state.

for a middle

Be able to apply the mechanical moves — collapse states that share all arrows into one plus a guard, lift shared arrows to a superstate — and explain why the state count multiplies while guards only add.

for a senior

Show judgement on splitting independent concerns into parallel models and on scoping pair coverage to the regions where history leaks. Be ready to say what you would delete and what you would refuse to delete.

for a principal

Own what the model is for and who maintains it. A model that drives generation is a coverage artefact and shrinking it is a coverage decision; a model built for conversation can be coarse. Assign ownership so it does not drift from the code.

### Where the explosion comes from A state model grows multiplicatively, not additively. Each independent condition someone folds into the state name multiplies the state count by its number of values. A warehouse stock ledger that starts as four clean lifecycle states — Draft, Reserved, Picked, Shipped — turns into sixteen the moment someone adds "partially picked" and "back-ordered" as states rather than as data, and into forty-eight when the channel (store, web, transfer) joins them. The arrows grow faster still, and the test count grows with the arrows. Once a model needs a wall to print on, nobody reviews it, and an unreviewed model is worse than no model because it carries authority it has not earned. ### Move data out of the state space The first question for any suspicious state is: does it have different arrows, or the same arrows with different data? If two states share every outgoing event and every target, they are one state and a data condition. Express the condition as a **guard** on the arrows and delete one of the states. This alone often halves a bloated model. The cost is honest and small: each guard splits one arrow into a satisfied and an unsatisfied case, which is additive, whereas the state it replaces was multiplicative. A related move is to take a dimension out of the state model entirely and cover it with a different design technique. The channel a reservation came from usually does not change *which arrows exist*; it changes values that flow through them. That belongs in the data dimension of the case, chosen by whatever technique the team uses for input values, not in the lifecycle model. ### Superstates and hierarchy When several states share the same outgoing arrow — every non-terminal ledger state accepts cancel and expire — draw the group once. A **superstate** (the hierarchical statechart idea) holds the common transitions, and the substates hold only what differs. Ten arrows collapse to four plus two group-level arrows, and the reader can see the lifecycle without the noise. For test derivation, a group-level arrow still expands: it must be exercised from each substate, or at least from a risk-chosen subset, and being explicit about which you chose is part of the design record. ### Orthogonal regions The most common cause of explosion is two concerns that do not interact being modelled over their cross-product. A ledger row's fulfilment lifecycle and its financial settlement lifecycle mostly proceed independently; modelling them together produces a machine over every combination. Modelling them as two **orthogonal** machines gives two small, reviewable models and a much smaller case count — plus an explicit, short list of the points where they genuinely do interact, which is where the interesting defects are. The judgement call is exactly that list: split too eagerly and you lose the interactions the single model would have shown. ### Scope the coverage instead of the model Sometimes the model is genuinely irreducible, and the lever is coverage rather than structure. Keep 0-switch as the floor across the whole model — every specified arrow taken once — and buy pair coverage only where history plausibly leaks: regions that allocate or release resources, that carry an amount forward, or that compensate a previous step. Publishing that as an explicit rule ("pairs on the settlement region, singles elsewhere") turns an unbounded suite into a bounded one, and makes the gap deliberate and reviewable rather than accidental. ### Prune what cannot happen Read the model for states that appear as no transition's target: they are unreachable, and unless an arrow is missing they should be deleted from the model and the code. Do the same for arrows whose guard cannot be satisfied. Dead structure is not free — it inflates every count, invites cases that need a back door to set up, and misleads the next reader. ### The organisational half Containment is also about who maintains the thing. On a team of eleven, a single hundred-state model owned by nobody drifts out of step with the code within a release, and a drifted model produces confidently wrong tests. Several small models, each owned by the group that owns that subsystem, each small enough to review in a stand-up, survive. Ask what the model is *for* before shrinking it: a model built to generate a regression suite must be complete over the behaviours the suite claims, while a model built to have the conversation with the requirement owner can be far coarser and still do its job. Shrinking the first kind carelessly deletes coverage; refusing to shrink the second kind wastes weeks.

  • When is splitting one model into two orthogonal models the wrong move?
    When the two concerns interact more than occasionally. Splitting hides exactly the cross-concern combinations that the single model would have forced you to look at, and those interactions are usually where defects sit. Split when the interaction points are a short, enumerable list you can test explicitly; keep one model when the two lifecycles constrain each other at most steps, and shrink it by other means instead.
  • How do you decide which region of a large model deserves pair coverage?
    Follow the resources and the carried-forward data. Regions that allocate or release something, compensate an earlier step, or accumulate an amount are where entering a state by two routes can differ. Pure navigation regions rarely repay it. Write the choice down as an explicit rule, so the gap is a decision on record rather than an accident nobody can defend in a review.
  • What is the risk of shrinking a model that a generator uses to produce the regression suite?
    You delete coverage silently. A model used for conversation can be coarse without harm, but a model that generates the suite defines what the suite claims to cover, so every state folded away removes cases nobody will notice are gone. Before simplifying, establish what the model is for; if it drives generation, treat a reduction as a coverage change and review it as one.

saying these in an interview costs you the question

  • Adding a state for every data variation
  • Never questioning whether two concerns are independent
  • Keeping unreachable states because deleting feels risky
  • Assuming a bigger model is a better model
  • Splitting into parallel models that interact constantly
  • Shrinking a generator's model without treating it as a coverage change

context