skip to content

Scenario Suite Maintenance

The main BDD frameworks across ecosystems, their trade-offs, and how runners, IDE plugins and CI reporting fit together. Asked to see whether you have run these suites for real, including what they cost to maintain.

on this pageshow

questions

4

What is step-definition sprawl in a scenario suite, and why does it raise maintenance cost?

level: juniorimportance: must knowfreq 62%

answer

  1. Two artefacts, one of them costs money
  2. Growth per sentence, not per behaviour
  3. Same fact, several phrasings, several definitions
  4. One rule change, many near-copies to edit
  5. Track definitions-per-scenario as a trend

basics

~20 s

Step-definition sprawl is a glue layer that grows one definition per sentence written instead of one per behaviour, so near-identical phrasings each get their own code. The library becomes large and duplicated, and every behaviour change touches many places.

solid answer

~40 s

Every distinct sentence in a scenario needs a registered piece of code that matches it. Sprawl is when that glue library grows with the number of sentences typed rather than the number of behaviours the product has, so "the member has 240 points", "the balance is 240 points" and "the account holds 240 points" each get their own definition. The cost is change amplification: one rule change means editing many near-copies, and a missed copy leaves a scenario passing against the old rule. Sprawl also brings ambiguous matches, dead definitions no scenario references, and a slow search that pushes the next author to write yet another copy. Watch the definitions-per-scenario ratio as a trend, keep a readable index of the sentences the suite already speaks, and consolidate duplicates as you touch them.

code

pseudocode · 8 lines
pseudocode
step("the member has 240 points"):
    ledger.setBalance(member, 240)

step("the account holds 240 points"):
    ledger.setBalance(member, 240)

step("the balance is 240 points"):
    ledger.setBalance(member, 240)

go deeper

for a junior

Be ready to say what the glue layer is and why writing a new sentence for a fact the suite already expresses adds code. Knowing that duplication, not scenario count, is what makes these suites expensive is enough at this level.

for a middle

Explain the mechanics: text matching means each phrasing needs its own registration, so one behaviour change fans out across near-copies, and overlapping patterns produce ambiguous matches. Be able to name the signals you would measure.

for a senior

Show that you have owned one of these libraries. Talk about the dead-glue report, near-duplicate clustering, consolidating on touch, and the missed-copy failure where a stale definition keeps a scenario green against a rule that no longer exists.

for a principal

Own the accountability question: who reviews the glue layer, what threshold triggers consolidation work, and how you argue for that budget when the sprawl is invisible to everyone outside the team.

## The two halves of a scenario suite A plain-language scenario suite is made of two artefacts that must stay in step. One is the set of scenarios themselves: sentences a business reader can follow, describing context, an event and an expected outcome. The other is the **glue layer** — for every distinct sentence the runner encounters there must be a registered piece of code that matches that sentence and carries it out. The sentences are the readable half; the glue is the half that costs money to own. **Step-definition sprawl** is the failure mode in which the glue layer grows in proportion to the number of *sentences that have been typed* rather than the number of *behaviours the product has*. A suite with three hundred scenarios covering forty business rules should not need eight hundred definitions; when it does, most of those definitions are near-copies of each other. ## Why it happens, even to careful teams Nobody decides to sprawl. Four ordinary forces produce it: - **The path of least resistance.** An author writes a new sentence, the run fails with "no matching definition", and the cheapest repair is to copy the nearest existing definition and edit the copy. - **Search cost.** Once the library is large, checking whether a sentence for "the member's balance is 240 points" already exists takes longer than writing a fresh one. The economics quietly favour duplication. - **Phrasing entropy.** "the member has 240 points", "the account holds 240 points" and "the balance is 240 points" are one fact in three costumes. Natural language invites variation; the runner matches text, so each costume needs its own code. - **No audience.** The scenario files have business readers and get read in review. The glue has no audience outside the team, so it is skimmed rather than reviewed, and duplication is never called. ## What it actually costs **Change amplification.** A rule changes in one place in the product and in many places in the glue. Every missed copy is a scenario that keeps asserting the old rule — passing, and therefore silent. **Onboarding and authoring cost.** New joiners cannot tell which sentence to reuse, so they add more. The ratio worsens monotonically unless somebody is accountable for it. **Ambiguity.** Two patterns that both match one sentence produce either a hard failure or, in the worse case, a match against the wrong definition, so the scenario tests something nobody wrote. **Dead glue.** Definitions no live scenario references still compile, still get refactored, still appear in searches, and still mislead the next reader about what the suite covers. **Erosion of the readability payoff.** The suite justified its translation layer on the claim that the files read as one language. Twelve dialects of the same sentence is exactly the claim failing. ## A concrete shape An eleven-person team owns a loyalty-points ledger. Their suite carries 340 scenarios and 812 step definitions — roughly 2.4 definitions per scenario, and rising quarter over quarter. A rule change ("points are awarded once per settled order, not once per settlement attempt") touched nine definitions that all expressed the same award step. Eight were updated. The ninth, reachable only from two older scenarios, kept the old behaviour, and the mismatch let a retried settlement produce a **duplicated side effect**: points credited twice on the same order, in scenarios that stayed green because they asserted the stale rule. The defect was not in the business logic; it was in a copy of the glue that had drifted out of sight. ## Seeing it before it hurts Useful signals, all cheap to compute: - **Definitions per scenario, tracked as a trend.** The absolute number means little across teams; a ratio climbing over two quarters means the library is outgrowing the behaviour set. - **Definitions never matched during a full run** — the dead-glue count. - **Near-duplicate clusters.** Normalise sentence text (lower-case, replace numbers and quoted values with placeholders) and cluster; the clusters are your consolidation backlog. - **Ambiguity failures per month**, which rise as overlapping patterns accumulate. - **The healthiest signal of all:** how often a new scenario can be written entirely from sentences that already exist. A team that adds scenarios without adding glue has a vocabulary, not a vocabulary problem. ## What you do about it At suite level the levers are ownership and vocabulary rather than cleverness. Give the glue layer a named owner and review it as production code. Agree a phrase vocabulary — a short, readable index of the sentences the suite already speaks — and make consulting it part of writing a scenario. Consolidate on touch instead of scheduling a big cleanup that never gets funded. Delete dead definitions the moment the report names them, and wire the dead-glue and ambiguity checks into the build so the count cannot silently climb again. How an individual binding should be shaped is its own subject; what the suite owner controls is whether the library grows per sentence or per behaviour, and whether anyone is watching the trend.

  • How would you find step definitions that no scenario references any more?
    Have the runner report which definitions were matched during a full run and subtract that from the registered set; the remainder is dead glue. Do it on the full pack rather than a filtered subset, or you will delete definitions that only a nightly scenario uses. Fail the build when the dead count rises, then delete rather than archive — an unreferenced definition still shows up in searches and still misleads the next reader about what the suite covers.
  • Is a rising definitions-per-scenario ratio always a problem?
    No. A suite covering genuinely new behaviour will add definitions legitimately, and a small suite has a noisy ratio. The signal is the trend against the behaviour set: if the number of business rules is flat while the library grows for two quarters, the growth is duplication. Pair the ratio with near-duplicate clustering and the dead-glue count before drawing a conclusion from it.
  • What happens when two definitions both match one sentence?
    Most runners treat an ambiguous match as an error and fail the scenario, which is the good case because it is loud. The bad case is a runner or configuration that resolves to the first match, so the scenario silently exercises a different piece of code than the author intended. Either way the cure is the same: consolidate the overlapping patterns rather than narrowing one of them until the clash goes away.

It is a phrasebook that gains a new page every time someone says the same thing slightly differently: eventually nobody can find the page they need, so they write another one.

saying these in an interview costs you the question

  • Thinks more step definitions means more test coverage
  • Copies an existing definition rather than searching for one
  • Treats the glue layer as throwaway code, not production code
  • Blames slow runs on the size of the glue library
  • Leaves unreferenced definitions in place "just in case"
  • Fixes an ambiguous match by narrowing a pattern instead of merging

context

open as a page

How do tags on scenarios keep feedback fast on a large scenario suite?

level: middleimportance: should knowfreq 48%

basics

~20 s

Tags are labels the runner selects on, so each pipeline stage runs only what it needs: a fast subset per change, the full pack nightly. The value comes from a small, enforced taxonomy, not the mechanism.

open as a page

Your scenario suite has drifted into a 47-minute interface-driven pack. How do you get feedback back?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Diagnose first: per-scenario durations, which layer each scenario drives, and how many scenarios restate the same rule. Then re-point most scenarios below the interface, delete duplicates, and keep only a thin set of interface-driven journeys.

open as a page

How do you judge whether a scenario suite's plain-language layer still pays for its maintenance cost?

level: principalimportance: should knowfreq 38%

basics

~10 s

Weigh the translation layer's cost — slower authoring, indirection when debugging, an owned glue library — against evidence it buys participation and shared vocabulary. If only engineers read the scenarios, the cost buys nothing.

open as a page