How do you judge whether a scenario suite's plain-language layer still pays for its maintenance cost?
answer
- Cost paid whether or not read
- Ask who edits the scenarios, and when
- General evidence for the practice is contested
- Split the suite rather than abolish it
- Keep the conversation even if the format goes
basics
~10 sWeigh the translation layer's cost — slower authoring, indirection when debugging, an owned glue library — against evidence it buys participation and shared vocabulary. If only engineers read the scenarios, the cost buys nothing.
solid answer
~40 sThe suite trades an extra artefact for readability, and the cost is incurred whether or not the readability is used. I look for evidence of use: has anyone outside engineering proposed or amended a scenario recently, are examples produced before implementation, are intent disputes settled by opening a scenario? Against that I set the cost: authoring speed, the indirection when triaging a failure, and the glue library's growth trend. The general evidence for the practice is contested, so the decision has to come from local observation rather than a claimed industry number. The answer is rarely all-or-nothing: keep the plain-language layer where a real reader and reviewer exist, convert the rest to ordinary coded tests, and keep the discovery workshops either way, since the conversation usually carries most of the value.
go deeper
Understand that the readable sentences are an extra artefact with a real upkeep cost, and that the cost is only worth paying if someone outside the team reads or writes them.
Be able to name both sides of the trade concretely: slower authoring and indirection when debugging on one side, shared vocabulary and non-engineer participation on the other.
Bring observations rather than opinions — who last edited a scenario, how failures are triaged, how the glue ratio is trending — and propose keeping the layer only where a reader demonstrably exists.
Own the decision and its reversal cost: a named kept subset, a migration that converts on touch, a review date, and a written rationale. Say plainly that the general evidence is contested and that you decided from local data.
## The trade the format makes A plain-language scenario suite buys readability with an extra artefact. Instead of a test that says what it does in code, you have sentences plus a translation layer that binds them to code. The translation layer is not free, and a principal owning this decision has to be able to name both sides of the trade without sentiment. **What it costs.** Authoring is slower, because writing a scenario means finding or adding sentences as well as writing the code beneath them. Debugging pays an indirection tax: a failure is read in one artefact and diagnosed in another. The glue library needs an owner and consolidation work, or it grows per sentence rather than per behaviour. Reporting and the surrounding tooling need upkeep. Review happens twice, once for the readable half and once for the code. And every new joiner pays a learning cost for a format they may not have used before. **What it buys.** A shared vocabulary between people who otherwise talk past each other. Genuine participation from people who do not read code — proposing examples, spotting a missing case, correcting a rule. A specification that can be read outside the team. Concrete examples produced before implementation, which is where most of the defect prevention actually happens. Regression protection expressed in terms of the domain rather than the implementation. The payoff is entirely contingent on the second list actually occurring. The cost is incurred whether it occurs or not, which is why this decision has to be made on evidence rather than on the fact that the suite exists. ## Signals that the payoff has stopped - Scenarios are written **after** the code, by engineers, as a formality — the discovery value is gone and only the tax remains. - No one outside the engineering team has proposed or amended a scenario in a quarter. - Scenario text has filled with interface nouns and technical detail, so the readability claim is nominal; a business reader cannot follow it even if willing. - Failures are triaged by reading the glue rather than the sentence, meaning the readable half is no longer the primary artefact. - The definitions-per-scenario ratio keeps climbing and consolidation work never gets funded. - The suite is described inside the team as "the regression pack" rather than as the specification. ## Signals that it still pays Someone outside engineering proposes examples before implementation starts. Disputes about intended behaviour are settled by opening a scenario. The vocabulary shows up in conversation, in tickets, in the product's own language. New rules arrive already expressed as examples. Where these hold, the translation layer is not overhead — it is the thing producing the alignment, and cutting it would cost more than it saves. ## Be honest about the evidence There is no dependable published number for what this format returns. The claims made for it are largely experiential, the studies are small and confounded, and the honest position in an interview is that the general evidence is contested and the decision has to be made from your own team's observed behaviour. Someone who quotes a defect-reduction percentage for the practice is quoting folklore. What you can measure locally is real: who edits the scenarios, how long a change takes, how often intent disputes reach implementation. ## The decision is not binary The eleven-person team on the loyalty-points ledger looked at six months of history and found three scenario edits by anyone outside engineering, all three in the earn-and-burn rules that the product owner reviews before each release. Their answer was a split, not an abolition: keep the plain-language layer for the 41 scenarios covering earn, burn, expiry and reversal, where there is a real reader and a real reviewer; convert the remaining scenarios — the ones no one outside the team had ever opened — into ordinary coded tests at the level they belong; and keep the discovery workshops unchanged, because the conversation that produces concrete examples was carrying most of the value and does not require the file format at all. ## Running the migration without a cliff Stop first, convert second. Freeze new scenarios outside the kept subset so the population stops growing while you decide. Convert on touch rather than in a big bang, so the effort lands where change is actually happening. Delete duplicates before converting anything — it is the cheapest work and it reduces what has to be migrated. Keep the published report alive until you can see the audience has moved, because switching it off is how you discover the reader you did not know about, loudly and at the wrong moment. Set a review date and say what you expect to see by then. The organisational point matters as much as the technical one: this decision is usually made silently, by attrition, when the last person who cared stops maintaining the glue. Making it explicitly — with the evidence written down, a named kept subset, and a review date — is what distinguishes a lead's call from a suite that simply rotted.
- If you drop the format, what would you keep regardless?The discovery conversation that produces concrete examples before implementation. That is where the misunderstandings get caught, and it needs no file format at all — the examples can live in the coded tests' names and data. Teams that abolish the workshops along with the syntax usually lose the part that was working and keep none of the savings they expected.
- What would you measure over the next quarter to make this call on evidence rather than opinion?Edits to scenario text by people outside engineering; how many new rules arrived as examples before implementation started; the definitions-per-scenario trend; median time to diagnose a failed scenario; and whether intent disputes were settled by reading a scenario. All five are observable from version history and triage notes, and together they tell you whether the readable half is a working artefact or a ceremony.
- How do you run the conversion without breaking a reader you did not know about?Freeze new scenarios outside the subset you intend to keep, convert on touch instead of in one sweep, and delete duplicates before migrating anything. Keep the published output alive after the conversion begins and watch whether anyone asks about it. Announce a review date with what you expect to see by then, so the decision stays reversible rather than becoming permanent by accident.
It is a translated edition of a document: worth the translator's salary only while somebody is actually reading the translation.
saying these in an interview costs you the question
- Defends the format on principle without evidence of readers
- Quotes a defect-reduction percentage for the practice as fact
- Treats the choice as abolish-everything or keep-everything
- Drops the discovery workshops along with the file format
- Lets the decision happen by attrition when maintenance stops
- Counts scenario volume as proof the layer is valuable