skip to content

A feature file's scenarios keep failing on shared setup they don't use — how do you restructure it?

level: seniorimportance: should knowfreq 44%

answer

  1. Coupling, not flakiness
  2. Least reliable shared step sets the ceiling
  3. Would every scenario be wrong without it
  4. Group by rule, or split the file
  5. Fail only for a reason it names

basics

~20 s

Treat it as coupling, not flakiness. Shrink the shared context to what is universally true, push the rest into the scenarios that need it, group by rule or split the file, and move technical setup out of the prose.

solid answer

~50 s

The symptom is structural: shared context at the top of the file re-runs before every scenario, so the file's reliability is the reliability of its least reliable shared step multiplied by the scenario count. First I shrink that context — for each step, would every scenario be *wrong* without it? If not, it moves down into the scenarios that need it, even at the cost of repeated lines. Then I group the scenarios that illustrate one business rule together, using a rule-level block where the dialect offers one so each rule scopes its own context, and splitting the file when its title needs the word "and". Purely technical setup — seeding data, stubbing a dependency — leaves the prose entirely for the automation layer. Where a heavy dependency genuinely is required, I make that visible with a tag instead of implicit at the top. The invariant I am restoring: **a scenario should fail only for a reason it names.**

go deeper

for a junior

Recognise the shape of the problem: shared context at the top of a file applies to every scenario below it, so a failure there is reported against scenarios that have nothing to do with it.

for a middle

Explain the mechanism and the first fix: shared context re-runs per scenario, so move any step that is not universally required down into the scenarios that actually need it, accepting repeated lines in the text.

for a senior

Demonstrate the diagnosis and a sequenced plan — shrink, group by rule or split, move technical setup into the automation layer, make remaining heavy dependencies explicit — plus the run-time and report-readability numbers that justify it, applied one file at a time.

for a principal

Own the standard the team works to: a scenario fails only for a reason it names, and specification text may repeat itself where code may not. Be ready to defend that trade-off against a reviewer who reads repetition as a defect, and to say who pays when a pack's reports stop being trusted.

### Reading the symptom "Scenarios fail on setup they do not use" is a structural diagnosis, not a flaky-test diagnosis. It says the feature file has accumulated shared context — typically a Background block, sometimes an ever-growing chain of Given steps copied into every scenario — that most of the scenarios below it neither need nor mention. The failures are real; they are just landing on the wrong scenarios. The mechanism is worth being precise about, because the fix follows from it. Shared context in a feature file re-runs before **every** scenario in that file, including every expanded row of a Scenario Outline. So the reliability of the whole file is the reliability of its least reliable shared step, multiplied by the number of scenarios. In a **hotel booking channel manager** whose file opens with "Given each distribution channel has acknowledged the current rate plan", a connector that develops **an intermittent timeout** does not fail one scenario about acknowledgement; it fails the rate-validation, currency-rounding and overbooking scenarios too. In a **340-case regression pack**, one such file turned a single flaky dependency into twenty-three red scenarios a morning, none of which named the actual fault. ### The restructuring moves, roughly in order **1. Shrink the shared context to what is universally true.** For each step, ask: would every scenario in this file be *wrong* without it? If the honest answer is "only about half of them", move it down into those scenarios. Duplicating a Given line across four scenarios is cheaper than a precondition that binds fourteen. This alone usually fixes the blast-radius problem. **2. Split by rule.** A feature file that needs five lines of shared context is usually covering more than one capability. Group the scenarios that illustrate a single business rule together — many dialects offer a rule-level block for exactly this, and where it exists each rule may carry its own scoped context, so the acknowledgement steps sit only above the acknowledgement scenarios. Where the dialect has no such block, split the file. The heuristic for both: if the feature's own title needs the word "and", it is two features. **3. Move technical setup out of the prose entirely.** Seeding a datastore, standing up a stub of an external dependency, resetting a clock — none of that is behaviour a business reader wants to read, and none of it belongs in the specification text. It belongs in the automation layer's setup, invisible to the file. What stays in the file is only the domain-level state a reader needs to judge the outcome. **4. Make the dependency honest where it must stay.** If some scenarios genuinely require a live channel connection, group and tag those so the fact is visible in the file and can be selected against, rather than leaving it implicit in a shared block at the top. **5. Push cases down a level.** A file that has grown heavy is often carrying enumeration rather than illustration — a large Examples table of near-identical rows. Keep the rows that illustrate the rule's boundaries; move the rest to a cheaper check that does not pay the scenario-level setup cost at all. ### What an interviewer is listening for The weak answer is a stabilisation answer: add a retry, extend a timeout, quarantine the reds. Those may be worth doing for the underlying dependency, but they do not touch the reason unrelated scenarios failed. The strong answer names the coupling, and states the invariant it is restoring: **a scenario should fail only for a reason it names.** Everything above is in service of that. The second thing being listened for is a cost argument that is not purely aesthetic. Shared context re-runs per scenario, so it is simultaneously the largest reliability multiplier and often the largest time multiplier in the pack; shrinking it makes the pack faster and the reports readable at the same time. Being able to quantify that — "four seconds of shared setup across a 340-case pack is roughly twenty-two minutes of wall clock spent re-establishing state" — is what separates an opinion from a plan. ### The sequencing, and what it costs Restructure incrementally: take the worst file, move its shared steps down, split it by rule, and measure the red count and the run time before and after. Doing it wholesale across a large pack in one change makes the diff unreviewable and, if something regresses, unbisectable. Expect the specification text to get *longer* — that is the trade being made. Repetition that a reader can see is worth more than economy that a reader has to reconstruct by scrolling, and a file whose scenarios each stand alone is one that a domain expert can review a scenario at a time. If the team resists on the grounds of duplication, the counter is that a feature file is a document for humans, and the automation layer beneath it is where duplication should actually be removed.

  • How would you quantify the cost of the shared setup before proposing the change?
    Measure the shared steps' own duration and multiply by the scenario count, since they re-run per scenario and per expanded Examples row. Across a 340-case pack, four seconds of shared setup is roughly twenty-two minutes of wall clock. Pair that with the red count attributable to shared steps rather than to the behaviour under test — that pair turns an aesthetic argument into a plan.
  • Isn't repeating the same Given line in six scenarios exactly the duplication we are told to remove?
    In code, yes; in specification text, no. A feature file is a document read a scenario at a time, and a precondition a reader must reconstruct by scrolling is worse than one stated twice. Deduplicate underneath instead — the automation layer is where a repeated line should map to a single implementation.
  • Would adding a retry around the unreliable shared step be a reasonable first move?
    It may be worth doing for the dependency itself, but it does not address why unrelated scenarios failed. Retrying makes the coupling quieter rather than removing it, and it slows every scenario in the file. Fix the structure first, then decide whether the underlying dependency still needs hardening or should be stubbed out of the path entirely.

saying these in an interview costs you the question

  • Calls it flakiness and adds a retry or a longer timeout
  • Quarantines the red scenarios without touching the coupling
  • Adds more shared steps so every scenario matches the setup
  • Reorders scenarios so the dependent ones run first
  • Refuses to repeat a context line because it is duplication
  • Leaves slow technical setup written as business-readable prose

context