skip to content

An outer acceptance scenario has stayed red for eleven days of a three-week release train. How do you split it?

level: seniorimportance: should knowfreq 40%

answer

  1. Long red is a sizing signal
  2. First ask why it is red
  3. Thinner end to end, not layer by layer
  4. Each piece acceptable on its own

basics

~20 s

A scenario still red after a couple of days signals a slice that is too large, not slow progress. Confirm it is failing for the right reason, then split it by behaviour into thinner end-to-end cases, never by architecture layer.

solid answer

~50 s

An outer loop open for eleven days has stopped driving anything: it no longer tells you what to do next, and it represents work nobody has accepted or demonstrated. Before splitting, check the scenario is failing for the right reason — a scenario that asserts an incidental detail can stay red long after the behaviour works. If the slice really is too big, split it vertically. Start with a walking skeleton: the narrowest submission-to-result path that is complete end to end, which lands the wiring everything else reuses. Then add one rule per scenario. The split to avoid is architectural — one scenario per layer — because none of those is a behaviour anyone can accept and they all stay red until the last lands. The test for a good split: could this scenario go green while the others stay unwritten, and would a stakeholder call it progress?

code

pseudocode · 12 lines
pseudocode
# one scenario, red for 11 days
scenario("operator submits a job and receives all requested renditions")

# split by behaviour: a walking skeleton, then one rule each
scenario("operator submits a job and receives the default rendition")
scenario("operator requests two extra renditions and receives all three")
scenario("operator submits a source in an unsupported codec and is told why")

# the split to avoid: by layer, none of these is acceptable alone
# scenario("submission handler accepts the request")
# scenario("worker picks up the queued job")
# scenario("storage records the finished rendition")

go deeper

for a junior

Know the basic idea: an acceptance scenario is meant to close within a day or two, and one that stays red much longer usually means the piece of work was chosen too large rather than that the team is slow.

for a middle

Be able to describe a vertical split — a walking skeleton first, then one rule per scenario — and explain why splitting into request handling, worker and storage scenarios produces three long-red scenarios instead of one.

for a senior

Demonstrate the diagnosis: check whether the scenario is red because the behaviour is missing or because the assertion encodes an incidental detail, then re-slice and say what you did with the original umbrella scenario.

for a principal

Own the sizing discipline across a team: how slice size is agreed before work starts, what the standing timebox is, and how you keep a release train forecastable when nobody has accepted the work in flight.

### The timebox is the diagnosis The outer loop is meant to be a driver: something you re-run repeatedly during a slice to see the failure message change. A scenario that has been red for eleven days has stopped driving anything. It is no longer telling you what to do next, because "everything" is what to do next. Most practitioners treat a day or two as the comfortable life of an open outer loop and anything beyond that as a signal — not that the team is slow, but that the slice is wrong. The reason to care is not tidiness. An outer loop open for eleven days of a three-week release train means eleven days of work that nobody has accepted, that has not been demonstrated, and whose remaining size is unknown. On a train with fourteen working days left, that is a slice you cannot honestly forecast and cannot safely drop. ### Check the oracle before you split Before concluding the slice is too big, confirm the scenario is red for the right reason. A scenario that encodes an incidental assumption stays red long after the behaviour works. On a video-transcoding queue, a scenario asserted that three requested renditions were reported in the order they were requested. The workers completed out of order and reported as they finished, so the scenario failed for six of the eleven days while the behaviour underneath had been correct since day four. The ordering assumption was never a requirement; it was an accident of how the example was first written. Fixing the oracle — asserting that all three renditions exist and are playable, without a claim about order — closed most of the gap. That failure mode is common enough to check first: an outer loop can be red because the behaviour is missing, or because the scenario is wrong. Only the first calls for a split. ### How to split: by behaviour, never by layer The tempting split is architectural: one scenario for submission handling, one for the worker, one for result storage. It is the wrong split, because none of those is a behaviour anyone can accept. They all stay red until the last one lands, so you have swapped one long-red scenario for three, and gained nothing but bookkeeping. The split that works is a thinner **vertical slice** — a smaller end-to-end behaviour that is complete on its own: * **The walking skeleton first.** The simplest submission that produces one default rendition end to end. Narrow, real, acceptable, and it lands the wiring that everything else reuses. * **Then one rule per scenario.** Requesting extra renditions. An unsupported source codec. A submission that exceeds the size limit. Each opens and closes its own outer loop in hours, and each can be demonstrated the day it lands. The test for a good split is blunt: *could this scenario go green while the others stay unwritten, and would a stakeholder recognise the result as progress?* If either answer is no, the split is architectural. ### Signals to split, before eleven days You rarely need to wait for the timebox to expire. The earlier signals are: * The number of inner cycles beneath the scenario keeps growing and the scenario's failure message stops changing. * You cannot demonstrate anything to anyone, because nothing observable has landed. * The diff has grown to the point where you are worrying about the merge rather than the behaviour. * You cannot answer "how much is left?" with anything better than "most of it". ### What to do with the eleven-day scenario Do not delete it — it is the statement of the whole behaviour and it is still the acceptance condition for the feature. Keep it, mark it as the umbrella case, and write the thinner scenarios in front of it. Land those one by one; the umbrella scenario goes green on its own when the last of them does, and at that point it either becomes a valuable end-to-end example worth keeping or a duplicate worth removing. Either way you get back the property the double loop depends on: a red outer loop that means *"today's work is unfinished"*, not *"this feature started sometime last month"*.

  • Why is splitting by architecture layer the wrong move?
    Because none of the pieces is a behaviour anyone can accept, and they all remain red until the last one lands. You replace one long-red scenario with three, gain no demonstrable progress, and lose the property that a green outer loop means somebody got something they asked for. Layer-shaped checks belong at the level beneath, not in the acceptance loop.
  • What signals tell you to split before the timebox expires?
    The failure message stops changing while inner cycles keep landing; the diff has grown until you worry about the merge rather than the behaviour; nothing observable can be demonstrated to anyone; and you cannot answer 'how much is left?' with better than 'most of it'. Any one of those is enough to stop and re-slice.
  • Do you delete the original oversized scenario once you have split it?
    Not immediately. Keep it as the umbrella acceptance condition for the whole behaviour and write the thinner scenarios in front of it. When the last thin scenario lands, the umbrella goes green on its own; then decide whether it remains a valuable end-to-end example or is now pure duplication and can be removed.

saying these in an interview costs you the question

  • Treats long-red as normal for a big feature
  • Splits by architectural layer instead of behaviour
  • Deletes the scenario rather than re-slicing the work
  • Never checks whether the scenario itself is wrong
  • Waits for the release train to end before re-slicing
  • Adds more parallel outer loops instead of thinner ones

context