skip to content

Behavior-Driven Development

Specifying software by behavior a business reader can check, and discovering those examples together before any code exists. Asked because it is a collaboration practice as much as an automation one.

on this pageshow

questions

29

In double-loop TDD, what are the outer and inner loops, and which turns green last?

level: juniorimportance: must knowfreq 58%

answer

  1. Two cycles, one nested inside the other
  2. One is slow and user-facing
  3. The fast one completes many times over
  4. Whatever opened the slice closes it

basics

~20 s

The outer loop is a failing acceptance scenario written in the user's language; the inner loop is the fast unit cycle beneath it. The inner loop goes green many times over; the outer scenario goes green last, when the behaviour is complete.

solid answer

~40 s

You open a slice of work by writing one acceptance scenario for the behaviour you are about to build and running it. It fails, and you leave it failing — that red is the statement of intent for the slice. Underneath it you run complete unit-level cycles: a failing unit test, the code that satisfies it, then restructuring while green. Each of those ends green and can be committed. You re-run the outer scenario as you go, and its failure message changes as the behaviour fills in. It turns green **last**, and that is the slice's definition of done. The point of nesting them is that the inner loop gives locality and design feedback while the outer loop gives direction and proves the parts compose into something a stakeholder asked for.

code

pseudocode · 13 lines
pseudocode
outer = scenario("operator submits a job and receives a playable rendition")

run(outer)                       # RED - opens the slice

while failing(outer):
    unit = next_missing_behaviour()
    write_failing_test(unit)     # inner RED
    make_it_pass(unit)           # inner GREEN
    restructure(unit)            # inner REFACTOR
    commit()                     # units green, outer still red
    run(outer)                   # failure message moves

run(outer)                       # GREEN - the last thing to turn

go deeper

for a junior

Be ready to name the two loops and their order out loud: the acceptance scenario opens the slice and fails, unit cycles run underneath, the scenario goes green last. That sequence is the whole answer at this level.

for a middle

Explain the mechanics: what re-running the outer scenario tells you mid-slice, why its failure message changing is the progress signal, and what each loop gives you that the other cannot — locality below, direction above.

for a senior

Show you have run this on real work: how you keep only one outer loop open, how the slice size is chosen so the loop closes in a day or two, and how you spot a scenario that went green too early because its assertion was weak.

for a principal

Own the argument for the practice: what a team loses when it adopts only the inner loop, or only the outer, and how you would introduce the double loop into a team that already has both kinds of tests but runs them as separate rituals.

### One workflow, two cadences Double-loop development nests two feedback cycles that run at very different speeds. The **outer loop** is a single acceptance scenario expressed in the language of the person who wants the behaviour: a concrete example of the system doing something useful, end to end. The **inner loop** is the fast developer cycle beneath it — a failing unit-level test, the code that satisfies it, then restructuring while green — repeated as many times as the slice needs. The sequence is what makes it a *double* loop rather than two separate habits: 1. Write the acceptance scenario for the next thin slice of behaviour and run it. It fails. That failure is not an accident to be cleaned up later; it is the statement of intent that opens the slice, and it is the thing that will tell you when you are finished. 2. Leave it failing. Drop down a level and run complete inner cycles against the units you need. Each inner cycle ends green. 3. Periodically re-run the outer scenario. Its failure message moves: first a missing entry point, then a missing behaviour, then a wrong value, then a passing run. 4. The outer scenario turns green **last**. That is the slice's definition of done, and it is the moment to look for the next slice. ### Why the outer loop must not go green early An outer scenario that passes before the inner work is done is a warning, not a win. It usually means the scenario asserts something the system already did — a screen appeared, a call returned without error — rather than the outcome the user cares about. The outer loop earns its place only if its green is expensive to obtain: it should be unreachable until the behaviour genuinely exists. Equally, the outer loop must not be the *only* loop. If you drive everything from the acceptance level, each defect arrives as a single coarse failure with no locality: you know the journey is broken, not which unit is wrong. The inner loop supplies locality and design pressure; the outer loop supplies direction and the claim that the parts add up. ### A worked slice Take a video-transcoding queue. The next slice of behaviour: an operator submits a source file and receives a playable rendition. The outer scenario is one example — a submitted job, a completed rendition, an observable result the operator could point at. It goes red on the first run because nothing accepts a submission yet. Beneath it, the developer runs a series of inner cycles: a job record that validates its source reference; a queue that hands one job to one worker; a worker that reports completion; a result the operator can read. Say six inner cycles across a day and a half, each ending green and each committed. The outer scenario is re-run after most of them, failing differently every time; the changing failure is the progress bar. When the last unit lands, the outer scenario goes green, the slice is done, and the next scenario opens the next slice. ### What the two loops are for The two loops answer different questions, and confusing them is the usual interview stumble: * The **inner loop** answers *"is this unit correct, and is its design bearable?"* It runs in seconds, is called constantly, and is thrown at the problem in large numbers. * The **outer loop** answers *"does the composed system do the thing that was asked for?"* It runs in seconds-to-minutes, is called far less often, and there is exactly one of it open at a time. Only one outer loop should be open at once. Two or three half-finished acceptance scenarios in flight means two or three unfinished slices in the working tree, which is the same problem as a long-lived branch by another name. ### The cadence in practice Nothing about the double loop requires a particular level of test or a particular way of writing the scenario. It requires only that a user-visible example is written first, kept failing, and used as the completion signal. Teams that adopt only the inner loop tend to produce well-tested units that do not compose into anything a stakeholder recognises. Teams that adopt only the outer loop tend to produce slow, coarse suites with no design feedback. The double loop is the claim that these two are one workflow, not two competing disciplines — which is exactly what an interviewer is checking when they ask about it.

  • How many outer loops should be open at the same time?
    One. Each open outer loop is an unfinished slice sitting in the working tree, so two or three at once is a long-lived branch by another name — unmerged work, unknown remaining size, nothing demonstrable. Finish the slice, let the scenario go green, then open the next one.
  • What does it mean if the outer scenario passes on its very first run?
    It is a warning that the scenario does not assert what you think. Usually it checks something the system already did — a call returned, a page rendered — rather than the outcome the user cares about. A useful outer loop should be unreachable until the behaviour genuinely exists, so strengthen the assertion before starting the inner work.
  • Why not drive everything from the acceptance level and skip the inner loop?
    Because a coarse failure has no locality: you learn the journey is broken, not which unit is wrong, so diagnosis time grows with the size of the system. The inner loop also applies design pressure at the level where the design actually lives. The outer loop proves composition; it is a poor debugger.

The outer scenario is the destination on the map and the inner cycles are the individual turns; you check the map repeatedly, but you only arrive once, at the end.

saying these in an interview costs you the question

  • Says the acceptance scenario is written after the code
  • Thinks the outer loop must stay green throughout
  • Describes the two loops as separate suites, not one workflow
  • Claims the inner loop is optional once scenarios exist
  • Cannot say which loop turns green last
  • Treats a first-run pass of the outer scenario as success

context

open as a page

What is a step definition in a scenario suite, and what binds it to a line of a scenario?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A step definition is a function holding the automation code for one scenario step. The runner matches the step's plain-language text against a pattern registered with that function, converts any captured values into arguments, and calls it.

open as a page

What does a Background section in a feature file do to the scenarios below it?

level: juniorimportance: must knowfreq 66%

basics

~10 s

A Background holds context steps that run again before every scenario in that file, so shared setup is written once. Every scenario inherits those steps, whether or not it needs them.

open as a page

In a behaviour scenario, what do the Given, When and Then steps each describe?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Given states the context that already holds before the behaviour under test. When names the single event that triggers it. Then states the observable outcome that must follow. All three describe behaviour, not interface mechanics.

open as a page

What makes a scenario report generated from the last run 'living documentation'?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The document is generated from scenarios that actually executed, so it describes only behaviour the suite exercised. When the system changes and the scenario does not, the scenario fails and the page visibly breaks instead of quietly going stale.

open as a page

Who are the Three Amigos in a discovery workshop, and what does each perspective contribute?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The Three Amigos are three perspectives that meet before a requirement is built: business, which owns what problem is being solved and why; development, which owns how it could work; and testing, which owns what could go wrong.

open as a page

What is step-definition sprawl in a scenario suite, and why does it raise maintenance cost?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Step-definition sprawl is a glue layer that grows one definition per sentence written instead of one per behaviour, so near-identical phrasings each get their own code. The library becomes large and duplicated, and every behaviour change touches many places.

open as a page

How does parameterizing a step definition stop the glue layer growing one definition per sentence?

level: middleimportance: must knowfreq 58%

basics

~20 s

Write one definition per behavior and capture the parts that vary as parameters. Sentences differing only by a value then reuse the same function, so the library grows with the number of behaviors, not sentences.

open as a page

When is a Scenario Outline with an Examples table better than separate scenarios?

level: middleimportance: must knowfreq 62%

basics

~20 s

A Scenario Outline states one behaviour once with placeholders, and its Examples table supplies a row per case, each running as its own scenario. Use it when rows differ only in data, not in meaning.

open as a page

How do you tell a declarative scenario step from an imperative one?

level: middleimportance: must knowfreq 62%

basics

~20 s

An imperative step spells out the interaction: open a page, type a value, press a control. A declarative step names the intent behind it — "a collaborator has editing rights". Declarative steps survive redesigns and stay readable to non-programmers.

open as a page

What may you commit while the outer acceptance scenario is still red?

level: middleimportance: should knowfreq 44%

basics

~20 s

Commit every inner cycle that leaves the build working and all unit-level tests green. The only permitted red is the acceptance scenario for the unfinished slice, and it must be visibly marked as in progress so nobody reads it as a broken build.

open as a page

A scenario has three When steps joined by And — what is wrong, and how do you fix it?

level: middleimportance: should knowfreq 48%

basics

~20 s

Three When steps mean three events, so a failure no longer says which one broke and the scenario documents a flow, not a rule. Demote the earlier events to Given facts and keep one event as the When.

open as a page

How do you trace a requirement to its scenarios and to those scenarios' latest run results?

level: middleimportance: should knowfreq 50%

basics

~20 s

Tag each scenario with a stable requirement identifier, then generate a matrix from those tags joined to the last run's results. Read it both ways: requirements with no scenario, scenarios with none, and requirements whose scenario failed.

open as a page

How does an Example Mapping session run, and what do its four card types capture?

level: middleimportance: should knowfreq 46%

basics

~20 s

Example Mapping is a timeboxed discovery technique using four card colours: yellow for the story, blue for each rule, green for a concrete example under a rule, and red for a question nobody in the room can answer.

open as a page

How do tags on scenarios keep feedback fast on a large scenario suite?

level: middleimportance: should knowfreq 48%

basics

~20 s

Tags are labels the runner selects on, so each pipeline stage runs only what it needs: a fast subset per change, the full pack nightly. The value comes from a small, enforced taxonomy, not the mechanism.

open as a page

An outer acceptance scenario has stayed red for eleven days of a three-week release train. How do you split it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A scenario still red after a couple of days signals a slice that is too large, not slow progress. Confirm it is failing for the right reason, then split it by behaviour into thinner end-to-end cases, never by architecture layer.

open as a page

How do you pass state between the step definitions of one scenario without using globals?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Put it in a context object the runner creates fresh for each scenario and hands to every step definition that asks for it. Process-wide or module-level variables survive between scenarios, which creates order dependence and breaks parallel runs.

open as a page

A feature file's scenarios keep failing on shared setup they don't use — how do you restructure it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Treat it as coupling, not flakiness. Shrink the shared context to what is universally true, push the rest into the scenarios that need it, group by rule or split the file, and move technical setup out of the prose.

open as a page

A scenario's Then only checks that the playlist screen shows the new order, and a silent reordering corruption survived nine nightly runs — how do you strengthen it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Assert the promised outcome where the effect lives, not the nearest visible proxy. Re-read the stored order through the service's own read path, check it from a second collaborator's view, and assert no track was lost or duplicated.

open as a page

A discovery workshop covered a utility billing tier rule, yet an off-by-one at the tier boundary reached production. How would you change how that workshop gathers examples?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Make boundaries a standing obligation of the session: for every rule with a threshold, the group states the value below, at and above it, plus the empty case, and writes down which side the threshold falls on in the business's own words.

open as a page

Your scenario suite has drifted into a 47-minute interface-driven pack. How do you get feedback back?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Diagnose first: per-scenario durations, which layer each scenario drives, and how many scenarios restate the same rule. Then re-point most scenarios below the interface, delete duplicates, and keep only a thin set of interface-driven journeys.

open as a page

How would you set and uphold a standard for scenario abstraction level across many teams?

level: principalimportance: should knowfreq 34%

basics

~20 s

Set testable heuristics, not word rules: a step should survive an interaction change, admit one behaviour only, and be judgeable by someone who has never used the product. Uphold it with exemplars, shared vocabulary and review.

open as a page

How do you decide whether publishing a scenario suite as the specification of record is worth its cost?

level: principalimportance: should knowfreq 40%

basics

~20 s

Establish a named audience with a recurring question first, then gather readership evidence over time. Generation is cheap; keeping every scenario worded for outsiders forever is the real cost, so narrow the published scope to where a reader actually exists.

open as a page

How do you judge whether Three Amigos discovery workshops repay their cost across a dozen teams?

level: principalimportance: should knowfreq 33%

basics

~20 s

Measure what the sessions catch, not that they happen. Watch questions raised before build, stories split or sent back, and requirement-misunderstanding defects found after build -- and vary the depth of discovery by how ambiguous and how costly the requirement is.

open as a page

How do you judge whether a scenario suite's plain-language layer still pays for its maintenance cost?

level: principalimportance: should knowfreq 38%

basics

~10 s

Weigh the translation layer's cost — slower authoring, indirection when debugging, an owned glue library — against evidence it buys participation and shared vocabulary. If only engineers read the scenarios, the cost buys nothing.

open as a page

In a feature file, how does a data table differ from a doc string as a step argument?

level: juniorimportance: nice to knowfreq 26%

basics

~20 s

Both attach data to the single step above them. A data table is pipe-delimited rows and columns, read as structured records; a doc string is a delimited block of free text passed through as one string, line breaks intact.

open as a page

A scenario run reports that two step definitions matched one step — what causes that and how do you fix it?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Two loaded patterns both match that sentence — usually a duplicate definition, an unanchored expression matching a substring, or a catch-all wildcard. Fix the glue: delete the duplicate, anchor the pattern, narrow the wildcard. Never reword the scenario.

open as a page

A generated specification report renders scenarios that never executed as documented behaviour. How do you fix it?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

Render every scenario with its run status and make not-executed visually distinct, stamp the revision, environment and any tag selection on the header, publish executed-versus-total counts, and fail the publishing step when too much of the suite did not run.

open as a page

Which level should the outer acceptance loop drive from when a full run is too slow to repeat?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Drive from the outermost level that still returns inside a development cycle — usually the service boundary just inside the delivered interface. Keep a small set of the same journeys running through the delivered interface on a slower cadence for wiring confidence.

open as a page