skip to content

What may you commit while the outer acceptance scenario is still red?

level: middleimportance: should knowfreq 44%

answer

  1. Only one thing is allowed to be red
  2. Distinguish unfinished from broken
  3. Every inner cycle ends committable
  4. Mark it pending, not merely known

basics

~20 s

Commit every inner cycle that leaves the build working and all unit-level tests green. The only permitted red is the acceptance scenario for the unfinished slice, and it must be visibly marked as in progress so nobody reads it as a broken build.

solid answer

~50 s

The rule is that exactly one failure is allowed in the tree — the acceptance scenario that opened the slice — and that failure has to be distinguishable from a genuine break. So each inner cycle that ends green gets committed: compiling code, all unit tests passing, nothing previously working now broken. What you never commit is a red unit test, half-written code, or a weakened assertion that makes the scenario pass. To keep the signal honest, teams either mark the unfinished scenario as work in progress so the gating run reports it as pending, or keep it out of the shared branch until the slice lands, or put the incomplete path behind a toggle so it is unreachable. Holding everything back until the scenario is green is the worse option: it turns a two-day slice into a two-day unmerged diff.

go deeper

for a junior

Remember the simple rule: you commit work whose unit tests all pass and whose build works, even though the acceptance scenario for the slice is still failing. Never commit a failing unit test to fix tomorrow.

for a middle

Explain the mechanics of keeping the signal honest — marking the scenario as pending, or a toggle over the incomplete path — and say why deleting an assertion to reach green destroys the completion signal the whole slice depends on.

for a senior

Show the failure mode you have lived through: a run that stayed red for days, a real regression hidden inside it, and the convention you introduced afterwards to keep unfinished and broken visually distinct.

for a principal

Own the tradeoff between an honest signal and continuous integration: which convention the organisation adopts, what it costs in review discipline, and how you would keep slices small enough that the question rarely arises.

### Two kinds of red A working tree in the middle of a slice contains one deliberate failure: the acceptance scenario that opened the slice. Everything else — every unit-level test, the compile, the static checks — is green, because each inner cycle ended green before the next one started. So the honest answer to "what may you commit?" is: *everything except a second kind of red*. The distinction that matters is between: * **Known red** — the outer scenario is failing because the behaviour it describes is not finished yet. This is expected, planned, and readable. * **Broken red** — something that used to work no longer does, or something was committed half-written. This is an emergency. A pipeline that renders both as the same colour destroys the value of the signal. Within a week people stop looking, and a genuine regression sits unnoticed behind "oh, that's just the in-progress scenario". Every workable answer to this question is really an answer to *how do you keep those two distinguishable*. ### Ways teams keep it distinguishable * **Mark the unfinished scenario as work in progress** so the gating run reports it as pending rather than failed, and re-enable it when the slice completes. The marker must be visible in review, and leaving one behind on a merged slice is a defect in its own right. * **Keep the unfinished scenario out of the shared branch** until the slice lands, committing only the inner work. Cheap, but it costs you the record of intent in the shared history and encourages the branch to live longer. * **Put the incomplete path behind a toggle** so the partial behaviour is unreachable in the running system. This is what lets the inner work ship continuously while the slice is still open. The choice is a team convention, not a right answer; what an interviewer wants is that you *have* one and can say what it costs. ### What never gets committed * A failing unit test "to be fixed tomorrow". The inner loop's contract is that it ends green; a red unit test committed is broken red. * A weakened assertion or a commented-out check that makes the outer scenario pass. Deleting the oracle to reach green destroys the completion signal the slice is built on. * Half-written code that does not compile, or a change that leaves a previously passing behaviour failing. ### A concrete case On a video-transcoding queue, a slice was opened by a scenario describing an operator submitting a job and receiving three renditions. It stayed red for two days while inner cycles landed: job validation, queue admission, worker dispatch, completion reporting. Eleven commits went in during those two days, each with the whole unit suite green, and the scenario marked in progress so the gating run reported one pending item rather than one failure. On the twelfth commit the marker came off and the scenario went green in the same change that completed the behaviour. Contrast the failure mode: on an earlier slice the same team left the scenario failing in the gating run "because everyone knows what it is". Four days later a genuine regression in the queue's retry path went unnoticed for a full day, because the run had been red the whole time and nobody read the detail. ### Reviewing a slice that is still open The commit stream of an open slice is readable in a way a single large drop is not. Each commit says which unit was added and why, and the acceptance scenario sitting red at the top says what the sequence is for. A reviewer joining mid-slice can see the intent before the implementation, which is one of the quieter benefits of the double loop: the failing scenario is documentation of work in flight, not just a gate. That only holds if the commits are honest. A commit that says "wire up worker dispatch" but also silently relaxes an unrelated check is worse than no commit at all, because it hides inside the one red result everyone has already agreed to ignore. ### Why committing at all matters The temptation is to hold everything back until the outer loop is green. That converts a two-day slice into a two-day unmerged diff, and the cost of integration grows faster than the size of the change. Committing each inner green keeps the change small, keeps others' work rebasing cleanly, and keeps the risk of the slice proportional to its size. The double loop only stays cheap if the outer loop's redness is allowed to coexist with a continuously integrated trunk — which is precisely why the bookkeeping around "known red" is worth being strict about.

  • What is the risk of leaving the unfinished scenario failing in the shared gating run?
    People stop reading red. Once the run is expected to be red for days, a real regression hides inside the same colour and can sit unnoticed for a long time. If the scenario stays in the run, it must be reported as pending or skipped-with-reason rather than failed, so a genuine failure is still visually distinct.
  • How do you stop a work-in-progress marker from being forgotten after the slice lands?
    Removing it is part of the same change that completes the behaviour — the slice is not done while the marker is there. Beyond discipline, teams make the markers visible: list them in the run summary, review them explicitly, and treat one older than a couple of days as a defect to raise rather than a nuisance to ignore.

saying these in an interview costs you the question

  • Says nothing may be committed until the scenario passes
  • Commits a red unit test to be fixed later
  • Weakens or deletes the scenario's assertion to reach green
  • Leaves the failing scenario in the gating run unmarked
  • Cannot distinguish unfinished-red from broken-red
  • Keeps the whole slice on a branch for days to avoid red

context