A scenario's Then only checks that the playlist screen shows the new order, and a silent reordering corruption survived nine nightly runs — how do you strengthen it?
answer
- The assertion was green the whole time
- Ask what the promise actually was
- The actor's own view is the least trustworthy
- Add the unchanged-contents invariant
- Quiet failures need aimed assertions
basics
~20 sAssert the promised outcome where the effect lives, not the nearest visible proxy. Re-read the stored order through the service's own read path, check it from a second collaborator's view, and assert no track was lost or duplicated.
solid answer
~50 sThe Then is the oracle — the statement that decides whether the behaviour was wrong — and a screen is only a proxy for it. A cached or optimistically updated view can show the intended order while the persisted order is scrambled, which is exactly how a corruption stays invisible through a 6-hour nightly run repeated night after night. Strengthen it in three moves. First, name the promise rather than the display: the order the service returns on a fresh read is the outcome. Second, observe it from an independent vantage point — a second collaborator, or a reload that cannot be served from the writer's own cached state. Third, assert an invariant alongside the ordering: the same 14 tracks, none lost, none duplicated. Keep the assertions in domain terms; reaching into storage internals inside a step buys sensitivity at the price of a scenario that breaks on every schema change.
code
pseudocode · 2 linesWhen the editor moves the ninth track to first position
Then the playlist view shows that track at the topgo deeper
Remember that the Then should name what the product promised, not what appears on a screen. If you can explain why a visible confirmation is not proof that anything was saved, you are at the expected level here.
Explain how an optimistic view can agree with intent while stored state disagrees, and be ready to rewrite a display assertion into a fresh-read assertion plus an invariant about unchanged contents.
This is your question. Show that you diagnose a false-green suite by reading the assertions rather than the failures, and that you can argue the coupling tradeoff between a public read path and direct inspection of internals.
Own the systemic version: how you audit a whole suite's oracles, which rule families you strengthen first, and how you keep the added assertions from turning into slow, bundled scenarios nobody can read.
### The Then is an oracle, and proxies make weak oracles An oracle is whatever decides that observed behaviour is wrong. In a scenario, the Then is it. The failure described here is not a missing assertion — there was an assertion, and it was green for nine consecutive nightly runs while roughly one reorder in 340 wrote a scrambled order to storage. The assertion was simply pointed at a proxy: the view the actor who made the change was looking at. That view was updated locally the moment the gesture completed, so it agreed with the intent regardless of what was stored. A green scenario, a broken promise, and a suite that produced false assurance for over a week. This is the specific reason "assert an observable outcome, not a screen state" is a rule rather than a stylistic preference. Screens are frequently *where* an outcome becomes visible, so the rule is not "never look at a screen". It is: name the promise in the Then, then check it where the promise actually lives, using the weakest coupling that still detects a violation. ### Move one — restate the promise Start with the text, before touching automation. "Then the playlist screen shows the new order" describes a display. "Then the playlist order starts with that track for every collaborator" describes what the product promised, and it immediately implies checks the first wording never suggested: someone other than the actor, and a state that outlives the gesture. Rewriting the sentence is what surfaces the missing checks; teams that go straight to the automation usually add a wait and call it fixed. ### Move two — observe from an independent vantage point The actor's own session is the least trustworthy place to confirm a persisted effect, because everything between the gesture and the pixels may be optimistic. Independence can be bought cheaply: re-read the playlist as a second collaborator, or re-request it in a way that cannot be served from the writer's local state. This is what would have caught the corruption on the first night — the scrambled order was in storage the whole time, visible to anyone who read it fresh. ### Move three — assert the invariant, not only the example Ordering rules have a companion invariant that is cheap to state and catches whole classes of corruption: the multiset of tracks is unchanged. A reorder that loses one track, duplicates another, or drops a track's attribution is a corruption the position assertion may not notice, because position one can be correct while position eleven is a duplicate. One extra assertion line — the playlist still holds the same 14 tracks, each exactly once — turns a narrow example into a check with real defect-detecting power. ### How far down should the assertion reach? There is a genuine tension. Reaching directly into storage from a step is the most sensitive check available and the most brittle: it couples business-facing prose to a schema, it breaks whenever storage changes for reasons no reader of the scenario cares about, and it lets a scenario pass while the read path that users depend on is broken. The usual resolution is to assert through the service's own public read path — the same route a collaborator's client would take — which is stable, observable and still sees the real stored state. Reserve direct inspection of internals for a narrow diagnostic check outside the business-facing suite, if at all. ### Silent versus loud failures The diagnostic question worth asking of any Then is: *what would this assertion look like if the behaviour were quietly wrong rather than loudly broken?* Loud failures — an error, a blank page, a rejected request — are caught by almost any assertion, including weak ones. Quiet failures are caught only by assertions aimed at the promise. A suite whose Thens are all proxies is tuned entirely for loud failures, which is why it can run for six hours a night, stay green, and still miss data being corrupted underneath it. ### What not to conclude Two overcorrections are common. The first is to bolt on many more assertions per scenario until each one covers everything, which reintroduces bundled Thens that stop at the first failure and scenarios nobody can read. Prefer one promise per line, and a separate scenario for a separate rule. The second is to conclude that the interface should never be observed. If a rule is about what a collaborator sees, seeing it is the outcome — the point is that the scenario must say so deliberately, rather than defaulting to the display because it was the easiest thing to check.
- Why not assert against storage directly, since that is where the corruption was?Because it buys sensitivity with coupling. A step that reads storage breaks whenever the schema changes for reasons no reader of the scenario cares about, and it can pass while the read path real collaborators depend on is broken. Asserting through the service's own public read path sees the same stored state, stays stable across internal changes, and keeps the prose in domain terms.
- How would you find the other scenarios with the same weakness?Read the Then steps alone, ignoring the rest of each scenario, and mark every one whose subject is a display element or a message rather than a promised outcome. That list is the backlog. Ordering, permission and money rules are worth doing first, because those are the ones where a quiet wrong answer is more damaging than an outage.
- Does adding these assertions risk making scenarios slower or flakier?A fresh read and a second collaborator's view do cost time, and a read that races the write introduces nondeterminism if it is not waited for properly. The mitigation is to assert on a state that has settled — a fresh request that must reflect a completed write — rather than sampling immediately after the gesture. If a rule cannot be asserted deterministically, that is a fact about the system worth surfacing.
Confirming a bank transfer by reading the screen you just typed into is not confirmation; asking the other account whether the money arrived is.
saying these in an interview costs you the question
- Says a green suite proves the behaviour was correct
- Asserts a confirmation message and calls it the outcome
- Confirms a persisted change from the writer's own view
- Adds a wait instead of aiming the assertion at the promise
- Reaches into storage tables from a business-facing step
- Bundles every new assertion into one long Then line