How do you decide how fine-grained the numbered steps in a stored manual test case should be, and what goes wrong at each extreme?
answer
- does the row carry distinct information
- click-level rows ripple on every redesign
- one row, six actions, no location
- setup above the table
- intent, not control names
basics
~20 sGive an action its own row when its expected result is separately observable and a divergence there would mean something different from a divergence in the row beside it. Too fine is unmaintainable noise; too coarse leaves a divergence naming nothing.
solid answer
~50 sThe rule that decides a row is whether it carries distinct information: **a step earns its own row when its expected result is separately observable and a failure there would mean something different from a failure in the neighbouring row.** Everything else is packaging. Steps written at click-and-keystroke level make every small interface change ripple across hundreds of rows, and executors start skimming a wall of rows that mostly say nothing — which defeats the point of writing them. Steps written at whole-journey level ("complete the checkout") produce a divergence that names the case and nothing inside it, and an expectation that is really four claims wearing one sentence. Setup goes above the table in preconditions, and rows are written at the level of intent rather than of the current layout, so the case survives a redesign that broke no behaviour.
go deeper
Be ready to say that a step should be one action a person can perform with one thing to observe afterwards, and that logging in and navigation belong in preconditions.
Explain both failure modes concretely: click-level rows ripple on every interface change, and journey-level rows leave a divergence that names the case but not the action.
Show the deciding rule and apply it. A row earns its place when its observation is distinct, and you should be able to diagnose a bloated case and say where you would split it.
Own that granularity varies by area and cannot be set once as house style. Defend finer rows where evidence is needed and coarser ones where speed matters, in the same repository.
## The rule that decides a row There is one test, and it is about information rather than about size: > **A step earns its own numbered row when its expected result is separately observable, and when a divergence there would mean something different from a divergence in the row beside it.** Everything else is packaging. Two actions that always succeed or fail together tell you the same thing whether they are one row or two, so the second row buys nothing and costs maintenance forever. ## What goes wrong when steps are too fine - **Maintenance ripple.** Steps written as "click the field, type the value, click next" are descriptions of the current interface. A layout change that broke no behaviour forces edits across every case that describes it, and those edits are usually done under time pressure and half-finished. - **Reading fatigue.** Executors skim. A wall of rows whose expectations are mostly non-observations trains people to stop reading the expectation column, and then the rows that carry a real assertion are skimmed with the rest. - **Authoring cost crowds out coverage.** Time spent decomposing one journey into forty rows is time not spent writing the case for the scenario nobody has covered yet. - **Fragile precision.** A row that names an exact control is wrong the moment the control is renamed, and the case then produces arguments about whether the product or the case is at fault. ## What goes wrong when steps are too coarse - **A divergence names nothing.** If one row covers six actions, the next reader has to reconstruct which of the six broke — usually from a free-text comment written in a hurry. - **Compound expectations.** "The order is placed, the confirmation is shown and the stock count drops" is three claims in one row. Two can hold while the third does not, and the row has no way to say so. - **Unrunnable by a newcomer.** A coarse row silently assumes domain knowledge. Whoever wrote it can run it; the person covering next month cannot. ## Setup belongs above the table A step table that opens with four rows of logging in, dismissing a banner and navigating to a screen is spending its most-read rows on getting to the starting line. Push that into preconditions. The dividing line: **if a row exists only to reach the interesting behaviour, it is setup; if reaching it is the thing being checked, it is a step with its own expectation.** The benefit is not only tidiness — it keeps the step table about the behaviour under test, so what the case asserts is visible at a glance. ## Intent over interface Write the action at the level of what the user is trying to do, and the expectation at the level of the domain rather than the current screen: - Brittle: "Click the green Continue button in the top-right, then click Pay." - Durable: "Pay for the basket with the saved card. — Expected: the order is confirmed and the basket is empty." The second survives a redesign; the first has to be rewritten by whoever notices, which is usually the executor at the worst possible moment. ## Symptoms and what they usually mean | Symptom | Likely cause | Usual fix | |---|---|---| | One layout change forced edits to dozens of cases | steps written at control level | rewrite the actions as intent, expectations in domain terms | | Executors keep asking which screen a row means | rows too coarse, knowledge assumed | split the row at the point where the observation changes | | The case runs to thirty rows | several cases wearing one title | split where the preconditions could honestly be restated | | Half the expected results are placeholders | rows that carry no distinct information | fold each into the following row's action | ## The judgment underneath Granularity is not a house style you can set once. A case covering a regulated flow that somebody may have to evidence later legitimately carries finer rows than a case covering a stable internal screen. The thing to keep constant is not the row count but the **question the row count answers**: when this diverges, will the record say something a reader can act on without asking the person who ran it? If the answer is yes, the granularity is right, whether the case has five rows or fifteen.
- A case has grown to thirty rows. What do you do with it?Look for the point where the preconditions could honestly be restated — that is usually the seam between two cases wearing one title. Splitting there gives each part a title that reads alone in a failure list and a shorter path from setup to the behaviour under test, and stops one divergence early on hiding everything after it.
- Does the right granularity differ between areas of the same repository?Yes. A flow whose evidence someone may have to produce later justifies finer rows than a stable internal screen. What stays constant is the question the rows answer: when this diverges, will the record say something a reader can act on without asking whoever ran it?
saying these in an interview costs you the question
- Writes each click and keystroke as its own row
- Puts six actions and three assertions in one row
- Names exact controls and colours in every step
- Starts the step table at logging in
- Applies one fixed row count as house style