When a keyword-driven failure names only a row, how do you make it diagnosable?
answer
- the stack describes plumbing, not intent
- rebuild what indirection removed
- push a frame per keyword entry
- fail at the precondition, not six steps later
- judge the layer on time-to-diagnose
basics
~20 sRebuild at the keyword boundary what the machine stack no longer says: print the keyword invocation path, the step index, the bound arguments, a subject identifier from the system, expected versus observed, and per-row artefacts.
solid answer
~50 sIndirection is what you pay for a table-driven suite: the failing frame is a keyword implementation and its caller is the engine's own loop, so the stack describes plumbing rather than intent. The fix is to construct the missing information deliberately. Push a frame on entry to every keyword and print that logical stack on failure, so the report reads `prepare envelope > send invitation` rather than naming one generic body. Add the row label and index, the step number, the arguments as bound after defaults, an identifier from the system under test for joining to its logs, and artefacts saved under a per-row path. Make each keyword check its preconditions on entry so a step that quietly did nothing fails immediately instead of surfacing as nonsense six steps later. Then judge the layer on time-to-diagnose, and collapse it if only engineers author rows.
code
pseudocode · 20 lineskeyword_stack = []
function invoke(keyword, row, step_index, arguments):
keyword_stack.push(keyword.name)
try:
keyword.run(arguments)
catch cause:
report(
"case : " + row.label + " (row " + row.index + ")",
"keywords : " + join(keyword_stack, " > "),
"step : " + step_index + " of " + row.step_count,
"arguments : " + render(arguments), # after defaults applied
"subject : envelope " + context.envelope_id,
"expected : " + row.expected,
"observed : " + cause.observed,
"artefacts : " + save_artefacts(row.index)
)
fail(cause)
finally:
keyword_stack.pop()go deeper
Understand why a failure in a table-driven suite is harder to read than one in a hand-written case: the code that failed is generic, so the report has to carry the identity of the row and step itself.
Be able to list what a good failure message includes — row label, step index, bound arguments, expected versus observed — and explain why precondition checks inside each keyword prevent misleading late failures.
Demonstrate you have lived through this: describe rebuilding a logical invocation stack, saving per-row artefacts, correlating a subject identifier with server-side logs, and diagnosing a defect that lived in keyword composition rather than in any single keyword.
Own the stopping condition. Decide when the vocabulary and its interpreter cost more than they return, measure time-to-diagnose rather than aesthetics, and be willing to collapse the layer back into ordinary reusable code.
## Why the stack trace disappears In a hand-written case, a failure points at a line you wrote for that case, in a file whose name says what it covers. Every layer of indirection erodes that. In a data-driven suite the failure points at one generic body plus a row. In a keyword-driven suite it points at the interpreter that was walking the table — the frame that fails is the keyword implementation, and the caller above it is a loop inside the engine, not the business intent. The machine stack is intact and useless: it describes the framework's own plumbing, not the case. That is the real cost of these frameworks and the reason interviewers ask about them. Everything below is about deliberately rebuilding, at the keyword boundary, the information the machine stack no longer carries. ## What a keyword-layer failure must print Treat the failure report as a designed artefact with a fixed shape: * **The case identity** — the row's label and its position, so the row is findable in a table of a thousand. * **The keyword invocation path** — the *logical* stack: `prepare envelope > send invitation`. Push a frame on entry to every keyword, pop on exit, and print the stack on failure. This one addition does more for diagnosis than any other, because it is what the machine stack should have said. * **The step index** — step 4 of 7, so the reader knows how far the case got. * **The arguments as bound**, after defaults and substitutions have been applied — not as written in the table. Most surprises live in the gap between the two. * **A subject identifier from the system under test** — the envelope id, the correlation id — so the failure can be joined to server-side logs. * **Expected and observed, in the domain's terms**, not a raw structural diff. * **Per-step artefacts**, saved under a path derived from the row, so no one re-runs the suite just to see what the screen looked like. ## Fail at the contract, not three steps later The second technique is to make each keyword check its preconditions on entry and fail there, naming itself and the precondition. A keyword suite's worst failures are the ones where step 3 quietly did nothing and step 6 reports a nonsense mismatch. A `send invitation` that asserts "an envelope exists and has at least one signer" turns a mystery at step 6 into a precise sentence at step 3. ## A worked incident A suite of 1,140 rows over 34 keywords covers a document e-signing flow. Nine days into a 3-week release train, 27 rows start failing with "expected 1 audit entry, found 2" — a duplicated side effect. The table shows nothing unusual: the failing rows use different roles, channels and expiry values, and each row invokes `send invitation` exactly once. The keyword invocation path is what breaks it open. On the failing rows the printed path reads `prepare envelope > send invitation` *and*, later in the same case, `send invitation` at the top level. A composite keyword had been extended that sprint to send the first invitation itself, so every row that also listed the step explicitly now sent two. Nothing in the table was wrong and nothing in either keyword was wrong in isolation; the defect lived in the composition, which is exactly the place indirection hides. Without the logical stack the same diagnosis is a bisect through a keyword library. Median time-to-diagnose for that suite ran about 41 minutes before the logical stack and per-row artefacts were added, and about 6 afterwards. Those numbers are one team's; the point is that time-to-diagnose is measurable, and it is the metric this layer should be judged on. ## Keeping it diagnosable as a standing discipline * **Log each keyword entry and exit with its bound arguments** at a level that is on by default in the suite's own runs. A keyword suite that only logs on failure cannot explain a case that passed for the wrong reason. * **Cap composition depth.** Two levels of composite keyword is usually enough; the third is where the incident above comes from. * **Version the vocabulary and review keyword changes like interface changes**, because a table row is a caller you cannot find with a code search across the repository. * **Keep a keyword reference with preconditions**, so a row author can predict a failure instead of discovering it. ## Knowing when the layer stops paying The honest senior answer includes a stopping condition. Ask who has authored a row in the last two release trains. If the answer is "engineers only", the table is a second, worse programming language that engineers maintain in addition to the real one, and the right move is to collapse it: keep the keyword implementations as ordinary reusable functions in code, delete the interpreter, and get the machine stack back. Keep the layer when domain experts genuinely author rows, when the same step sequence is reused across many cases, or when the same table drives more than one interface. Judge it on time-to-diagnose and on who authors, never on how elegant the table looks.
- Twenty-seven rows fail with "expected one audit entry, found two", and each row lists the sending step exactly once. Where do you look?At the composition, not at the rows. Print the keyword invocation path for a failing row: if it shows the sending action reached both from inside a composite preparation keyword and again at the top level, the duplicate side effect comes from a composite that was extended to do the step itself. Nothing in the table is wrong and neither keyword is wrong alone; the defect lives in the nesting, which is precisely what indirection hides.
- What does a precondition check inside a keyword buy that an assertion at the end of the case does not?It fails at the point where the state first became wrong, and it names the keyword and the missing precondition. Without it, a step that quietly did nothing surfaces as an unrelated mismatch several steps later, and the reader has to reconstruct which step was responsible. Checking on entry converts a mystery at step six into a precise sentence at step three, and it costs almost nothing at run time.
- How do you decide that the keyword layer is no longer worth maintaining?Ask who actually authored a row in the last two release cycles. If the answer is engineers only, the table is a second, worse language maintained alongside the real one; keep the keyword implementations as ordinary reusable functions, delete the interpreter, and recover the machine stack. Keep the layer when domain experts genuinely author rows, when step sequences are reused widely, or when the same table drives more than one interface.
saying these in an interview costs you the question
- Says just read the stack trace, without noticing it names the engine
- Adds retries around a keyword instead of reporting what it observed
- Prints the arguments as written rather than as bound after defaults
- Re-runs the whole suite to see what a failed row looked like
- Lets composite keywords nest without any depth limit
- Defends the layer on elegance rather than on who authors rows