skip to content

Team Consequences

What a machine-assisted suite does to the people who own it: where the effort moves once drafting is cheap, and which judgements stay with a person. Interviewers probe whether the trade paid.

on this pageshow

questions

6

Which decisions stay with a person when a machine drafts test cases from a running application?

level: middleimportance: must knowfreq 62%

answer

  1. Some judgements never delegate
  2. Ask where correct comes from
  3. Drafting from behaviour copies defects too
  4. Intent lives outside the running system
  5. Ranking consequence needs product knowledge

basics

~20 s

Two: what correct behaviour actually is, and which behaviours carry real consequence. A tool drafting from a running system can only describe what it observes, so it will happily record a defect as the expected result.

solid answer

~40 s

A drafting tool reads the system as it exists and generalises from it. That makes it strong at mechanical breadth — fields, paths, boundary values — and structurally unable to supply the **oracle**, the statement of what the system *should* do. Intended behaviour comes from a requirement, a rule, an agreement or the person who asked for the feature; if current behaviour is wrong, a case drafted from it locks the defect in as expected. The second reserved decision is **weight**: which behaviours cost something real when they break. That rests on knowledge the tool does not hold — what the business loses, what changed last week, what an incident cost. Delegate the drafting and the enumeration; keep the definition of correct and the ranking of consequence with the people who own the product.

code

pseudocode · 10 lines
pseudocode
# drafted from observation: records what the system did
case "coupon reduces the order total":
    order = place_order(goods = 80.00, delivery = 10.00, coupon = "SAVE10")
    assert order.total == 81.00          # literal copied from one live run

# after a person supplies the oracle: the rule is stated in the case
case "coupon reduces the goods subtotal before delivery":
    order = place_order(goods = 80.00, delivery = 10.00, coupon = "SAVE10")
    expected_total = 80.00 * (1 - COUPON_RATE) + 10.00   # rule, not observation
    assert order.total == expected_total   # fails today, and should

go deeper

for a junior

Be ready to say plainly that a tool drafting from a working system records what it does, not what it should do. Know that an expected value in a case has to trace back to a rule somebody stated.

for a middle

Explain the mechanics: where the oracle comes from, why an expectation recomputed from a stated rule survives legitimate change better than a copied literal, and which parts of a case a tool can safely widen.

for a senior

Show how you keep the two reserved decisions visible in practice — expectations that cite their source, review attention spent on the claim rather than the scaffolding, and drafted cases that are never allowed to define correctness on their own.

for a principal

Own the position that intent and the ranking of consequence are not delegable, and say what your team does when the pressure is to ship whatever the tooling produced. Name what you would give up first if throughput demanded it.

## Two halves of a test case Every automated case has a mechanical half and a judgement half. The mechanical half reaches a starting state in the application, supplies inputs, and captures what came back. The judgement half is the **oracle**: the statement of what the result should have been. A tool that drafts cases by reading an application — its code, its recorded traffic, its interface shapes — is genuinely strong at the mechanical half and structurally unable to supply the judgement half. Everything it can observe describes what the system *does*. Correctness is a claim about what the system *ought* to do, and that claim is not present in the artefact being read. It lives in a requirement, a pricing rule, an agreement with another team, a regulation, or the head of the person who asked for the feature. This is not a limitation that better tooling removes. Give a drafting tool a perfect view of the implementation and it produces a perfect description of the implementation, defect included. ## Decision one: what "correct" means | What the tool reads | What it can produce | What it cannot decide | |---|---|---| | The current implementation | Cases that pin present behaviour | Whether present behaviour is right | | Recorded live traffic | Realistic inputs and sequences | Which observed outcomes were intended | | An interface or screen shape | Field-level and boundary coverage | Which fields carry real consequence | | A written requirement | A case restating that requirement | Whether the requirement is still current | Work a concrete case. Suppose a coupon should take ten percent off the goods subtotal, before delivery charges. The implementation applies it after the delivery charge. A drafting tool observes an order of 80.00 in goods plus 10.00 delivery, sees a total of 81.00, and writes `assert total == 81.00`. The case is green forever, it is genuinely repeatable, and it has pinned the defect as the specification. A person applying the rule computes `expected_total = 80.00 * 0.9 + 10.00`, gets 82.00, and the case fails on its first run — which is the case doing its job. The practical form of the reserved decision is therefore small and checkable: **an expected value must trace to a stated rule, not to an observed run.** Where the case recomputes the expectation from the rule, the rule is visible inside the case and a reviewer can disagree with it. Where the case carries a literal copied from a run, the only claim it makes is "this has not changed" — worth having when that is exactly the claim you meant, and misleading when it is not. ## Decision two: which behaviours carry consequence The second reserved decision is about weight. Drafting spreads effort evenly across the surface it can see; consequence is not evenly distributed. Which behaviours matter turns on knowledge that is nowhere in the artefact being read: - What it costs when this is wrong — a refund, a regulatory report, an unrecoverable data change, a queue of angry customers. - What changed last week, and what is about to be redesigned, so effort is not spent freezing a surface that is moving. - What has broken here before, and how it was noticed — by a check, or by a customer. - Which behaviours other teams depend on, and which are internal detail that should stay free to change. None of that is recoverable from the running system. A tool can enumerate; only a person can rank. Keep the scope narrow when you answer: the reserved judgement is *which behaviours carry consequence*, which is a different question from how a given check is best run. ## What the machine legitimately buys Holding the two reserved decisions is not a refusal of the tool. Once a person has fixed the claim and the ranking, drafting earns its place at: 1. **Enumeration** — the combinations, boundary values and error paths implied by a rule someone already stated. 2. **Mechanical bulk** — navigation, arrangement, cleanup, and the repetitive scaffolding of a scenario. 3. **Shape consistency** — new cases written the same way as the ones already in the suite. 4. **First drafts of the unloved** — the error paths and odd states people skip when writing by hand. The division is stable and easy to state out loud: **the person states the claim, the machine widens it.** A team that keeps the claim and the ranking and delegates the widening gets a real speed-up. A team that delegates the claim gets a suite that is large, green, fast, and asserts that the product should behave exactly as it currently misbehaves. ## How the split shows up in review Attention should follow the same division. Time spent checking that a drafted case runs, navigates and tidies up after itself is time spent on the half the machine is good at. The half worth a person's attention is the assertion: what does this case claim, where did the expected value come from, and would a reasonable change to the product make it fail for the right reason? A case that cannot answer the second of those has not been reviewed, however long someone looked at it.

  • A drafted case asserts the exact total one live run produced. What do you do with it?
    Treat the number as evidence, not as the specification. Recompute the expectation inside the case from the rule that governs it, or trace the literal back to a requirement and record where it came from. If nobody can say why that total is right, the case proves only that behaviour has not changed — worth keeping when that is exactly the claim you wanted, and misleading otherwise.
  • Where can drafting genuinely shorten the work without touching either reserved decision?
    In enumeration and scaffolding. Once a person has fixed the claim and the ranking, a tool is good at spelling out combinations, boundary values and error paths that follow from the stated rule, and at writing the arrangement and cleanup around them. The person states the claim; the machine widens it. That is a real saving and it leaves the judgement where it belongs.

A transcription machine can write down every note a band actually played, exactly as played. It still cannot tell you the song was meant to be in a different key.

saying these in an interview costs you the question

  • Says a tool can infer intended behaviour from the code
  • Treats observed output as the specification
  • Cannot say where an expected value came from
  • Delegates the ranking of risk to whatever wrote the cases
  • Assumes a green drafted case proves the behaviour is right
open as a page

When a drafting model writes most of a suite's cases, where does the team's effort move instead of vanishing?

level: middleimportance: must knowfreq 55%

basics

~20 s

Authoring effort falls; review, triage and repair rise. Cheap drafting increases how many cases a team must read, explain and keep truthful, so the cost moves downstream and is charged against a bigger standing suite.

open as a page

An approval gate admitting cases into a regression suite now receives ten times as many machine-drafted candidates. How do you design it so refusal stays real?

level: principalimportance: should knowfreq 46%

basics

~20 s

Bound intake to review capacity, default to refusal when a candidate is undecided, and make rejection cheap and reasoned. Tier review depth by what a case can stop, then watch the refusal rate: a gate that never refuses is decoration.

open as a page

Leadership asks whether machine-assisted case authoring paid for itself. What must that before-and-after comparison measure to be honest?

level: principalimportance: should knowfreq 40%

basics

~20 s

An honest comparison prices a case's whole lifetime on both sides - drafting, review, triage, repair - over a window long enough for upkeep to appear, states the expected value in advance, and names the confounders that could explain the result.

open as a page

Why should accepting a self-repaired element locator, or promoting a drafted case into the regression set, be recorded rather than automatic?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Both acts change what the suite claims, and a silent default means nobody owns the claim. A recorded acceptance gives the change an author, a reason and a reversible history; a default gives it none of those.

open as a page

Which symptoms show a machine-grown test suite has outgrown the upkeep capacity of the team that owns it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The suite stops being read. Failures are answered by re-running rather than diagnosing, nobody can say what a case protects, review turns into sampling, and repairs are deferred. All of it happens while the suite still passes.

open as a page