skip to content

AI-Augmented Testing

AI as the tool that tests an ordinary product, not a developer drafting unit tests: reviewing machine-written cases, self-repairing locators, judged verdicts. Interviewers ask what those cases bought.

on this pageshow

questions

28

Machine-drafted cases triple a regression suite's case count in a week. What has that growth not proven?

level: juniorimportance: must knowfreq 55%

answer

  1. Volume is not signal
  2. Reach is a union, not a sum
  3. Repeats change the count only
  4. Ask what became newly detectable

basics

~20 s

Case count measures how much was written, not how much can be detected. Drafted cases usually re-walk journeys existing cases already walk and check outcomes already checked, so the set of failures the suite can catch may not grow.

solid answer

~50 s

A suite's value is the **set of distinct failures it can detect**, and that set is a union, not a sum. Drafted cases come from the same requirement text and the same code the existing cases came from, so most of them repeat a journey that is already walked and check an outcome that is already checked; the count grows with drafting speed while the union barely moves. Structural coverage behaves the same way - it can only rise if something previously unexecuted now runs, and volume alone does not find unexecuted code. The honest reading of a batch is its **marginal contribution**: which behaviours became reachable that were not reachable before, and which wrong behaviour would now be caught that would previously have shipped. Growth in count without growth in either is inflation, and it is paid for on every run.

code

pseudocode · 11 lines
pseudocode
reach_before = union(executed_units(c) for c in existing_cases)
checked_before = union(asserted_outcomes(c) for c in existing_cases)

reach_after = reach_before | union(executed_units(c) for c in drafted_batch)
new_reach   = reach_after - reach_before        # empty => batch added no reach

for c in drafted_batch:
    if executed_units(c) <= reach_before and asserted_outcomes(c) <= checked_before:
        mark(c, "count only")

report(size(drafted_batch), size(new_reach), count_marked("count only"))

go deeper

for a junior

Be ready to say plainly that a case count counts artefacts rather than measuring what a suite can catch, and to give one concrete reason a brand-new case might add nothing at all.

for a middle

An interviewer expects the union effect explained: reach is the set of behaviours exercised and outcomes checked, so a batch adds nothing unless it enlarges that set. Show how you would compute the difference before and after.

for a senior

Demonstrate that you check a batch's marginal contribution before it merges, and that you can state its ongoing cost in run time and triage width rather than arguing from an intuition about size.

for a principal

Own the definition of progress your teams work to. Decide what a drafted batch must demonstrate - reach that did not exist, or a failure that would previously have shipped - and hold that line when volume is offered as evidence instead.

## What a case count counts A suite's **case count** is a count of artefacts: files, functions, rows in a plan. What a suite is *worth* is a different kind of object entirely - a **set**, namely the set of distinct wrong behaviours that would make at least one case fail. Call that set the suite's **reach**. Every case contributes two things to reach: the parts of the system its run causes to execute, and the outcomes it actually checks about that execution. The suite's reach is the **union** of those contributions across all cases, not their sum. A case whose executed parts and checked outcomes both already sit inside that union is adding a member to a set that already contains it. The count goes up by one. The reach goes up by nothing. That is the whole mechanism of a coverage illusion under bulk drafting, and it fits in one line: **counts add, reach unions.** ## Why drafted volume repeats itself Redundancy here is not carelessness; it is what the setup produces by default. - Drafts are made from the same sources the existing cases were made from - the same requirement text, the same interface description, the same code - so they rediscover the same journeys through the system. - The variation a drafting pass produces most cheaply is surface variation: different names, different step ordering, different sample values drawn from **inside one equivalence class**. - The drafting step usually has no picture of what the suite already reaches, so it cannot aim at the gap. Without that, heavy overlap with the existing suite is the expected outcome rather than a bad day. - Volume arrives faster than review capacity, so the moment where somebody once said "we already have that one" simply does not happen for most of a batch. - Where a description is ambiguous, a drafting pass tends to produce several cases for the *same* reading of it, rather than one case per reading - the opposite of what the ambiguity called for. ## Three numbers, and which of them can move | Number | Moves only when | Under a redundant batch | | --- | --- | --- | | Case count | any case at all is added | rises by the full size of the batch | | Structural coverage | code that no case executed now runs | flat, or drifts slightly on setup paths | | Reach | a new outcome is checked on behaviour the suite exercises | unchanged | The middle row is the one that fools people in both directions. A structural figure can only rise when some part of the system that nothing reached is now reached, so adding traffic to already-executed code cannot move it: a batch that leaves the figure flat has genuinely told you something. But a batch that *does* move it has told you only that something new **ran**. Reach needs the second ingredient as well - an outcome checked - before a wrong behaviour in that newly executed code could fail anything. ## Reading a batch honestly The honest read of a drafted batch is its **marginal contribution**, and getting it is mechanical: 1. Before the batch, record the union of executed units across the suite and the union of asserted outcomes. 2. Run with the batch and record both unions again, then subtract. 3. Mark every case whose executed set and asserted set both fall inside the old unions as **count only**. 4. For each case that remains, name the wrong behaviour it would catch, and check whether an existing case already catches it. The output is one small honest sentence. Not "200 cases added", but "six behaviours became reachable, four of which were already checked elsewhere". That sentence is what the batch actually did, and it is usually a great deal shorter than the batch. ## What the inflation costs Cases are cheap to produce and are never cheap to keep. - **Run time.** Every redundant case is paid for on every run, for as long as it exists, in wall-clock time and in machine capacity. - **Triage width.** One defect sitting under forty redundant cases produces forty failures to read, and forty failures look like a large problem rather than a single one. - **Reading cost.** Anyone changing the code beneath those cases must first work out what each of them claims, and forty near-identical claims take far longer to understand than one. - **False confidence.** The dangerous cost is the belief that an area is well covered because many cases mention it, because that belief stops the work that would actually have closed the gap. ## When volume genuinely pays Bulk drafting is not the problem; **unaimed** bulk drafting is. Point the same capacity at a gap that has already been identified - a list of branches nothing executes, an inventory of error paths nobody checks, the input classes a stated rule defines but no case supplies - and every draft lands outside the union by construction. Then count and reach rise together, which is the only condition under which the count was ever worth quoting.

  • How do you tell a redundant drafted batch from a deliberate set of equivalence-class variations?
    By what varies and why. A deliberate set walks one journey with inputs chosen from different classes - a boundary, an empty value, an over-long value - and every variation can fail on its own while the others pass. Redundant drafts vary surface detail such as names, ordering and phrasing while the inputs stay inside a single class, so no variation is capable of failing alone.
  • The batch cost nothing to write. Why not keep all of it and see whether it ever catches something?
    Because writing is the one cost that was cheap. Each case is paid for on every run in wall-clock time, in triage whenever it fails for an unrelated reason, and in the reading time of whoever must understand it before changing the code beneath it. A case with no path to a failure of its own charges that rent indefinitely and never repays it.

Photocopying a map does not enlarge the territory it shows. You end up carrying much more paper over exactly the same ground.

saying these in an interview costs you the question

  • Treats a larger case count as a stronger suite
  • Calls the batch a coverage win without checking what newly executed
  • Assumes drafted cases must be new because their names differ
  • Says volume costs nothing because nobody wrote it by hand
  • Cannot name a single wrong behaviour the new cases would catch
open as a page

When a suite repairs its own broken locators and still passes, which product defects does that green run hide?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A repair hides every change that broke the original anchor: a control that moved, was relabelled, lost its announced name, or was replaced by a different one. The suite reports the flow works while the interface changed.

open as a page

In a machine-drafted end-to-end case, how do you tell an outcome assertion from one that restates the steps just performed?

level: middleimportance: must knowfreq 62%

basics

~20 s

An outcome assertion checks a fact the system decided or stored on its own. A restating assertion only re-reads what the case supplied or already watched happen. Ask whether it would still pass with the behaviour removed.

open as a page

Which decisions stay with a person when a machine drafts test cases from a running application?

level: middleimportance: must knowfreq 62%

basics

~20 s

Two: what correct behaviour actually is, and which behaviours carry real consequence. A tool drafting from a running system can only describe what it observes, so it will happily record a defect as the expected result.

open as a page

When does delegating a test's pass/fail verdict to a model beat an explicit assertion, and when is it strictly worse?

level: middleimportance: must knowfreq 46%

basics

~20 s

Delegate the verdict only when the acceptable result is a family you cannot enumerate: free-form text, wording shown to a user, perceived quality. Where an exact value or a structural rule exists, an explicit comparison is strictly better.

open as a page

When a drafting model writes most of a suite's cases, where does the team's effort move instead of vanishing?

level: middleimportance: must knowfreq 55%

basics

~20 s

Authoring effort falls; review, triage and repair rise. Cheap drafting increases how many cases a team must read, explain and keep truthful, so the cost moves downstream and is charged against a bigger standing suite.

open as a page

Why hand a drafting model the acceptance criteria rather than the implementation when drafting a test?

level: middleimportance: must knowfreq 58%

basics

~20 s

Acceptance criteria describe the behaviour a test must prove; the implementation describes what the code happens to do today. A model given the code writes assertions that mirror it, so the case stays green even when the behaviour is wrong.

open as a page

When reviewing a machine-drafted service case, how do you judge whether its setup and assumed data are honest?

level: juniorimportance: should knowfreq 40%

basics

~20 s

Trace every fact the case depends on back to a line that created it. Anything it reads but never creates is a borrowed assumption that will fail wherever that data is absent, for reasons unrelated to the behaviour.

open as a page

Why can a check that asks a model to judge a page's error text pass one run and fail the next on unchanged output?

level: juniorimportance: should knowfreq 38%

basics

~20 s

A judged step generates its answer by sampling, so identical input can produce different verdicts. Borderline outputs flip most, wording changes in the standard move verdicts, and the model doing the judging can be updated without a commit anywhere.

open as a page

What must be stripped from a production record before it goes into a drafting prompt, and what must survive?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Strip whatever identifies a person or grants access: names, contact details, account numbers, tokens and keys. Keep the shape - field set, formats, lengths, encodings, boundary values - because the shape is why a real record drafts better cases.

open as a page

A drafted batch of service-level cases asserts the exact responses captured when the drafts were made. What can that suite not detect?

level: middleimportance: should knowfreq 47%

basics

~20 s

Expectations copied from observed output encode whatever the system does today, defects included. Such a suite detects change rather than wrongness: it stays green over a defect that was present at capture, and reddens on every intended change.

open as a page

A run reported every case green after repairing forty locators - how should its output surface those repairs?

level: middleimportance: should knowfreq 38%

basics

~20 s

Give a repaired case its own outcome, distinct from a clean pass; put the substituted anchors on the run summary a reviewer actually reads; and set a volume above which the run stops being green at all.

open as a page

What house conventions should you pin into the context you give a drafting model before it writes a case?

level: middleimportance: should knowfreq 42%

basics

~20 s

Pin whatever a reviewer would otherwise repeat every time: case naming and placement, the domain vocabulary, the level of abstraction, what a failure must record, and the constructs the team bans. Attach a short rulebook plus two exemplar cases.

open as a page

How does a self-healing locator score candidate substitutes when its original element is gone?

level: middleimportance: should knowfreq 42%

basics

~20 s

A self-healing locator ranks the elements now on screen against a description of the vanished element captured on the last clean run - attributes, visible text, kind of control, neighbouring labels, position - and substitutes the highest-scoring candidate.

open as a page

A machine-drafted end-to-end case arrives with a green run; why is that weak evidence, and what do you read first?

level: seniorimportance: should knowfreq 52%

basics

~10 s

Green proves only that the case executes here today; a case that can never fail is green too. Read the claim first, then the assertions, the setup, and whether the pack already proves it.

open as a page

Why should a model-judged verdict never be the only thing that fails a build, and what deterministic check sits behind it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A build gate must be reproducible, reviewable and stable in its wording; a judged verdict is none of these. Let explicit assertions on statuses, amounts, identifiers, required fields and forbidden content fail the build, and report judged findings as signals.

open as a page

Which steps must a self-repairing suite refuse to heal and fail outright on, whatever the substitute's confidence?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Steps where the control's identity is part of what you are testing: the exact action a case exists to prove is reachable, irreversible or money-spending operations, and any check of a control's announced name, enabled state or permission gating.

open as a page

What confidence threshold lets a self-healing locator repair silently, and who owns that number?

level: seniorimportance: should knowfreq 38%

basics

~20 s

No universal number exists. Calibrate the threshold on a labelled sample of the mechanism's own past repairs, trading wrong substitutions against needless failures. It is a team-owned setting with evidence behind it, not a shipped default left untouched.

open as a page

An approval gate admitting cases into a regression suite now receives ten times as many machine-drafted candidates. How do you design it so refusal stays real?

level: principalimportance: should knowfreq 46%

basics

~20 s

Bound intake to review capacity, default to refusal when a candidate is undecided, and make rejection cheap and reasoned. Tier review depth by what a case can stop, then watch the refusal rate: a gate that never refuses is decoration.

open as a page

Leadership asks whether machine-assisted case authoring paid for itself. What must that before-and-after comparison measure to be honest?

level: principalimportance: should knowfreq 40%

basics

~20 s

An honest comparison prices a case's whole lifetime on both sides - drafting, review, triage, repair - over a window long enough for upkeep to appear, states the expected value in advance, and names the confounders that could explain the result.

open as a page

In a suite grown by bulk drafting, how do you find cases that duplicate an existing case's intent under a new name?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Compare what cases do rather than what they are called: the units each case executes, the outcomes it checks, and the requirement it claims to cover. Cases whose signatures coincide share one intent whatever their names say.

open as a page

How do you catch a machine-drafted end-to-end case that duplicates existing coverage under a new name?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Reduce the draft to a signature — precondition, trigger, and the fact its assertions prove — then search the pack for that fact rather than the name. Matching signatures mean a duplicate however different the wording is.

open as a page

Why should accepting a self-repaired element locator, or promoting a drafted case into the regression set, be recorded rather than automatic?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Both acts change what the suite claims, and a silent default means nobody owns the claim. A recorded acceptance gives the change an author, a reason and a reversible history; a default gives it none of those.

open as a page

Before a suite trusts a model-judged check, how do you calibrate it against human-labelled examples?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Run the judged check over outputs people have already labelled against the written standard, then read every disagreement. Trust it only once it catches known past defects and the subtly-wrong near-misses, not just obviously good and obviously broken examples.

open as a page

Which symptoms show a machine-grown test suite has outgrown the upkeep capacity of the team that owns it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The suite stops being read. Failures are answered by re-running rather than diagnosing, nobody can say what a case protects, review turns into sampling, and repairs are deferred. All of it happens while the suite still passes.

open as a page

When a control's accessible name disappears and a self-repairing suite substitutes a positional locator, what defect goes unreported?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

An accessibility regression. The control no longer exposes a name assistive technology can announce, so some users cannot reach it - and the repair that routed around the missing name turned the one automated signal for it into a pass.

open as a page

What has to be recorded about a drafting run for a generated case to be explainable months later?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Record the inputs, not the conversation: the request text, every attached artefact with its revision, which build and settings produced the draft, and the untouched output before human edits. Link the merged case to that record.

open as a page

A suite's locator healing rate has risen from 2% to 30% of steps this quarter - what does that tell you?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

A climbing healing rate says the suite's locators are drifting from the product faster than anyone repairs them at source. Repair has become the maintenance strategy: at thirty percent, one located step in three operates an element nobody chose.

open as a page