skip to content

Test Generation

Prompting for tests that genuinely cover edge cases, and spotting the tautological ones that just assert whatever the code already does. Interviewers ask because generated tests can raise the coverage number while lowering real confidence.

part ofAI-assisted developmentoverview, primer and where to startread it →
on this pageshow

questions

5

A generated unit test asserts exactly what the function already returns. What is it unable to catch?

level: juniorimportance: must knowfreq 60%

answer

  1. Nothing in it disagrees with the code
  2. It is green the day it is written
  3. Ask what wrong answer it would reject
  4. The expectation reuses the function's own calculation

basics

~20 s

It cannot catch the function being wrong. An expectation taken from the implementation agrees with that implementation by construction, so the test reports that behaviour has changed rather than whether the behaviour was ever right.

solid answer

~50 s

Such a test is **tautological**: its expectation and the code it judges have the same origin, so the comparison is an identity dressed as a check. Three shapes turn up. The expectation is produced by calling the subject itself, or read back off a stub the test configured - that one cannot fail for any reason to do with correctness. The test recomputes the function's own calculation, often through the same helper - that one moves whenever the code moves, so a rule that was wrong from the start stays invisible. Or the expectation is a literal captured from a run - that one goes red on any change that moves the value, wanted or not, and still says nothing about whether the captured value was right. The reading test is one question: what would have to be wrong for this to go red?

code

pseudocode · 18 lines
pseudocode
# the rule, as the leave policy states it:
#   1.5 days accrue for each COMPLETED month of service,
#   counted up to the last day of the policy year and no further.

# generated after being shown accrue()'s body
test "accrual for a mid-year joiner":
    joined, as_of = MAR_01, DEC_31
    expected = months_between(joined, as_of) * ACCRUAL_RATE   # the body's own calculation, restated
    assert accrue(joined, as_of) == expected
    # green whatever months_between() counts - started months, completed months,
    # or months past the year end - because both sides use it

# derived from the policy instead of from the code
test "a joiner on 1 March has ten completed months by 31 December":
    assert accrue(MAR_01, DEC_31) == 15.0     # 10 completed months x 1.5

test "service that has not completed a month accrues nothing":
    assert accrue(DEC_15, DEC_31) == 0.0      # red if part-months are counted

go deeper

for a junior

Know the shape and the name: a test whose expected value came out of the code it is testing. Be able to say why it is green on the day it is written and what that therefore proves.

for a middle

Explain the three shapes - self-comparison, restated calculation, recorded literal - and why the one that goes red most often is still reporting change rather than correctness.

for a senior

Show how you find these in review without running anything: name a wrong result, point at the line that would reject it, and say what you would put in its place.

for a principal

Own the difference between pinning behaviour on purpose and mistaking a pin for a check, and make sure the suite's naming lets any reader tell which one a test is.

## What you are holding A unit test that arrives from a generator looks exactly like one a colleague wrote. It sets up an input, calls one function, compares a result, and it is green. The question worth asking of it is not *"is this assertion correct?"* - it usually is - but **"what would have to be wrong for this to go red?"** When the only honest answer is *"the function would have to stop doing what it does today"*, the test is **tautological**: its expectation and the thing it judges have the same origin, so the comparison is an identity wearing the costume of a check. (*Why* drafts drift this way is a question about what the model was handed up front, and that is a different subject; this is about reading the artefact on the page.) ## Three shapes, and what each can report | shape | where the expectation comes from | what turns it red | what it cannot report | | --- | --- | --- | --- | | **Self-comparison** | a call to the subject itself, or a value read back off a stub the test configured | non-determinism, and little else | anything at all about the rule | | **Restated expression** | the test recomputes the function's own calculation, often through the same helper | one side changing without the other | a shared rule that was wrong from the start | | **Recorded literal** | a value captured from a run on the day the test was written | a change to the value it recorded, intended or not | whether the captured value was ever right | The three are not equally bad, and saying which one you are looking at is most of the review comment. - The **self-comparison** is the pure form and the easiest to miss when the two sides are separated by a fixture or a double: the test configures a stub to return a value, then asserts that the value comes out. It reports nothing except that the code is deterministic. - The **restated expression** is the common one and the most durable, because it looks like careful work. Both sides move together, so it stays green through a change that alters what people actually receive - the event a test existed to prevent. - The **recorded literal** fails often, and that is its problem rather than its defence. It goes red on legitimate changes as readily as on regressions, trains the team to update it without thinking, and still never says whether the number it holds was correct. ## Why it survives a pull request Nothing about it is alarming. It is green; it reads like the tests around it; the diff shows tests being **added**, which reviewers rarely argue with; and the failure it should have produced lies in the future, on a change nobody has made yet. When it does finally go red, the cheapest repair - paste in whatever the function now returns - quietly converts it back into a pin. That is how a suite fills up with these without anyone deciding to build one. ## The reading test that finds them 1. **Name a wrong answer.** Not "a bug" - a specific wrong result the function could plausibly return: a boundary counted one day early, a part-month rounded up, a period running backwards that yields a positive figure. 2. **Point at the line that would reject it.** Read the assertions and find the one that goes red on that wrong answer. 3. **If there is no such line, you have found one.** The test may still be worth keeping, but it is not evidence about correctness and should not be counted as though it were. 4. **Then check the expectation's own arithmetic** - its unit, its scale, which side of the boundary it falls on. An expectation can be independent of the code and still wrong, which is a separate failure worth its own minute. ## The same artefact, honestly used Pinning current behaviour on purpose is a real technique. Before restructuring code whose rules nobody wrote down, a test that says *whatever this did yesterday, it must still do* is exactly what you want, and generating a pile of them is a good use of a generator. The artefact is identical to the defect. What differs is that the team knows the expectation's authority is the code, and the test's **name** says so - `pins_current_behaviour_for_a_mid_year_joiner` rather than `mid_year_joiner_accrues_correctly`. Treating every pin as a defect is as wrong as trusting every pin as a check; what a suite cannot survive is nobody being able to tell which is which. ## What to do with one you find Replace the expectation on the one or two cases you can state as a rule - service that has not completed a month accrues nothing, a period that ends before it begins is refused - and name each test after the rule it encodes. Keep the rest as pins where they are useful, renamed to say what they are. The edit is small and it pays forward: the next person reading the suite inherits a rule instead of a number, and the next draft written against this file starts from a better example.

  • Is a test that pins the function's current output ever the right thing to write?
    Yes - deliberately, when you are about to restructure code whose rules nobody wrote down and you want any change in behaviour reported. The artefact is the same one; what makes it honest is that the team knows the expectation's authority is the code, and the test's name says so rather than implying the value was checked against a rule.
  • A generated test goes red after an intended behaviour change. What is the wrong repair?
    Pasting the new output in as the expectation without deciding whether the new output is correct. That turns a failure into a rubber stamp and leaves the next reader a value nobody has ever checked. Decide what the rule now says, set the expectation from the rule, and rename the test after it if the rule moved.

It is a spell-checker built out of the document it is checking: every word comes back confirmed, including the misspelt ones.

saying these in an interview costs you the question

  • If it passes and the coverage number moved, the test is earning its place.
  • The model read the function, so its expected values must be right.
  • A test is fine as long as it can fail for some reason.
  • When it goes red, update the expected value to whatever the code now returns.
  • Generated tests are tautological, so a generated test is never worth keeping.
open as a page

How do you check a generated test's expected values when the rule they encode is written down nowhere?

level: middleimportance: must knowfreq 55%

basics

~20 s

Split it in two. Settle what you can alone - unit, scale, sign, which side of the boundary the value falls on - then take what is left to whoever owns the rule, as one concrete case with a yes-or-no answer.

open as a page

You asked for unit tests on a leave-accrual calculator and got only ordinary dates. What should the request have named?

level: middleimportance: should knowfreq 48%

basics

~20 s

Name the situations, not the quantity. A model drafting from the code produces cases the code already implies. The awkward ones live in the rule, so list them yourself: the period that ends before it starts, the joiner on the boundary.

open as a page

Your coverage number rose after a batch of generated unit tests. Why might the team be no safer?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The number counts code that ran. Generation made running code cheap and left the hard half - deciding what the right answer is - exactly as expensive, so a figure that stood in for effort no longer does.

open as a page

What do you weigh before a batch of generated unit tests joins the build the whole team waits on?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Weigh what each test would report that nothing else would against what everybody pays for it on every run: minutes of build time, failures that are not about the product, and an obligation to maintain something nobody can explain.

open as a page