What do you weigh before a batch of generated unit tests joins the build the whole team waits on?
answer
- The build is a shared resource
- Price used to do the filtering for you
- Somebody must be able to state the claim
- Cheap to produce, unchanged to own
basics
~20 sWeigh what each test would report that nothing else would against what everybody pays for it on every run: minutes of build time, failures that are not about the product, and an obligation to maintain something nobody can explain.
solid answer
~50 sProducing a test became nearly free; owning one did not. Admission used to decide itself, because nobody paid to write a test they had no use for, and that filter is gone. So decide it out loud, on four things: **marginal detection** - what failure would this report that nothing else would; **stability** - whether it can go red for reasons other than the behaviour, such as the clock, ordering or shared state the drafting model was never shown; **run time**, paid by everyone on every run; and whether a person can say in one sentence what it protects, because that sentence is what the next failure gets triaged against. The last one is the part a person owns: you can generate the body of a test, but the claim it makes is about the product, and somebody has to stand behind it.
go deeper
Remember that anything you add to the shared build is paid for by everyone on every run, so a test needs a reason to be there beyond having been easy to produce.
Be able to weigh a test on what it would report that nothing else would, and to spot the shapes that go red for reasons other than the product.
Show how you review a batch as a batch rather than test by test, and what you do with one that fails intermittently and has no statable claim.
Own the admission bar and argue it on cost: what the build is worth to the team, what the filter used to be, and why a rule about what a test carries beats a rule about where it came from.
## The decision that used to make itself A shared build is a resource the whole team draws on. Every test in it is paid for repeatedly - in minutes on every run, in attention on every failure, and in maintenance for as long as the code lives. Against that, a test earns its place by reporting a wrong behaviour that nothing else would report. That trade used to be made implicitly, by price: writing a test cost enough that nobody wrote one without a use for it, so a suite stayed roughly proportionate to the thinking behind it. Generation removes the price on one side and leaves the other where it was, so the filter that used to be free now has to be applied deliberately - and a team that does not acquires a large suite it never decided to have. ## What to weigh | what you weigh | the question | what a bad answer looks like | | --- | --- | --- | | **Marginal detection** | what wrong behaviour would this report that no existing test would? | "it covers the same path from another angle" | | **Stability** | can it go red for reasons that are not about the product? | it reads the clock, depends on ordering, or assumes state a shared run does not guarantee | | **Run cost** | what does it add to every run, for everyone? | a slow setup duplicated across a batch of near-identical tests | | **A statable claim** | can a person say in one sentence what it protects? | the honest answer is "it was in the batch" | The stability row has a note specific to drafted tests. A model writes against what it was shown, and your build's constraints - parallel runs, a fixture that is not isolated, a clock that is not yours - are rarely in that material unless somebody put them there. So a batch is likelier to carry environment-dependent shapes than tests written by somebody who has watched that build go red at three in the morning. That is about what the model could see, not what it is capable of. ## The part a person has to supply You can generate the body of a test, and you can get a generator to write a sentence about why it exists - but that sentence will be derived from the code, which is the thing in question. The reason a test deserves to exist is a claim about what the product must do, and somebody has to stand behind it. So the thing a human owns at admission is **the claim**, expressed in the one place that survives: the test's name, and an expectation derived from the rule rather than read off the implementation. This is cheap to demand and it changes what happens later. A test whose name states a rule can be triaged by anybody who reads it. A test named after the function it calls has to be triaged by inferring the intent from its body, and the person best placed to do that is whoever generated it - if they are still on the team, and if they remember. ## What an unstable generated test actually costs Not one red build. The cost lands in three places, and the third is the expensive one: - **The run everybody waited on**, restarted. - **The investigation**, longer than for a hand-written test because there is no author to ask what it was protecting. - **The team's willingness to believe a red build.** Once re-running is the normal response to red, the suite has stopped being a signal and become a toll, and every test in it - including the good ones - is worth less. And the repair is harder in one specific way. When nobody can say what a test protects, the honest moves are to establish the claim, to repair the instability where its cause is plain to see, or to drop the test. Making it pass by re-running it is the option that looks cheapest, and it is a decision to stop reading the signal - it just does not feel like one at the time. ## Reviewing a batch as a batch A batch is not a pile of individual decisions, and reviewing it test by test is both slow and misleading. 1. **Ask what the batch is for** before reading any of it. A batch with no stated purpose rarely acquires one during review. 2. **Read a sample closely** and judge the batch's shape from it - where it clustered, what it asserts, what it never touches. 3. **Judge marginal detection at the batch level**: which behaviours would now be caught that would previously have shipped? 4. **Admit, revise or decline the batch as a unit.** Declining one is a normal outcome, and cheap now that producing it was cheap. ## What is worth standardising A lead's version of this question is not *which tool* but *what bar*. A rule about origin - generated tests banned, or generated tests fine - aims at the wrong thing and is hard to enforce; a rule about what a test must carry applies to every test in the repository and catches exactly the failure mode cheap production creates. The bar that works is short: **a test with no statable claim does not enter the shared build.** How a team reviews AI-assisted changes in general is a broader subject with its own answer; this is the one part of it the build itself can hold.
- Is declining a batch of generated tests wasteful, given the work already exists?The work cost almost nothing, which is the point: the sunk cost is small and the ongoing cost is not. Judge the batch on what it would report and what it would cost every run from now on. Saying no to a batch and asking for a smaller one aimed at the risky paths is usually the cheaper move on both counts.
- How would you tell whether a team's generated tests are paying for themselves?Look for failures that found something. Over a few months, ask which real defects were first reported by a test from these batches, and set that against the build time they add and the investigations they caused. That comparison is about your own suite, so it is answerable, unlike a general claim about whether generated tests work.
saying these in an interview costs you the question
- Tests are always worth having, so let the whole batch in.
- A generated test costs nothing, since nobody had to write it.
- If it goes red intermittently, re-run it and move on.
- Ban generated tests and the problem goes away.
- Build time is an infrastructure problem, not a review decision.