Which symptoms show a machine-grown test suite has outgrown the upkeep capacity of the team that owns it?
answer
- A human limit, not a runtime one
- A green suite hides it completely
- Failures answered by re-running, not reading
- Cases per competent repairer, and its direction
basics
~20 sThe suite stops being read. Failures are answered by re-running rather than diagnosing, nobody can say what a case protects, review turns into sampling, and repairs are deferred. All of it happens while the suite still passes.
solid answer
~50 sCapacity here is a human limit, not a runtime one, so the first symptoms are behavioural rather than numeric. Cases appear whose purpose nobody can state, because the person who accepted one never held its intent. Failures are answered by running the case again instead of reading it, since reading an unfamiliar case is the expensive act. Review quietly becomes sampling: nobody decides to sample, the queue simply outgrows the day. Repairs after a legitimate product change get batched and deferred, and eventually a product change is slowed because bringing the suite along is costly. The compact diagnostic is a ratio: cases owned per person who could competently repair one, and which way that ratio has moved. The suite can be entirely green while all of this is true, which is exactly why teams admit it late.
code
yaml · 10 linescase:
id: declined-payment-keeps-cart
protects: "a declined payment leaves the basket intact"
drafted_by: model
accepted_by: person
owner: payments-team
accepted_on: 2026-04-11
last_repaired: 2026-04-11
failures_last_90d: 7
failures_actually_diagnosed: 1go deeper
Be ready to recall that someone has to look after each automated case after it exists. If a case fails and nobody knows what it was checking, that is a cost the team already carries, whatever the run says.
Explain the mechanism: a team can absorb only so many cases because upkeep needs knowledge of what each one protects, and cheap drafting removes the effort that used to keep suite size and team size in step.
Demonstrate that you can spot it early. Name the behavioural symptoms, say why a passing run is no evidence either way, and give the ratio you would actually look at rather than a general worry about suite size.
Own the throttle. Decide what bounds intake now that authoring effort no longer does, and be able to defend to a sponsor why the team is declining cases it could produce in seconds.
A suite's real limit is not how long it takes to run. It is the number of cases the people who own it can read, explain and repair, and that number is set by human attention rather than by machinery. Machine drafting removes the thing that used to keep suite size and team size roughly in step, namely the effort of writing a case, so a suite can pass its owners' limit without anything visibly breaking. That is what makes the failure mode hard to catch: the suite is green, the run is quick, and the team is already past the point where it can honestly stand behind everything it owns. ## Why the limit is a human number The capacity that matters is how many cases the team can competently repair, and competence here means knowing what a case was meant to protect. That knowledge is built by writing cases, reviewing them attentively, and repairing them after real product changes. None of it arrives with a drafted case. So a team's absorbable count is roughly its people multiplied by the cases each can hold in working memory, and it moves slowly. The suite, once drafting is cheap, does not move slowly at all. ## The symptoms, roughly in the order they appear 1. **Purpose amnesia.** Someone asks what a case protects and the only answer available is to read the whole body of it. This appears first, at acceptance, before any failure has happened. 2. **The re-run reflex.** A failure is answered by running the case again rather than by reading it. Re-running is cheap and reading an unfamiliar case is not, so the team is behaving rationally under a constraint it has not named. 3. **Review becomes sampling.** Nobody decides to sample. The queue simply grows past what the day holds, and a fraction gets a real read while the rest get a glance. 4. **Repair debt after product change.** A legitimate product change breaks a batch of cases; the batch is fixed in a rush, or worse, the change is delayed because the suite is expensive to bring along. 5. **Ownership diffuses.** Cases stop having a person who would notice their absence. Nobody objects to a case, and nobody defends one either. 6. **Growth decouples from the product.** The suite grows in quarters when the product barely changed, which means the count is now driven by drafting supply rather than by anything the product did. ## Reading each symptom honestly | Symptom | The comfortable reading | What it actually indicates | | --- | --- | --- | | Nobody can state what a case protects | the case is self-documenting | intent was never captured at acceptance | | Failures answered by re-running | the failure was transient | reading the case costs more than the team can spend | | Review queue grows steadily | the team is busy this sprint | intake exceeds the absorbable rate, permanently | | Repairs batched after a change | efficient batching | upkeep is being deferred rather than paid | | Suite grew while the product did not | better coverage | supply, not risk, is choosing what gets covered | ## Why teams admit it late Three things delay the admission. The suite is green, and green is emotionally read as healthy even though it only ever reports on the product, not on whether anyone can maintain the cases. Volume looks like progress, and a chart of case count climbing is easy to show and pleasant to present. And the cost is paid in interruptions that are never booked anywhere, so no budget line moves until a release is late. The clean diagnostic question is a ratio rather than a number: how many cases exist per person who could competently repair one, and which direction has that ratio moved over the last two quarters? A team of four owning 300 cases they wrote is in a different position from the same team owning 1,200 cases a drafter produced, even though both suites pass. ## What changes once the limit is accepted Accepting the limit is mostly about putting a deliberate throttle where authoring effort used to be one. Intake is bounded by what the team can absorb rather than by what the drafter can produce. Each accepted case carries a one-sentence statement of the behaviour it protects and a named owner, written at acceptance because that is the only moment the intent exists cheaply. Failure work is treated as work, with someone accountable for reading rather than re-running. And growth of the suite is expected to track product change, so a quarter where the count climbed and the product did not becomes a question rather than an achievement. None of this is a reason to refuse machine drafting. It is the recognition that drafting speed and upkeep capacity are two different budgets, and that only one of them changed.
- Why can a suite be entirely green and still be past the team's capacity?Green reports on the product under test, not on the suite's maintainability. A case nobody understands, nobody owns and nobody could repair passes exactly as easily as a well-understood one. Capacity is about whether a person can act on a failure when it eventually arrives, and a passing run says nothing about that.
- Which of these symptoms would you expect to appear first, and why?Purpose amnesia, because it is created at acceptance rather than at failure. The moment a case is accepted without anyone recording what behaviour it protects, the debt exists; every other symptom needs a failure or a product change to surface it, which may be weeks away.
- Does an unowned drafted case differ from an unowned hand-written one?Yes, in what can be recovered. A hand-written case has a person who once held its intent and can usually reconstruct it. A drafted case that was accepted without a recorded purpose has no such person, so the intent has to be inferred from the assertions themselves, which is circular.
A warehouse that keeps accepting deliveries looks fine until someone needs one specific item and nobody knows which shelf it is on.
saying these in an interview costs you the question
- Says a green suite proves the team is keeping up
- Measures capacity by pipeline runtime instead of people
- Treats re-running a failure as diagnosing it
- Assumes drafted cases need no owner
- Reads a rising case count as rising protection