Leadership asks whether machine-assisted case authoring paid for itself. What must that before-and-after comparison measure to be honest?
answer
- Define paid before reading any numbers
- Lifetime hours, never authoring hours alone
- Window must outlast one product change
- Compare like areas, not calendar halves
- Commit in advance to what failure looks like
basics
~20 sAn honest comparison prices a case's whole lifetime on both sides - drafting, review, triage, repair - over a window long enough for upkeep to appear, states the expected value in advance, and names the confounders that could explain the result.
solid answer
~50 sFix what "paid" means before looking at any data: a cost side, a value side, a window, and a decision rule you would honour either way. The cost side has to be lifetime hours across drafting, review, triage and repair, because the change moves cost downstream and a study that stops at authoring always reports a win. The value side must be chosen up front, whether that is behaviours now covered that previously had none, regressions caught before a release, or time to make a change confidently. The window must span at least one real product change, since upkeep does not exist until the product moves. Then name the confounders, team maturation, product churn, cadence changes and which area was picked, and control what you can by comparing two comparable product areas over the same period rather than one team across calendar halves.
code
yaml · 17 linesevaluation:
asks: "did machine drafting lower cost per protected behaviour?"
window: "two release cycles, so upkeep has appeared"
unit: "two comparable product areas, same period"
cost_hours:
- drafting
- review
- triage
- repair
value:
- behaviours_covered_first_time
- regressions_caught_before_release
confounders:
- team_maturation
- product_churn
- area_selection
decision_rule: "no gain in cost per protected behaviour -> re-scope"go deeper
Recall that a saving in one place can be a cost in another. If drafting a case got faster but someone still has to read, run and fix it, the saving is only real once those hours are counted too.
Be able to list what belongs on the cost side of such a comparison and why authoring alone is misleading. Explaining that upkeep only appears after the product changes, so a short window flatters the result, is the mechanics expected here.
Show you can run the study. Pick a window that has survived a real product change, choose comparable areas rather than calendar halves, and be specific about the confounder you could not control and how it limits the conclusion.
Own the judgement about evidence. Commit to the cost side, the value side, the window and the stopping rule before data exists, defend the choice to a sponsor who wanted a simpler number, and be willing to report a split or inconclusive verdict.
"Did it pay?" is not yet a question anyone can answer. Before any data is worth collecting it has to be turned into four commitments made in advance: what counts as cost, what counts as value, over what window, and what result would make you stop. Making all four before looking is what separates an evaluation from a justification, and the order matters, because every one of these choices can be made after the fact in a way that produces whichever answer the person making it already prefers. ## Turning the question into something decidable The first commitment is the **cost side**, and the honest version is lifetime hours, not keyboard hours. Drafting, review, triage and repair are all costs of the same suite, and only the first of them falls when a model does the drafting. An evaluation that stops at the term that fell is guaranteed to report a win, which is why it is the most common shape and the least informative one. The second is the **value side**, stated before the numbers arrive. Reasonable candidates include the count of behaviours that now have a case and previously had none, the regressions the suite caught before they reached a release, and the elapsed time between a developer wanting to make a change and being confident it is safe. Any of these can be defended. What cannot be defended is choosing among them after seeing which one happens to look good. The third is the **window**. Upkeep only exists after the product has moved, so a window shorter than a full change cycle prices the honeymoon and nothing else. Two release cycles is a defensible floor; the test is whether the cases in the study have already survived at least one real product change. The fourth is a **decision rule**: the result that would make you narrow, re-scope or stop. Without one, an inconclusive study is always read as a mild endorsement. ## Two shapes of the same comparison | | The naive comparison | The honest comparison | | --- | --- | --- | | Cost counted | hours saved authoring | drafting plus review plus triage plus repair | | Value counted | number of cases added | behaviours protected that were not before | | Window | the quarter after adoption | at least one full product change cycle | | Unit compared | calendar halves for one team | comparable product areas over one period | | Result if unclear | reported as a success | reported as unclear, with the rule applied | The unit line is the one that does most of the work. A before-and-after over calendar time attributes to the practice everything else that changed in those months, and something always changed: the team learned the domain, the product's release cadence moved, two people joined, an area stabilised. Comparing two comparable areas of the product over the same period, one using machine drafting and one not, is imperfect but it holds most of that constant. ## The confounders you are obliged to name - **Team maturation.** Teams get better at their own product; some of any improvement is simply the calendar. - **Product churn.** A quiet period makes upkeep look cheap and a redesign makes it look ruinous, independent of how the cases were written. - **Selection of the area.** The area chosen for machine drafting is often the one that was easiest to cover, which flatters the result. - **Volume as a disguised cost.** More cases is an input, not an outcome. Reporting it as a benefit double counts, since every added case also raises the recurring bill. - **The definition of a caught regression.** Whether a failure counted as a real catch is a judgement, and it must be made by the same rule on both sides of the comparison. ## Reporting it honestly The strongest report states the four commitments, gives the numbers, names the confounders it could not remove, and then states its confidence plainly. An outcome worth being prepared for is a **split verdict**: cost per case roughly flat, coverage of previously untested behaviour clearly up. That is a real and defensible result, and it usually means the practice paid in reach rather than in efficiency, which changes what the organisation should do next: keep drafting, but move the throttle from authoring speed to review capacity. An inconclusive result is also a result. Saying so costs credibility once and buys it back permanently, whereas a study built to confirm a decision already made is discovered eventually and taints every measurement the team publishes afterwards. The judgement a lead genuinely owns here is what evidence they will accept, decided while it is still possible to be wrong about it, and whether the organisation is prepared to act on a verdict it did not want.
- What would make you conclude the trade did not pay even though authoring hours fell sharply?If the suite grew substantially while the set of behaviours actually protected stayed roughly where it was, and review and repair hours absorbed the whole authoring saving. That is volume bought at flat spend with no additional protection, and the extra cases will keep charging rent every release.
- Why is a calendar before-and-after weaker than comparing two comparable product areas?Calendar comparison attributes every change in the period to the practice, and teams, products and cadences all change over months. Two comparable areas in the same period share the calendar, the team's maturity and the product's churn, so most of what would otherwise be confounded is held constant.
- How do you report an inconclusive result without it being read as a failure?State the four commitments made in advance, give the numbers, and say which confounder you could not remove and what it would take to remove it. An inconclusive study that names its own limits is evidence of method; the alternative, a confident verdict the data does not support, is discovered later and costs far more.
saying these in an interview costs you the question
- Counts hours saved at the keyboard and stops there
- Uses a window shorter than one product change cycle
- Chooses the value measure after seeing the data
- Ignores that the team also matured over the period
- Reports the number of cases added as the benefit
- Has no result that would count as a failure