Six months in, a manager asks whether AI coding assistants made the team faster — what can you conclude?
answer
- Compared with what, exactly?
- Six months changed more than the tool
- The claim's boundary stops at authoring
- Report where effort moved, not whether it shrank
- Refuse the conclusion in both directions
basics
~10 sThroughput alone settles nothing: over six months the people, the codebase and the work all changed alongside the tool. Report where effort moved and which narrow claims you are willing to defend.
solid answer
~40 sStart by refusing the bare comparison, and say why. Six months ago the team, the codebase and the backlog were all different, so a change in merged volume has several candidate causes and the assistant is only one of them — say which you can rule out and which you cannot. Then report what you can defend: where effort *moved*, into reviewing, into repairs after merge, into the defects that surface weeks later, rather than whether the total shrank. Refuse in both directions, too. Flat throughput does not show the tool failed, because a real gain can be absorbed into quality, into a harder backlog, or into the same deadline met with less overtime. A narrow claim about a named kind of work survives being quoted back at you; "we are faster" does not.
go deeper
Ask what a speed-up claim is being compared with before you repeat it. Writing a change is one step of getting it live, so time saved there is not automatically time saved overall.
Explain what else moved over the same period, and name the steps a speed-up claim usually leaves outside its boundary: reviewing, repairs after merge, and the defects that arrive later.
Show that you refuse the conclusion in both directions and still deliver something useful — where effort moved, what you can rule out, and the narrow claims you are prepared to defend by name.
Own what the organisation will and will not measure, and what measuring costs. Choosing the unit and stating in advance what result would stop a rollout cannot be added retroactively without being chosen by someone who already knows the answer they want.
## What the manager is actually asking A request for a number is usually a request for a decision. *Did we get faster?* sits on top of *should we keep paying for this, extend it, or stop?* — and the second question is answerable even when the first is not. So answer the decision, and be explicit that you are declining the number and why. A confident figure offered here is the least durable thing you can say: it will be repeated for years by people who were not in the room, and it cannot survive the first person who asks how it was measured. The answer an interviewer is listening for is not pessimism. It is whether you can tell the difference between **what happened** and **what the tool caused**, and still leave the manager with something to act on. ## Why there is no clean before-and-after Six months is long enough for everything else to move as well. Before attributing anything to the assistant, write down what else changed: - **The people.** Joiners, leavers, and the ordinary fact that everyone is six months more familiar with the system than they were. - **The codebase.** The parts under work now are not the parts that were under work then, and no system is uniformly hard. - **The work.** A quarter of new features and a quarter of migration produce different numbers with no change in skill or tooling at all. - **The process.** A change to the pipeline, to who reviews what, to the team's shape, or a deadline that concentrated everyone's attention. - **The reporting.** Once a team believes it is faster, its estimates and its self-reports move with the belief, and they are the inputs to most of the numbers anyone has. None of these can simply be subtracted afterwards from a count of merged changes. **A measurement with no counterfactual is a description, not a comparison**: it tells you what happened, not what would have happened otherwise, and every attribution you make on top of it is a guess about that difference. Saying so is not a dodge. It is the part of the answer that makes the rest of it credible. ## The claim's boundary is narrower than the cycle The second problem is narrower and more common than confounding: a saving observed on one step, reported as a saving on the whole. Most speed-up claims are drawn with a boundary around authoring, because that is the step people can feel. | what the claim usually counts | what sits outside its boundary | |---|---| | time from starting a change to opening it | the time somebody else spends reading it | | changes merged in a period | changes repaired or reverted afterwards | | the author's hours | the reviewer's hours, and the next reader's | | the version that passed the checks | the defects that surface weeks later | | work in familiar areas | the cost of getting into unfamiliar ones | The rule underneath the table is simple: **a claim is only as honest as the widest boundary you drew**, and a step that gets cheaper tends to push work into the steps around it. So the first thing to establish is not whether total effort fell but where it went — that is observable, it is defensible, and a manager can act on it. How a team should handle the load that lands downstream is its own question. ## What you can honestly report - **Where effort moved**, with the evidence you actually have and its limits stated. - **What you can rule out and what you cannot.** Naming your own confounders is what distinguishes an evaluation from a justification. - **Narrow claims instead of broad ones.** *Drafting the repetitive parts of this kind of change is quicker, and those changes take longer to review* can be defended. *We are faster* cannot. - **The decision you recommend, and what would change it.** Say what you would expect to see if you were wrong, and say it before anyone tests you on it. If the manager needs a comparison rather than a description, that is a different exercise — one that has to be set up in advance, with the unit and the measure chosen before anybody has an opinion about the answer. Six months in, that option is gone for the period already behind you, and pretending otherwise is how a bad number gets into a slide. ## The two refusals Refuse in both directions, or you are not measuring, you are campaigning. 1. You may not conclude the tool worked because throughput rose. The causes are not separable after the fact. 2. You may not conclude it failed because throughput stayed flat. A genuine gain can be spent on quality, on harder work, or on a deadline met without the usual overtime — none of which appear as more changes merged. **Scepticism that points only at the claims you dislike is advocacy wearing the costume of rigour.** On a subject where everyone already has a position, the candidate who applies the same standard to the result they were hoping for is the one worth hiring.
- The manager still wants one number for the quarterly review. What do you give them?A claim about a named kind of work with its boundary stated — what was counted, over what period, and whose hours. If that cannot be produced honestly, give the decision instead: what you recommend, on what evidence, and what would change your mind. A qualified answer survives being quoted; a bare number is quoted without its qualifications.
- Developers report saving time themselves. Why is that not enough?Self-reports measure the step people notice, which is the step the tool changed, and they are given by people who already hold a view about the tool. They also cannot see work that landed on somebody else. Treat them as a hypothesis about where to look rather than as the measurement.
- Which result would you treat as the strongest evidence available to you?A narrow, repeated kind of work that the team did before and still does, where the change in how it goes can be described by the people doing it and checked against something outside their impressions. It is weak evidence, and it is honest about being weak, which is more than a headline figure manages.
saying these in an interview costs you the question
- Merged volume is up, so the team is faster
- The assistant is the only thing that changed in six months
- Time saved writing code is time saved overall
- Flat throughput proves the tool did not help
- Nobody can measure this, so the question is not worth asking
- Each developer's own estimate of time saved settles it