skip to content

How do you justify pair programming's cost when the evidence for it is contested?

level: principalimportance: should knowfreq 38%

answer

  1. Meet the two-salaries objection directly
  2. Refuse to quote a multiplier
  3. Say why the studies disagree
  4. Bounded trial, signals agreed first
  5. Throughput alone rigs the answer

basics

~20 s

Do not quote a multiplier — published evidence on effort and defect rates is mixed and study conditions vary widely. Argue from a bounded local trial with signals agreed up front, and be explicit about what pairing does not replace.

solid answer

~50 s

The naive frame is two salaries for one output, and it is worth meeting head-on rather than dodging. The claimed offsets are fewer defects escaping to later stages, faster onboarding, fewer areas with a single knowledgeable person, less rework from a design discovered too late, and feedback that arrives while the code is being written instead of days after. The honest complication is that the research is genuinely mixed: studies differ in task size, participant experience, pairing style and duration, and the reported effects on both effort and defects vary enough that no single credible multiplier exists to quote. So argue locally: run a bounded trial on a defined class of work, agree the signals in advance, and accept that throughput alone is the one measure guaranteed to make pairing look bad. Pair the argument with a clear statement of what pairing does not replace.

go deeper

for a junior

Be able to state the obvious objection — two people, one change — and name two or three benefits that are claimed against it, without pretending the question is settled.

for a middle

Explain why published studies disagree: participant experience, task size, pairing style and which outcome was measured all vary, so no single multiplier is quotable.

for a senior

An interviewer expects you to design an honest local trial: a defined class of work, signals agreed up front, a stated duration, and a straight reading of a result that is not controlled.

for a principal

Own the organisational side: how to make the case without leaning on contested numbers, why throughput alone rigs the outcome, and what you commit to keeping regardless — checks, outside review, written decisions.

## Meet the objection as stated The objection is arithmetic: two engineers, one change, so the change costs twice as much. Any answer that dodges this loses the room. The correct move is to accept the framing and then argue that the denominator is wrong — you are not buying keystrokes, you are buying a smaller amount of rework, defects caught earlier, and knowledge held in more than one head. ## Be honest that the evidence is contested There is a real research literature on pair programming, and it does not settle the question. Studies vary in ways that plausibly change the result: - **Participants.** Much of the early work uses students on short exercises; teams doing sustained work on a large unfamiliar system are a different population. - **Task size and duration.** A one-hour exercise and a three-month feature exercise different mechanisms. - **Style.** Ping-pong, strong-style and "two people at one screen with no rules" are not the same intervention, and are often not distinguished. - **Outcome definition.** Defect counts, defect escape rate, effort, elapsed time and quality ratings are all reported, and a result can be positive on one while negative on another. The result is that reported effects on effort and on defect rates run in both directions across studies. A candidate who confidently quotes a fixed figure — "pairing costs fifteen percent more and halves defects" — is quoting one study as if it were a law. The stronger position is to say the literature is mixed, name why, and move the argument onto ground you control. The same discipline applies to related folklore. The claim that defects get an order of magnitude more expensive at each later stage is widely repeated and widely disputed in its precise form; leaning on it to sell pairing puts a contested number underneath a contested practice. ## Argue from a bounded local trial What a lead can actually offer is evidence from this team, on this system, over a bounded period, with the signals agreed before it starts. Worked shape, on an 11-person team owning a hospital appointment scheduler: - **Scope.** Pair on the two highest-risk items per iteration — reversibility and blast radius decide which — and leave everything else alone. Not a blanket mandate. - **Duration.** Nine iterations, decided up front, so it cannot be cancelled the first time a week looks slow. - **Signals, agreed in advance.** Defects escaping from the paired area; elapsed days from a new joiner's start to their first unaided change there; the number of areas exactly one person can safely change; how long a finished change waits before someone else looks at it. - **Read the result honestly.** Suppose escapes in that area fall from 17 to 6 while total items delivered drop about 12%. That is suggestive and it is not proof — nothing was controlled, the team also got more familiar with the area over nine iterations, and the count is small. Say so. A lead who presents an uncontrolled before-and-after as causal has just taught their organisation to distrust the next measurement too. ## Do not let throughput be the only measure Items completed is the measure most readily available and the one pairing is structurally worst on, because two people are producing one stream. Choosing it alone answers the question before it is asked. The counter is not to hide it — report it — but to insist that the decision is made against the full set of agreed signals, including the ones that only appear later. ## Say what pairing does not replace This is the part that separates an advocate from an engineer with judgement: - **Automated checks and the pipeline.** A pair is two people with one shared blind spot and no memory of last month's regressions. - **An independent look by someone who was not in the room.** The pair's shared context is exactly the thing an outside reader does not have, which is why they see different problems. - **Written decisions.** The pair's shared understanding is undocumented and evaporates, including for the two of them. - **Specialist attention.** Depth in security, performance or accessibility does not appear because two generalists sat together. - **Clear requirements and a working environment.** Two people can build the wrong thing efficiently, and pairing on a broken environment simply wastes two people. ## Set it as a default, not a mandate A mandate produces compliance behaviour: sessions held because they are required, with one person driving and the other on their own screen. A default for a named class of work, with a stated reason and a review date, keeps the practice honest and keeps the argument open — which, given that the general evidence really is contested, is the intellectually defensible place to stand.

  • Leadership asks you for a number: how much does pairing cost and how much does it save?
    Give the honest answer, then give them something usable. Say that published results disagree because study populations, task sizes, pairing styles and outcome definitions differ, so any single multiplier would be a study quoted as a law. Then offer a bounded trial on the riskiest class of work, with the signals and the review date agreed in advance, and commit to reporting throughput alongside the rest rather than hiding it. That converts an unanswerable general question into a decision this organisation can actually make.
  • Which measurement would you refuse to make the deciding one, and why?
    Items completed per iteration. It is the easiest number to get and the one pairing is structurally worst on, because two people deliberately produce one stream, so choosing it as the deciding measure settles the question before the trial runs. Report it, because hiding it destroys trust, but decide against the full agreed set: escaped defects in the paired area, time for a new joiner to work unaided, and how many areas still have exactly one person who can change them.
  • What does pairing fail to replace, even when it is working well?
    Automated checks and the pipeline, because a pair shares one blind spot and has no memory of past regressions. An independent look from someone outside the pair, precisely because they lack the shared context. Written decisions, since the pair's understanding is undocumented and fades. Specialist depth in areas like security or performance. And it fixes neither an unclear requirement nor a broken environment — it just puts two people in front of them instead of one.
  • How would you introduce pairing to a team that is openly sceptical of it?
    Concede the cost argument first, and make the scope small enough to be safe: a named class of high-risk work, a stated end date, and signals agreed with the sceptics rather than presented to them. Let people choose their rhythm and rotation rather than prescribing one. Then report the result straight, including the part that looks bad. A trial that ends with an honest "this did not help here" buys more credibility for the next proposal than an adopted practice nobody believes in.

saying these in an interview costs you the question

  • Quotes a precise cost or defect multiplier as settled fact
  • Treats early student-exercise studies as evidence for sustained team work
  • Claims paired code needs no automated checks or outside reader
  • Sells pairing on the disputed late-defect cost curve
  • Measures the trial on delivered items alone
  • Imposes pairing as a mandate with no review date

context