How do you judge whether Three Amigos discovery workshops repay their cost across a dozen teams?
answer
- Prevention leaves no visible trace
- Never measure attendance or session count
- Count what the sessions catch
- Split defects by cause before claiming credit
- Depth varies with ambiguity and blast radius
basics
~20 sMeasure what the sessions catch, not that they happen. Watch questions raised before build, stories split or sent back, and requirement-misunderstanding defects found after build -- and vary the depth of discovery by how ambiguous and how costly the requirement is.
solid answer
~50 sCount outcomes, never attendance. The signals worth tracking are leading ones: open questions raised and answered before a story enters development, stories split or held at the board, and rules whose wording changed during the session -- each is a misunderstanding that did not become code. The lagging signal is the share of post-build defects classified as *requirement misunderstanding* rather than implementation error; that share falling is the claim the practice is actually making. Then treat depth as a dial rather than a policy: a novel pricing rule with money and regulation attached deserves the full triad, while a well-understood change may need only an asynchronous example review by one other perspective. Watch for the failure signature of ritual -- sessions that finish early with no questions, maps with one example per rule, examples written afterwards by one person. Avoid defending the practice with contested cost-of-late-defect multipliers; argue from your own defect classification instead.
go deeper
Understand that these sessions cost real hours and that their benefit is prevention, which is invisible by nature -- so the case for them has to be made with evidence, not with how diligently they are attended.
Be able to name what a session catches that would otherwise cost more later: questions closed before build, stories split, ambiguous wording fixed. Those catches are the observable part of the return.
Show you can spot ritual -- sessions with no questions, one example per rule, write-ups done afterwards by one person -- and that you would fix the structure rather than exhort the team to care more.
Own the allocation. Set depth by ambiguity and blast radius, instrument outcomes rather than ceremony, argue from your own defect-cause data instead of contested industry multipliers, and be willing to withdraw the practice where it is not paying.
### The question behind the question A dozen teams running a thirty-minute discovery workshop for two or three stories a week is a real and recurring spend, and it is a spend on prevention, whose return is invisible by construction: the defects it stops leave no trace. So the practice is defended with anecdotes and attacked with arithmetic, and it survives or dies on whoever argues better -- unless you instrument it. ### Measure the catch, not the ceremony The instinctive metric -- percentage of stories that had a session -- is the worst one available. It is trivially gamed, it rewards ritual, and a team that holds every session and asks nothing scores perfectly. Measure what the sessions *catch*: - **Questions raised and resolved before build.** Each open question closed before a story enters development is a decision that would otherwise have been made silently by one person mid-code. - **Stories split or held at the board.** A story that entered discovery as one item and left as three is discovery working exactly as intended. - **Rules whose wording changed during the session.** Ambiguity found and removed at its cheapest point. - **Rules that gained an example nobody had considered.** The clearest evidence that a second and third perspective were present. Those are leading indicators, available weekly, and they come free if the group already writes cards. ### The lagging indicator that carries the argument Classify post-build defects by cause, with at least the split between *implementation error* (the code does not do what the team agreed) and *requirement misunderstanding* (the code does what one person thought was agreed). Only the second class is what discovery claims to reduce. If that share is falling on teams running real sessions and flat on teams that stopped, you have an argument that survives contact with a finance conversation. If both are flat, the honest conclusion is that the sessions are ritual and need fixing or dropping. Be careful with the arithmetic in both directions. The widely repeated multipliers for how much more a defect costs in production than at requirements time are contested, and their original evidence base has been challenged; quoting them invites the wrong argument. Your own classified defect data is weaker in theory and far stronger in the room. ### Depth as a dial, not a policy The strongest principal-level answer is that not every requirement deserves the same discovery. Two factors set the depth: how much *ambiguity* the requirement carries, and how much a misunderstanding would *cost*. - **High ambiguity, high cost** -- a new tiering rule on a **utility billing run**, where money, regulation and re-issued statements are involved: the full triad, plus whoever owns the downstream consequence, and boundary examples on every threshold. - **Low ambiguity, high cost** -- a well-understood change with real blast radius, such as raising the concurrency of a rating service that peaks near 1,200 requests per minute: a short session focused on the failure modes rather than the rules. - **High ambiguity, low cost** -- explore it cheaply, or build a throwaway version and learn from it. - **Low ambiguity, low cost** -- one other perspective reviews the examples asynchronously, and that is enough. Insisting on a triad here is the thing that teaches teams the practice is bureaucracy. Publishing that dial is what stops the practice from being applied uniformly and resented uniformly. ### Reading the failure signature Ritual has a recognisable shape, and it is worth naming so leads can spot it without a survey: sessions that finish in eight minutes with no questions raised; maps with exactly one example under every rule; examples written up afterwards by one person for the others to approve; the same perspective missing every week. Any of these means the meeting is still happening and the mechanism is not. The fix is usually structural -- attendance, timing before commitment, a facilitator, a smaller requirement -- rather than exhortation. ### Scaling across teams without mandating Twelve teams will not adopt an identical practice, and forcing one is how you get compliance without effect. What travels well is a shared vocabulary and a shared expectation of *outputs*: whatever the session looked like, a story arriving in development carries concrete examples and no unowned open questions. How a team gets there -- a board, a call, an asynchronous thread -- is theirs. Pair that with one or two teams running the practice well and visibly, because a team that can show a boundary caught at the board persuades peers in a way a policy document does not. ### Knowing when to stop Be willing to say the practice is not paying on a given team and to withdraw it there. A team with a stable domain, a long-tenured group and a defect profile that is almost entirely implementation error is not the audience for heavy discovery, and defending the ritual there costs credibility you will need on the teams where it matters. The principal's job is the allocation, not the advocacy.
- A team wants to drop the sessions entirely. What do you ask for before agreeing?A short trial with the defect classification already in place: keep splitting post-build defects into implementation error and requirement misunderstanding for a few cycles without the sessions, and compare. If the misunderstanding share does not move, the team is right and I would rather spend the credibility elsewhere. If it climbs, the data makes the case without me arguing it -- and either way the team owns the conclusion.
- How do you avoid the metrics themselves becoming the thing that is gamed?Keep them diagnostic rather than targets, review them with the teams instead of reporting them upward, and never set a number to hit. The moment a question-card count becomes a goal, cards appear for questions nobody had. What resists gaming best is the lagging defect-cause split, because inflating it hurts the team that reports it.
- Does the practice still pay when the same three people have worked together for years?Less, and that is a legitimate reason to lighten it. Long-tenured groups on a stable domain share a lot of context implicitly, so the sessions catch fewer misunderstandings -- their defect profile shows it. The place it still pays is novelty: an unfamiliar rule, a new regulatory constraint, a new joiner. Varying depth by requirement rather than by team keeps the value without the ceremony.
saying these in an interview costs you the question
- Measures the practice by how many sessions were held
- Quotes a fixed late-defect cost multiplier as settled fact
- Mandates one identical process across every team
- Claims discovery workshops remove production escapes
- Never classifies defects by cause before claiming credit
- Refuses to withdraw the practice from a team it is not helping