Which suite-health indicators would you publish as team targets, and how do you keep them from being gamed?
answer
- An indicator is not a gate
- Leading steers, lagging confirms
- Name the cheapest dishonest move
- Publish the counterweight in the same view
- Direction, team level, never compensation
basics
~20 sPublish a small paired set: flake-rate trend beside quarantine size and age, runtime beside the count of blocking cases, pass rate beside flake rate and time to repair. Keep them as team-level indicators, never thresholds with consequences.
solid answer
~50 sFirst separate a gate, which blocks something and must be few and unambiguous, from an indicator, which is published to start a conversation. Suite health is almost entirely the second. Then choose a set of three to five that mixes leading indicators you can steer - flake-rate trend, quarantine size and age, runtime against budget, review latency, coverage trend on changed code - with lagging ones that are trustworthy but late, such as pipeline pass rate and time to repair. For each, name the cheapest dishonest way to move it and publish the counterweight in the same view, so gaming stops being invisible: excluding cases shows up as a growing quarantine list, retrying to green shows up in the flake rate. Aggregate at team level, report direction rather than thresholds, compare a team only against itself, and never attach any of it to compensation.
go deeper
You are unlikely to be asked to design this, but be able to say why a number that becomes a target stops describing reality, and to give one example such as excluding cases to improve a reliability figure.
Explain the difference between an indicator and a blocking gate, and give the cheap dishonest move for two or three concrete suite-health numbers. Knowing that trends beat levels is expected here.
Show that you have chosen a set for a real team, including what you deliberately left out and why. Be able to walk a specific pairing end to end and describe how you spotted a number moving for the wrong reason.
Own the organisational design: aggregation level, cadence of review, how the set is presented to leadership, and the honest admission that you are choosing which distortion to accept rather than eliminating gaming. Be ready to refuse a single cross-team score and to say what you offer instead.
### Start by separating an indicator from a gate Two different objects get called "a metric" and conflating them causes most of the damage. A **gate** is a threshold that blocks something: a run, a merge, a release. Gates must be few, unambiguous and cheap to satisfy honestly, because a gate that is expensive to satisfy honestly gets satisfied dishonestly. An **indicator** is a number published to start a conversation. It is not a threshold, nothing is blocked by it, and its job is to make a trend visible to the people who can act on it. Suite health belongs almost entirely in the second category. The instant an indicator becomes a target with consequences attached, you have invoked **Goodhart's law** — in Strathern's widely quoted formulation, *when a measure becomes a target, it ceases to be a good measure* — and every number below has a cheap, dishonest way to move it. ### Leading and lagging **Lagging indicators** report an outcome after the fact. They are trustworthy but arrive too late to steer: pipeline pass rate, time to repair a broken shared branch, how much rework a release caused. **Leading indicators** move before the outcome does, so they are steerable, and precisely because they are steerable they are gameable: flake-rate trend, quarantine list size and age, suite runtime against its budget, review latency, the share of changes that arrive with tests, and coverage *trend* on changed code rather than coverage level over the whole codebase. Publish both. A leading set alone tells you what people are doing; a lagging set alone tells you what happened but not what to change. ### Pair every indicator with the thing its gaming would break This is the practical answer to Goodhart, and it is what separates a principal answer from a list of metrics. For each number, name the cheap distortion and publish the counterweight in the same view: | Indicator | Cheapest way to move it dishonestly | Publish alongside | | --- | --- | --- | | Flake-rate trend | Exclude cases from the counted set | Quarantine list size and age of oldest entry | | Suite runtime vs budget | Delete or skip cases; weaken assertions | Count of blocking cases; coverage trend on changed code | | Pipeline pass rate | Retry until green; move checks off the blocking path | Flake rate; number of blocking checks; escapes past the pipeline | | Coverage trend | Execute code without asserting anything | Change-level review; trend restricted to changed lines | | Review latency | Approve without reading | Change size distribution; rework rate on merged changes | | Time to repair | Reclassify a red branch as "known issue" | Count and age of known-issue exclusions | No pairing is airtight. What pairing buys is that the dishonest move stops being *invisible*: it shows up as an odd shape somewhere else in the same dashboard, and someone asks about it. ### The governance that decides whether any of this works - **Aggregate at the team level, never the individual.** Individual attribution converts every indicator into a performance instrument within a quarter, and the distortion follows immediately. - **Direction, not thresholds.** "Flake rate falling over the last two quarters" is a conversation. "Flake rate below 1%" is a target, and targets get met by relabelling. - **Comparison against yourself, not across teams.** A team owning an integration-heavy dispatcher service and a team owning a small library have incomparable numbers, and ranking them teaches both to optimise presentation. - **Review the metric set itself on a cadence.** Retire indicators that have saturated or gone flat. A set that never changes is a set nobody is reading. - **Never attach the numbers to compensation or promotion.** This is the single strongest predictor of whether the set survives contact with the organisation. ### A concrete read Nine teams on a ride-hailing dispatcher platform report against a 27-minute shared-branch budget. One team's flake rate falls from 4.1% to 1.2% across a quarter — the best movement on the board. Read alongside its pair, the quarantine list went from 14 entries to 63, oldest entry 71 days. The suite got quieter, not better, and one of those entries was the case asserting that a support-desk role cannot reassign another operator's ride; a **permission escalation** shipped and was found 12 days later. Nobody involved acted in bad faith: they were told flake rate was the number that mattered, and they moved it the cheapest available way. That is Goodhart operating exactly as advertised, and the fix is structural — publish the pair — not a conversation about integrity. ### What to say when pushed Be honest that no metric set is un-gameable, and that the choice is *which distortion you can live with*. Then say what you would actually do first: pick a small set — three to five — that a team can hold in its head, pair each one, review them in a retrospective rather than a status report, and change the set when it stops producing arguments. An indicator that never provokes a disagreement has either been solved or is being managed for presentation, and both mean it is time to retire it.
- Leadership wants a single suite-health score across nine teams. What do you tell them?That a single score across incomparable systems will be optimised for presentation within a quarter. A team owning an integration-heavy service and a team owning a small library have structurally different numbers, so a composite mostly ranks their architecture. Offer instead a per-team view compared against its own trend, with the same handful of paired indicators, and a written narrative for anything moving sharply. If a single number is unavoidable, make it a count of teams improving rather than an average of levels.
- One team's flake rate improved more than everyone else's this quarter. What do you check before praising it?The paired indicator. The cheapest way to lower a flake rate is to stop counting cases, so look at the quarantine list size and the age of its oldest entry over the same period. If exclusions grew alongside the improvement, the suite got quieter rather than better, and the unguarded behaviour is now invisible. Also check whether retries were introduced or widened, since a retry that overwrites the first result erases the evidence the metric depends on.
- Why is coverage trend on changed code preferable to coverage level over the whole codebase as an indicator?Because a whole-codebase level is dominated by history nobody is touching, so it moves too slowly to steer anything and invites bulk gaming. Trend restricted to the code a change actually touches is responsive, is attributable to a decision someone just made, and is a reasonable conversation in review. It remains gameable by executing code without asserting on it, which is why it is published as a discussion input rather than as a gate.
- How do you know when to retire an indicator from the set?When it stops producing arguments. An indicator that has been flat and uncontested for several review cycles has either been genuinely solved, in which case keeping it costs attention for nothing, or it is being managed for presentation, in which case it is actively misleading. Review the set itself on a fixed cadence, retire the quiet ones, and add one that reflects whatever the team is currently struggling with.
saying these in an interview costs you the question
- Turns every suite-health indicator into a blocking threshold
- Publishes a single composite score across teams with different systems
- Attaches the numbers to individual performance or compensation
- Reports flake rate with no view of exclusions that could explain it
- Assumes a well-chosen metric cannot be gamed at all
- Never revisits the metric set once it is on a dashboard