Leadership asks for one number proving the design system works; what would you tell them, and what would you report instead?
answer
- a target stops measuring
- a small balanced scorecard
- adoption, fit, sentiment, service
- every number tied to a decision
- trends per product, not snapshots
basics
~20 sNo single number proves a design system works, and any number made a target gets gamed. Report a small scorecard: weighted UI coverage, usage and version lag, detach and override rates, satisfaction, and the system team's service health, per product over time.
solid answer
~50 sI'd say one number would be easy to give and misleading: import counts can be inflated by wrapping, and coverage can be pushed by mandate while teams resent the system. Once a single number becomes the target, it stops measuring anything. Instead I'd report a **small scorecard**, each measure tied to a decision: **weighted UI coverage** of the main flows, for reach; **usage and version lag** per product, for health of the dependency; the **detach and override rate**, for fit; a **satisfaction survey** of consuming designers and engineers, for whether adoption is willing; and **service health**, such as response time to requests and age of open defects, for whether the system team is keeping up. I'd show trends per product and platform, with a short narrative of what changed and why. If pressed for a headline, I'd give weighted coverage of the core flows alongside the satisfaction trend, never alone.
go deeper
Recall that no single number shows a system is working, and name a few signals a system team reports, such as coverage, usage and satisfaction.
Explain what each scorecard measure answers and how each can be gamed when it becomes a target on its own.
Show how you would build and run the scorecard: sources, per-product trends, survey design, and which decision each measure feeds.
Defend the refusal of a single number to leadership, choose a paired headline if forced, and decide how mandate versus voluntary adoption changes which measures carry weight.
## Why one number fails A **design system** exists to make many product teams faster and their products more consistent and accessible. No single metric captures all of that, and any single metric that becomes a target changes the behaviour it was meant to observe, an effect usually summarised as **Goodhart's law**: when a measure becomes a target, it ceases to be a good measure. - **Import or usage counts** rise when teams wrap local work in thin system shells. - **UI coverage** rises under a mandate even when teams work around the system and resent it. - **Satisfaction** can be high among a small group of enthusiastic early adopters while most products ignore the system. - **Detach rates** fall when teams stop detaching and start hand-building instead. The honest answer to 'one number' is a short explanation of that, followed by something better. ## A balanced scorecard A scorecard of four to six measures covers the dimensions that can each be gamed alone but not all at once. | Dimension | Measure | Question it answers | Decision it feeds | |---|---|---|---| | **Reach** | UI coverage of core flows, weighted by use | How much of what users see does the system supply? | Where to invest in new components or patterns | | **Dependency health** | Usage per product and **version lag** | Are products on current releases? | Which teams need upgrade help | | **Fit** | Detach and override rate by component | Where does the system fail real needs? | Which variants, defaults or defects to fix | | **Sentiment** | Recurring satisfaction survey | Is adoption willing or forced? | Whether governance or support needs to change | | **Service health** | Response time to requests, age of open defects, time to ship a fix | Is the system team keeping up? | Staffing and prioritisation of the system team | Each row is tied to a decision. A measure that feeds no decision should be dropped, because collecting it costs the consuming teams effort and goodwill. ## Designing the satisfaction survey The survey is the only row that measures sentiment directly, so its design matters: 1. **Keep it short** and recurring, for example twice a year, so response rates stay high. 2. **Hold a fixed core** of the same few questions every round so trends are comparable; rotate one or two topical questions. 3. **Survey both designers and engineers**, on every platform, and report them separately; their pain points differ. 4. **Include free text**, such as 'what did you last have to build yourself?', which often names the missing variant directly. 5. **Close the loop** by publishing what the team changed in response; otherwise response rates fall. ## Service health indicators Service health measures the system team as a service provider rather than the products. Typical indicators: - Time from a request or bug report to first response, and to resolution. - Age and count of open defects, especially accessibility defects. - Time from a fix landing to it being available in a release consumers can take. - Share of consumers on the latest major version, which partly reflects how easy upgrades are. A system with excellent coverage but a months-long queue of unanswered requests is heading for forks. ## Presenting it to leadership - Show **trends per product and platform**, not a single snapshot; the direction is the story. - Lead with a **short narrative**: what moved, why, and what the team will do next. - If a headline is demanded, pair two numbers that check each other, such as weighted coverage of core flows and the satisfaction trend. Coverage rising with satisfaction falling is a warning; both rising is the result leadership wants. - Keep the business case separate: cost savings and delivery speed arguments rest on their own evidence and methods, and mixing them into the adoption scorecard muddies both. ## An example For a customer-support ticketing tool, the scorecard might show coverage of the agent flows rising from half to two-thirds over a year, version lag shrinking after an assisted upgrade push, a detach spike on status badges traced to a missing status and fixed, satisfaction flat among engineers but up among designers, and median response time to requests at two days. That picture tells leadership far more than '87% adoption', and each line tells the system team what to do next.
- If leadership insists on a single headline number, which would you choose?Weighted UI coverage of the product's core flows is the closest to what users experience, so I would lead with it, but always presented next to the satisfaction trend. Coverage rising while satisfaction falls signals forced adoption; the pairing makes the headline hard to game without being noticed.
- How does measurement differ for a mandated system versus a voluntary one?Under a mandate, usage and coverage rise regardless of quality, so they say little; satisfaction, detach rates and service health carry the signal. In a voluntary system, usage and coverage are themselves evidence of value, because teams chose them, so reach metrics carry more weight.
saying these in an interview costs you the question
- A single adoption percentage is enough to prove the system works.
- Making coverage a team target will not change how teams behave.
- One snapshot of the numbers is enough; trends over time add nothing.
- Cost savings and adoption metrics belong in the same headline figure.
- A survey with entirely new questions each round still shows a comparable trend.