Before a design system's first release reaches a pilot team at a freelance job marketplace, what success criteria would you define, and why set them in advance?
answer
- decide before, not after
- outcome over output
- a shipped feature, not a count
- quality, effort, satisfaction
- agree what happens on a miss
basics
~20 sDefine criteria before launch so the decision to widen, fix or pivot is not rationalised afterwards. Measure outcomes: the pilot ships features on the system without forking, components meet the quality bar, build effort drops, and the pilot would recommend it.
solid answer
~50 sCriteria set in advance keep the next decision honest; without them, any result can be declared a success. I would define a small set covering **outcome** — the pilot ships its planned proposal-flow feature using system components, with no forks; **quality** — the included components pass the system's release-readiness bar and the pilot's new screens have no critical accessibility defects against WCAG 2.2 Level AA; **effort** — building a representative screen takes less time than the recorded baseline; and **experience** — the pilot team rates the system well and would recommend it to other teams. I would also agree a response time for blocking bugs, and what happens on a miss: usually fix and extend the pilot rather than widen anyway. I avoid output measures like the number of components shipped, which say nothing about value. Tracking adoption across the organisation later is a separate, ongoing measurement.
go deeper
Recall that a design system's first release should have agreed goals, and that they describe what changed for the pilot team, not just what was built.
Explain the difference between output and outcome criteria, and why a baseline and a small set of measures make the result trustworthy.
Show a concrete starter set — outcome, quality, effort, experience, responsiveness — with owners, a baseline and a decision rule for what a miss triggers.
Treat the criteria as a contract with leadership: they decide when the system earns wider investment and protect the team from widening before it is ready.
## Why set criteria before the release A **design system's** first release goes to a **pilot team** — the first product team to build with its shared foundations and components. Whatever happens next — widening to more teams, fixing and extending the pilot, or rethinking the approach — should follow from evidence. **Success criteria agreed in advance** make that possible: - They stop goalposts moving after the results arrive. - They force the system team and the pilot to agree on what the release is for. - They tell everyone which data to collect from day one, including a baseline. - They make a miss discussable without blame, because the rule was known. ## Outcomes, not outputs An **output** is what the system team produced: components shipped, pages of documentation written. An **outcome** is what changed for consumers and users: a feature shipped faster, fewer defects, a team that wants more. Output criteria are easy to hit while the release fails — ten components nobody uses score perfectly. First-release criteria should be mostly outcomes. ## A starter set | Dimension | Measure | Example target | Why it matters | |---|---|---|---| | **Outcome** | Pilot feature shipped on system components | Planned proposal-flow redesign ships with no forked components | Proves the system is usable for real work | | **Quality** | Components meet the system's readiness bar; pilot screens checked for accessibility | No critical defects against WCAG 2.2 Level AA on the new screens | Early trust depends on components that work | | **Effort** | Time to build a representative screen, compared with a baseline | Noticeably less than the recorded baseline | Shows the system saves effort | | **Experience** | Pilot team survey and interviews | Most of the team would recommend it to another team | Predicts how the story spreads | | **Responsiveness** | Time for the system team to fix blocking bugs | Within the agreed window | Tests whether support can scale | The targets are examples; each organisation sets its own, and a baseline measured before the pilot matters more than the exact number. ## Making them workable 1. **Measure a baseline first** — how long the pilot currently takes to build a comparable screen, and how many UI defects it currently ships. 2. **Keep the set small** — four or five criteria that someone will actually track. 3. **Name an owner** for each measurement. 4. **Agree the decision rule** — which misses mean extend the pilot, which mean rethink, and what must be true before widening. 5. **Review at a set date** rather than whenever it feels finished. ## What these criteria are not - They are not **ongoing adoption metrics** across the whole organisation, which are tracked continuously once the system is live. - They are not **release cadence** goals for later versions. - They are not a component checklist; component readiness is one input, not the verdict. ## Example: a freelance job marketplace Before the pilot, the system team measures how long the search-and-proposals team takes to build a typical screen and how many UI bugs its last two releases carried. They agree: the proposal-flow redesign ships on system components without forks; the new screens have no critical accessibility defects against WCAG 2.2 Level AA; a comparable screen takes clearly less time than the baseline; most of the team would recommend the system; blocking bugs are fixed within two working days. Eight weeks later, all but the effort target are met — building took about as long as before, because the documentation was thin. The agreed rule says to extend the pilot and fix the documentation before widening, and that is what happens.
- Why is the number of components shipped a weak success criterion for a design system's first release?It measures the system team's output, not whether anyone benefited. Ten components nobody uses score well while the pilot forks everything. Criteria should describe what changed for consumers and users — a feature shipped on the system, less effort per screen, fewer defects — so that a rising count cannot hide a failing release.
- What should happen if a design system's first release misses its success criteria?Apply the decision rule agreed in advance: usually extend the pilot, fix the causes the misses reveal and re-measure, rather than widening the rollout on schedule. A miss is cheap, valuable information at pilot scale; widening anyway spreads the same problems to every team and spends trust that is hard to regain.
saying these in an interview costs you the question
- Success criteria can be chosen after launch, once the data is in.
- Shipping the planned number of components means the release succeeded.
- If the pilot misses its criteria, widen the rollout anyway to keep momentum.
- Pilot team satisfaction is too subjective to be worth measuring.
- A baseline is unnecessary; improvement will be obvious without one.