As the architect for a large system, how would you derive a fitness-function suite from the system's architectural characteristics, and how do you govern it over time — including what to do when two fitness functions conflict?
answer
- drivers → few characteristics → forced priority order
- measurable statement: threshold + context + method
- cheapest sufficient mechanism, placed by cost
- each rule linked to an ADR with an owner
- conflicts: priority order, renegotiate openly, record, add holistic check
basics
~20 sPick the few characteristics that really drive the design, make each measurable with a threshold, and automate it in the pipeline. Review the suite as business needs change. When two conflict, make the trade-off explicit and decide by business priority — do not silently weaken one.
solid answer
~60 sStart from drivers, not tools: elicit the handful of architectural characteristics the business actually depends on (say availability, checkout latency, regulatory auditability, team-level deployability) and force a priority order — everything cannot be top priority. Turn each into a measurable statement with a threshold and a context ('p99 checkout under 400 ms at 2,000 concurrent users'), then choose the cheapest mechanism that assesses it: dependency test, pipeline gate, scanner, load test, production monitor, chaos experiment, or a scheduled manual review when automation is impossible. Classify each as atomic/holistic and triggered/continuous to place it in the pipeline by cost. Govern it like code: every rule traceable to an architecture decision record with an owner and rationale, changes reviewed like production changes, suppressions visible and counted, flaky functions fixed or removed because ignored red destroys the whole mechanism, and the suite revisited when business drivers shift. Conflicts — encryption versus latency, consistency versus availability — are not bugs; they are the trade-off surfacing. Resolve by re-checking the priority order, adjust the losing threshold deliberately, record the decision, and add a holistic function covering the interaction.
go deeper
Say the checks should come from what the business needs most, that each needs a concrete number to test against, and that they run automatically.
Walk the derivation: characteristics → measurable statements with thresholds → mechanism and pipeline placement, and note that conflicting checks need an explicit decision rather than silently lowering a threshold.
Add classification by cost, pipeline budget, ADR traceability, suppression/baseline metrics, flakiness policy, and a structured conflict resolution using the agreed priority order.
Cover the whole governance system: eliciting and prioritising a small set of characteristics with the business, budgeting verification cost, versioning the suite alongside decisions, retiring obsolete functions, treating conflicts as trade-off negotiations with recorded outcomes, and replacing review-board gatekeeping with automated governance co-owned by teams.
## Step 1 — Derive characteristics from drivers, not from a checklist Architectural characteristics come from business goals, service-level objectives, regulatory obligations, operational realities, and team topology — never from a generic '-ilities' list. Practical sources: the domain's cost of downtime, the revenue sensitivity to latency, the audit regime, the growth curve, the number of teams that must ship independently. Two disciplines matter here: - **Keep the list short.** Seven or eight characteristics is already a lot; every one you add costs verification and constrains design. - **Force a priority order.** Characteristics conflict by nature, so an unordered list is not a decision. 'Availability > auditability > latency' is a decision; 'all of them are critical' is an abdication that will be resolved later, badly, by whoever is on call. ## Step 2 — Make each characteristic measurable A characteristic without a number is not enforceable. Convert: - 'The system should be fast' → 'p99 latency of `POST /checkout` ≤ 400 ms with 2,000 concurrent users on production-sized data.' - 'The system should be modular' → 'no cycles between modules; only `api` packages are exported; no module depends on more than N others.' - 'The system should be secure' → 'no dependency with a known critical CVE older than 7 days; every endpoint declares authorization; no secret literal in source.' - 'The system should be resilient' → 'with any single instance terminated, error rate stays below 0.1% and no order is lost.' Each statement needs a **threshold**, a **context** (load, data volume, environment), and a **measurement method**. ## Step 3 — Choose the mechanism and cadence For each measurable statement pick the cheapest sufficient mechanism, then classify it: - **Atomic + triggered** (dependency tests, scanners, lint gates): cheap, run on every commit. - **Holistic + triggered** (combined load + security scenario, resilience scenario): expensive, run nightly or pre-release. - **Atomic/holistic + continuous** (production monitors, synthetic transactions, always-on chaos): run in production with alerting and named owners. - **Manual** (legal/data-residency review, accessibility audit): scheduled, with an explicit checklist — a fallback, never the default. Budget the total: a suite whose pipeline cost pushes feedback past ~10 minutes for the fast stage will cause batching, which destroys incremental change — the very thing the suite exists to protect. ## Step 4 — Govern the suite like production code - **Traceability.** Every fitness function links to the architecture decision record it enforces, with rationale and owner. A rule nobody can justify will eventually be deleted under delivery pressure. - **Change control.** Weakening a threshold is an architectural decision, not a build fix; it goes through the same review as any production change and updates the ADR. - **Suppressions and baselines are metrics.** Track the count and the trend. A shrinking frozen baseline is healthy; a growing allow-list means the rule is decorative. - **Flakiness is fatal.** An unreliable fitness function teaches people that red means 'rerun'. Fix or delete it; there is no third option. - **Periodic review.** Revisit the suite when drivers change (new market, new regulation, 10× traffic, a reorg that changes team boundaries). Retire functions whose characteristic no longer matters; the suite should not only grow. - **Onboarding value.** The suite is executable documentation — new engineers learn the constraints from failure messages, which is why messages must explain *why*, not just *what*. ## Step 5 — Handling conflicts between fitness functions Conflict is the normal state, because characteristics trade off: - Security (encrypt everything, deep validation) vs performance/latency. - Strong consistency vs availability under partition. - Elasticity/dynamic scaling vs cost ceilings. - Fine-grained modularity/deployability vs end-to-end latency and operational overhead. - Auditability (retain everything) vs privacy/data-minimisation regulation. A disciplined resolution: 1. **Surface it explicitly** — ideally the conflict is discovered by a *holistic* fitness function rather than by an incident. 2. **Return to the priority order** agreed with the business. The higher-priority characteristic keeps its threshold. 3. **Renegotiate the loser's threshold consciously**, with the business owner, rather than quietly disabling the check. 'p99 rises from 200 ms to 320 ms because payload encryption is mandatory' is a recorded decision. 4. **Look for a design change that dissolves the conflict** before conceding — caching, session-level rather than per-request crypto, hardware acceleration, asynchrony, read models. 5. **Record it** in an ADR and update both functions so the suite stays truthful. 6. **Add a holistic function** covering the interaction so the trade-off stays verified as the system evolves. The anti-pattern is a fitness function that is 'temporarily' disabled or has its threshold relaxed by whoever hit it. At that point the suite stops describing the architecture and starts describing whatever the code happens to do — which is where the organisation was before it had fitness functions. ## Organisational dimension Fitness functions replace the architecture-review-board bottleneck with automated governance, letting teams move independently while the constraints hold. That only works if teams see the rules as *theirs*: co-author them with the teams, expose the suite's health as a first-class dashboard, and treat the architect's job as curating the suite rather than approving individual changes.
- How many fitness functions should a system have?As few as cover the prioritised characteristics and no more. Each one costs runtime, maintenance, and credibility when it misfires. A large suite that people routinely bypass is worse than a small suite that is always trusted and never disabled without a recorded decision.
- A performance fitness function starts failing after a mandatory security change. What is the correct response?Treat it as the trade-off surfacing. Check the agreed priority order; look for a design change that dissolves the conflict (caching, session-level crypto, asynchrony); if none exists, renegotiate the latency threshold explicitly with the business owner, record the decision in an ADR, update the function, and add a holistic check covering the interaction. Never silently disable it.
- How do fitness functions change the role of an architecture review board?They shift governance from synchronous, per-change human approval to automated, always-on enforcement. The architect's work moves upstream: choosing and prioritising characteristics, co-authoring the rules with teams, and curating the suite — while teams ship independently because compliance is verified continuously rather than in a meeting.
saying these in an interview costs you the question
- Deriving characteristics from a generic '-ilities' checklist instead of business and regulatory drivers
- Refusing to prioritise characteristics, leaving conflicts to be resolved ad hoc during incidents
- Writing subjective rules with no threshold, context, or measurement method
- Relaxing a threshold or disabling a check to make the build green without recording an architectural decision
- Tolerating flaky fitness functions, which trains everyone to treat red as 'rerun'
- Letting the suite only grow — never retiring functions whose characteristic no longer matters
- Assuming the suite replaces architectural design or business conversation rather than encoding their outcomes