Stakeholders demand strong consistency, sub-100 ms latency, 99.99% availability, and fast delivery of new features — all at once. How do you prioritize conflicting quality attributes and make the trade-offs explicit?
answer
- Scenario first: no measure, no trade
- Utility tree: goal → attribute → scenario (importance × risk)
- Sensitivity point vs trade-off point
- Decompose per journey/operation, not per system
- ADR + error budget + fitness functions
basics
~20 sYou cannot maximize all of them, so make the conflict visible: turn each demand into a measured scenario, rank them by business value and risk with the stakeholders who own the money, then decide per user journey — not for the whole system — and write down what you gave up.
solid answer
~1 minFirst, refuse the framing that these are independent wishes. Convert each into a **quality-attribute scenario** with a response measure, so "fast" and "consistent" become numbers per journey. Second, build a **utility tree**: business goals → attributes → refined scenarios, each rated (business importance × technical risk). Run it as a workshop with the stakeholders who own the budget and the incident pager, so the ranking is theirs, not mine. Third, evaluate candidate architectures against the top scenarios — this is essentially **ATAM**: find **sensitivity points** (a decision that strongly affects one attribute), **trade-off points** (a decision that affects two attributes in opposite directions — e.g. synchronous cross-region replication: consistency ↑, latency ↑, availability ↓), risks, and non-risks. Prototype/spike the risky ones; measure rather than argue. Fourth, **decompose the demand**: payments can be strongly consistent while the feed is eventually consistent; the checkout journey gets 99.99% while the admin console gets 99.5%. Almost every "all four" demand dissolves once scoped per journey. Finally, make it durable: record the decision and the rejected alternatives in an ADR, encode targets as SLOs with error budgets (which explicitly price availability against delivery speed), and add fitness functions so the trade-off is re-checked continuously.
code
text · 14 linesUTILITY TREE (importance, risk)
Availability
AZ loss at peak, checkout keeps serving (H,H) <- design here
Admin console down 30 min (L,L) <- accept
Consistency
Seat never double-sold during 5-min partition (H,H) <- design here
Feed item order across devices (L,M) <- accept eventual
Performance
Checkout p99 <= 800 ms @ 6k req/s (H,M)
Modifiability
New payment provider <= 3 person-days (H,M)
TRADE-OFF POINT: synchronous cross-region replication
consistency ++ RPO ++ latency -- availability -- cost --go deeper
Say that these goals conflict, give one concrete pair (consistency versus latency), and that the fix is to ask which matters most for which feature rather than promising all of them.
Turn each demand into a measured scenario, propose ranking with stakeholders, and split requirements per feature — strict for payments, eventual for the feed. Name one mechanism per attribute.
Run the evaluation: utility tree with importance × risk, sensitivity and trade-off points, spikes to measure the uncertain ones, degraded modes, and SLOs with error budgets. Give a real trade-off you made and what you accepted as worse.
Own the process and its durability: elicit business goals and translate them into a maintained scenario set, price each nine and each round trip, push consistency requirements into business processes (escrow, compensation) where cheaper, ensure the ranking is stakeholder-owned, and institutionalize it with ADRs, error budgets, fitness functions, and a review cadence tied to business change.
## Step 0 — name the conflicts out loud The four demands are pairwise antagonistic in known ways: | Pair | Mechanism of conflict | |---|---| | Strong consistency ↔ low latency | Coordination costs a round trip; cross-region quorum adds 50–150 ms (PACELC's "else" branch) | | Strong consistency ↔ availability | During a partition you must refuse or diverge (CAP) | | High availability ↔ delivery speed | Most incidents are caused by change; more deploys, more risk | | Low latency ↔ cost | Headroom, caching tiers, more regions, better hardware | | Delivery speed ↔ security/compliance | Reviews, controls, audit gates | | Availability ↔ simplicity | Failover machinery adds its own failure modes (split brain, flapping health checks) | Saying this explicitly reframes the conversation from "do everything" to "choose where to spend", which is the actual deliverable. ## Step 1 — make every demand measurable Each becomes a six-part scenario (source, stimulus, artifact, environment, response, **response measure**). "Strong consistency" must become e.g. *"a seat, once reserved, is never sold twice, even during a 5-minute inter-region partition"*. "Fast delivery" becomes *"a new payment provider is integrated in ≤ 3 person-days touching only the payments module"* and/or DORA metrics (lead time, deploy frequency, change-failure rate, MTTR). Unmeasured demands cannot be traded, only shouted about. ## Step 2 — the utility tree A utility tree (from the SEI's ATAM/QAW practice) is a four-level structure: ``` Utility ├── Availability │ ├── "AZ loss during peak" (H, H) <- (business importance, technical risk) │ └── "Admin console outage" (L, L) ├── Performance │ ├── "Checkout p99 <= 800 ms @ 6k/s" (H, M) │ └── "Report export < 60 s" (M, L) ├── Consistency │ └── "No double-sold seat under partition" (H, H) └── Modifiability └── "New payment provider <= 3 days" (H, M) ``` Each leaf is a concrete scenario with a (importance, difficulty/risk) pair. **Design against the (H,H) leaves first**; the (L,·) leaves are explicitly permitted to be mediocre. The rating must come from stakeholders — product/business owners for importance, engineering for risk. The moment engineers assign business importance themselves, the exercise loses its authority. A useful forcing device: give stakeholders a fixed budget of "high" votes, or ask them to rank rather than rate. If everything is critical, nothing is prioritized. ## Step 3 — evaluate architectures, not opinions (ATAM in miniature) For each top scenario, walk the candidate architecture and classify decisions: - **Sensitivity point** — a decision to which one attribute is highly sensitive (e.g. replication factor → durability). - **Trade-off point** — a decision that is a sensitivity point for *two or more* attributes in opposite directions (e.g. *synchronous* cross-region replication: consistency ↑, RPO ↓, but write latency ↑ and availability ↓). Trade-off points are the architecture's real decision list. - **Risk** — a decision (or absence of one) that endangers a high-priority scenario. - **Non-risk** — a decision that is fine given current assumptions; record the assumption, because it may expire. Then **buy information** where uncertainty is expensive: spike the cross-region write latency with a real prototype, run a partition drill, load-test the fan-out. "Set-based" design — carrying two options until measurement kills one — is cheaper than defending a guess. ## Step 4 — decompose the demand (the move that usually dissolves it) Global "all four" requirements are almost always an artifact of not scoping. Split by: - **Journey / bounded context**: payments strongly consistent; catalogue and feed eventually consistent; analytics best-effort. - **Operation**: writes coordinated, reads served from replicas with bounded staleness or read-your-writes sessions. - **Tier of user**: enterprise tenants get isolated cells and stricter SLOs; free tier gets shared capacity and shed first under overload. - **Time**: strict during business hours or peak season; relaxed otherwise. - **Failure mode**: full functionality normally, defined **degraded mode** during partition/overload (read-only, cached, queued-and-confirm-later). Also look for **business-process escapes**: many consistency demands are really "never embarrass us", and can be met by optimistic reservation plus compensation (overbooking + rebooking, provisional charges, idempotent reconciliation). That converts an expensive technical constraint into a cheap operational one — the highest-leverage move available. ## Step 5 — price it Make the cost curve visible: "99.9% is the current design; 99.99% adds multi-region active-active, doubles infra spend, requires 24/7 on-call and quarterly game days, and slows deploys because every change needs staged rollout. Is that worth $X?" Stakeholders almost never insist once availability is a line item rather than an adjective. The same applies to latency (headroom is idle capacity you are paying for) and to consistency (a round trip per write). ## Step 6 — institutionalize the decision A one-time decision decays. Make it live: - **ADRs** (architecture decision records): context, decision, consequences, **alternatives rejected and why**. The rejected options are the valuable part when someone re-litigates in a year. - **SLOs + error budgets**: the canonical mechanism for the availability ↔ delivery-speed conflict. Budget remaining → ship fast, take risks. Budget burned → freeze features and spend on stability. It removes the argument by pre-agreeing the rule. - **Fitness functions**: automated, continuously running checks for the attributes you claimed — architecture/dependency tests for modifiability, load tests in CI for latency, chaos/partition drills for availability, security scans for the security scenarios. Without them, "we prioritized modifiability" is unverified sentiment. - **Review cadence**: revisit the utility tree when the business changes (new market, new regulation, 10× growth). Priorities are not permanent; assumptions expire. ## Anti-patterns to call out - **Gold-plating the unranked**: five nines on an internal admin tool while checkout has none. - **Averaging the conflict**: choosing a middling design that satisfies no scenario well, instead of being deliberately excellent at the top two and deliberately mediocre elsewhere. - **Architect-decides**: making the value call yourself. Your job is to expose consequences and cost; the business owns the ranking. (Where the business cannot decide, you decide *and write down* the assumed ranking so it can be contradicted.) - **Deciding once, never measuring**: no SLO, no fitness function, so drift is invisible until an incident. - **Ignoring development-time attributes**: teams trade away modifiability and testability silently because no stakeholder names them, then pay for it every sprint. Represent them in the tree explicitly. - **Treating constraints as attributes**: a regulatory requirement is not tradeable; put it in the constraints list and design within it. ## What a strong interview answer sounds like A sentence of framing ("these conflict in known ways, so the deliverable is a ranking, not a promise"), the mechanism (scenarios → utility tree → sensitivity/trade-off points → measured spikes), the decomposition move (per journey/operation), the pricing move (cost curve, error budget), and the durability move (ADR + fitness functions) — ideally illustrated with one real trade-off you made, including what you deliberately let be worse.
- The business owner insists all four requirements are non-negotiable. What do you actually do?Stop arguing in adjectives and produce two artifacts: (1) each demand as a measured scenario per journey, and (2) a cost/consequence table for satisfying each at the stated level, including money, staffing, and delivery impact. Then present two or three concrete architecture options with what each is deliberately worse at. Faced with priced options rather than a refusal, stakeholders almost always rank. If they still will not, you write down the ranking you are assuming, get silence-implies-consent on record in an ADR, and design to it — an explicit assumption can be corrected, an implicit one cannot.
- How does an error budget resolve the availability-versus-delivery-speed argument specifically?It converts a values dispute into an accounting rule agreed in advance. A 99.9% SLO grants 0.1% unreliability per window as a budget. While the budget is unspent, the team ships aggressively — that unreliability is what the budget is for. When it is exhausted, an automatic policy takes effect: feature freeze, stability work, stricter rollouts. Nobody has to win an argument during an incident, and it also disciplines the reliability side, since an unspent budget signals over-investment in stability and under-shipping.
- What is the difference between a sensitivity point and a trade-off point in an architecture evaluation?A sensitivity point is a decision to which one quality attribute is strongly sensitive — change the replication factor and durability changes sharply. A trade-off point is a decision that is a sensitivity point for two or more attributes that move in opposite directions — synchronous cross-region replication raises consistency and durability while lowering write latency and availability. Trade-off points are where the real architectural decisions live and where evaluations concentrate their attention; sensitivity points that affect only one attribute are usually just tuning.
A city planning meeting where everyone wants wider roads, more parkland, cheaper housing, and lower taxes. No design satisfies all four; the planner's job is not to invent one but to publish the trade curves — this many hectares of park costs this much housing — and let the elected owners choose, then zone it so the choice survives the next council.
saying these in an interview costs you the question
- Promising all conflicting attributes can be satisfied with "good engineering" instead of surfacing the trade-offs.
- Making the business value call yourself instead of getting the ranking from the stakeholders who own the budget and the pager.
- Ranking every scenario as high priority, which nullifies the prioritization.
- Choosing a compromise design that satisfies no top scenario well rather than being deliberately excellent at the top two.
- Applying one global consistency and availability target instead of decomposing per journey, operation, or tenant tier.
- Presenting availability as an adjective rather than a priced line item with infrastructure, staffing, and delivery-speed costs.
- Deciding once with no ADR, no SLO, and no fitness function, so the trade-off silently erodes.
- Silently trading away development-time attributes (modifiability, testability) because no stakeholder names them.
- Treating a legal or platform constraint as a tradeable quality attribute.