How do you evaluate competing architectural style candidates rigorously rather than by opinion? Describe a concrete method and what its outputs are.
answer
- scenario = stimulus + environment + response measure
- utility tree: attributes → prioritised scenarios
- ATAM outputs: risks, non-risks, sensitivity, tradeoff points
- tradeoff point = one decision, two attributes, opposite directions
- fitness functions verify the chosen attribute continuously
basics
~20 sTurn vague goals into concrete scenarios ('traffic grows tenfold in a year', 'a payment provider fails'), score each candidate style against them, and note where a choice helps one quality but hurts another. Record the winner and the trade-offs.
solid answer
~50 sUse scenario-based evaluation, of which ATAM (Architecture Tradeoff Analysis Method) is the canonical form. Build a utility tree: quality attributes branch into concrete, measurable scenarios with a stimulus, an environment and a response measure — 'peak traffic rises tenfold within 12 months; p99 stays under 300 ms'. Prioritise each scenario by business importance and technical difficulty. Walk the top scenarios against each candidate style and record risks, non-risks, sensitivity points (a decision that strongly affects one attribute) and tradeoff points (a decision that affects two or more attributes in opposite directions — these are the real architectural decisions). Complement this with lightweight techniques: spikes and prototypes on the riskiest assumption, back-of-envelope capacity maths, and pre-mortems. Then encode the winning attributes as automated fitness functions — latency budgets in load tests, dependency rules in architecture tests — so the property is continuously verified rather than assumed. Output: a decision plus an explicit risk and trade-off list, captured in an ADR.
go deeper
Say you turn vague goals into concrete situations with numbers and check each option against them, and that you write down what each option is bad at, not just what it is good at.
Name scenario-based evaluation, describe stimulus, environment and response measure, and mention comparing at least two candidates plus prototyping the riskiest assumption.
Reference ATAM and the utility tree, distinguish sensitivity from tradeoff points, and connect the result to fitness functions that keep the property verified over time.
Match evaluation depth to reversibility and blast radius, run it with cross-functional stakeholders, track assumptions behind non-risks as live tripwires, and show how the outputs feed governance and roadmap funding.
### Why rigour is needed Style debates default to preference because the arguments are about the future and nobody can falsify them in a meeting. The fix is to convert abstract attributes into *scenarios* — statements concrete enough that two people can disagree and then check. ### Scenarios and the utility tree A well-formed scenario has three parts: a **stimulus** (what happens), an **environment** (under what conditions), and a **response measure** (the observable threshold). "The system is scalable" is useless; "peak order rate rises from 500 to 5,000 per second over 12 months while p99 checkout latency stays under 300 ms" is testable. A **utility tree** organises these: the root is overall utility, branches are quality attributes (performance, availability, modifiability, security, deployability), leaves are scenarios, each tagged with two priorities — business importance and technical difficulty. High/high leaves are where evaluation effort belongs. ### ATAM in brief The Architecture Tradeoff Analysis Method (from the Carnegie Mellon Software Engineering Institute) is a structured workshop: present the business drivers, present the candidate architecture, identify its architectural approaches, build the utility tree with stakeholders, analyse the top scenarios against each approach, brainstorm and re-prioritise scenarios with a wider group, analyse again, and present results. Its four output categories are the part worth memorising: - **Risks** — decisions that may not meet a target ("a single shared datastore may not sustain the 5,000/s write target"). - **Non-risks** — decisions confirmed safe *given stated assumptions*; recording the assumption matters, because if it changes the non-risk becomes a risk. - **Sensitivity points** — a decision that strongly moves one attribute (replica count drives availability). - **Tradeoff points** — a decision that moves two or more attributes in opposite directions (synchronous replication raises consistency and durability while lowering write latency and availability). These are the genuine architectural decisions; everything else is detail. Full ATAM is heavyweight. Most teams run a compressed version: half a day with the top five to eight scenarios, using the same four output categories. ### Complementary techniques - **Back-of-envelope maths.** Estimate request rate, payload size, storage growth and fan-out before choosing. Many debates collapse once someone computes that the whole dataset fits in memory on one machine. - **Spikes and prototypes.** Time-boxed experiments against the single riskiest assumption — not a full build, just enough to falsify it. - **Pre-mortem.** Imagine it is 18 months later and the architecture failed; have everyone write down why. Surfaces risks that optimism suppresses. - **Decision matrix.** Candidates as rows, prioritised attributes as columns, honest scores. Its value is forcing explicit reasoning; beware fake precision from arbitrary weights. - **Fitness functions.** Automated, continuously executed checks that a chosen characteristic still holds: a load test asserting a latency budget, an architecture test asserting module dependency rules, a chaos experiment asserting failover time, a budget on startup time or bundle size. This is what converts "we chose this for scalability" from a claim into a verified property, and it is the core idea of evolutionary architecture. ### Pitfalls Evaluating one candidate in isolation produces confirmation, not comparison — always score at least two. Scenarios written only by architects miss operational and business realities, so include operations, security and product people. Analysis can also become its own delay: if a decision is cheap to reverse, prototype instead of analysing. And record the assumptions behind every non-risk, since most architectures fail when an assumption silently expires rather than when a known risk fires.
- What is the difference between a sensitivity point and a tradeoff point in an architecture evaluation?A sensitivity point is a decision that strongly affects a single quality attribute — for example replica count driving availability. A tradeoff point is a decision that affects two or more attributes in opposing directions — for example synchronous cross-region replication improving durability and consistency while worsening write latency and availability. Tradeoff points are where the real architectural judgement lies.
- How do fitness functions relate to the style you selected?They make the selected characteristics continuously verifiable: a load test asserting the p99 budget, an architecture test asserting that no module reaches into another's internals, a chaos experiment asserting failover within a target time. Without them a system silently drifts away from the properties it was chosen for.
Choosing a building design by running fire drills, load calculations and storm simulations on the blueprints, rather than by asking which facade the committee likes best.
saying these in an interview costs you the question
- Scoring only the preferred candidate, so the evaluation confirms rather than compares
- Writing scenarios as adjectives ('must be scalable') with no measurable response
- Producing a weighted decision matrix with invented weights and treating the total as objective truth
- Recording non-risks without recording the assumptions they depend on
- Running a heavyweight evaluation on a decision that is cheap to reverse, when a spike would settle it faster