skip to content

Architecture Evaluation

Structured ways to review an architecture before you have built it: ATAM and SAAM scenario workshops, plus lighter-weight RFC and review-board processes. The output that matters is the list of risks, sensitivity points and trade-offs the design implies.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In architecture evaluation, what is a "quality attribute scenario", and why is a goal like "the system must be scalable" not usable as an evaluation criterion?

level: juniorimportance: must knowfreq 40%

answer

  1. six parts: source, stimulus, artifact, environment, response, measure
  2. adjective → number
  3. utility tree: importance × difficulty
  4. growth + exploratory scenarios
  5. modifiability measured in effort, not ms

basics

~20 s

A quality attribute scenario is a concrete, testable sentence: some source sends a stimulus to the system in a given environment, and the system responds with a measurable result. "Scalable" has no measure, so nobody can agree whether a design meets it.

solid answer

~50 s

Evaluation compares a design against required quality attributes (performance, availability, modifiability, security). Bare adjectives like "scalable" are untestable: two reviewers can stare at the same design and disagree forever. A quality attribute scenario replaces the adjective with six parts: source of stimulus, stimulus, artifact, environment, response, response measure. Example: "During a peak-sales window (environment) 10,000 concurrent shoppers (source) submit orders (stimulus) to the order service (artifact); orders are accepted and persisted (response) with p99 latency under 500 ms and zero lost orders (response measure)." Now the design can be argued against evidence: which components sit on that path, what queues/caches/replicas exist, what measurements or models say. Scenarios are elicited from stakeholders, organised into a utility tree grouped by quality attribute, and prioritised by business importance and technical difficulty, because you can only analyse a handful deeply.

go deeper

for a junior

Name the six parts and give one concrete example with a number in it.

for a middle

Add the utility tree, the importance × difficulty prioritisation, and distinguish use-case / growth / exploratory scenarios.

for a senior

Show scenarios driving analysis: which components lie on the path, which decisions are sensitive, and how a measure maps onto an SLO or a load test you'd actually run.

for a principal

Tie scenarios to business drivers and money, keep them alive as fitness functions/SLOs, and use conflicting scenarios to force explicit trade-off decisions with owners.

## Why adjectives fail **Architecture evaluation** = judging whether a system's structure (its components, the connections between them, and the decisions behind them) will deliver the properties the business needs. Those properties are **quality attributes** (also "non-functional requirements"): performance, availability, modifiability, security, testability, usability, cost. Stated as adjectives they are useless for review: - unmeasurable: what number makes a system "scalable"? - context-free: scalable in read traffic? in data volume? in number of tenants? - unfalsifiable: no design can be shown to fail them, so a review produces opinions, not findings. ## The six-part scenario The standard form (from the SEI's work on ATAM/quality attribute workshops) has six parts: | Part | Meaning | Example | |---|---|---| | **Source of stimulus** | who/what triggers it | 10,000 shoppers; a developer; an attacker; a failing disk | | **Stimulus** | the event arriving | submit orders; add a payment provider; SQL injection attempt; node dies | | **Artifact** | what part is stimulated | order service; whole system; auth module | | **Environment** | the state at that moment | peak sale; normal load; degraded mode; during a deploy | | **Response** | what the system does | accepts and persists; change is made and shipped; request rejected and logged | | **Response measure** | the number that decides pass/fail | p99 < 500 ms, 0 lost orders; ≤ 3 person-days, no change to other modules; failover < 30 s | A scenario is *not* a use case: use cases capture function; scenarios capture how well, under what conditions. ## Kinds of scenarios - **Use-case scenarios** — normal expected operation. - **Growth scenarios** — anticipated change ("traffic 5x in a year", "add a second payment provider"). - **Exploratory / stress scenarios** — deliberately extreme, used to find where the design breaks ("the primary region disappears", "the schema owner leaves"). Modifiability scenarios are measured in change effort (files touched, modules affected, days, whether a deploy of other services is needed), not milliseconds — a common blind spot. ## The utility tree Scenarios are collected in a **utility tree**: root "utility", children = quality attributes, then refinements (performance → latency, throughput), leaves = concrete scenarios. Each leaf is scored twice, usually High/Medium/Low: **business importance** and **technical difficulty/risk**. (H,H) leaves are analysed first. This is the mechanism that keeps evaluation finite — a workshop may collect 40 scenarios and deeply analyse 6–10. ## Where the measures come from Good measures are ones you could later automate: SLO dashboards, load-test results, chaos-experiment outcomes, or a tracked metric like "time to add a new report type". If nobody can say how the number would be obtained, the scenario is still too vague. ## Edge cases and traps - Two scenarios can conflict (encrypt everything vs p99 latency) — that conflict *is* the finding; it marks a trade-off point. - A response measure of "as fast as possible" is not a measure. - Scenarios owned only by architects are worthless; stakeholder authorship is what makes the priorities real. - Scenarios should be revisited: last year's peak is this year's Tuesday.

  • Give a modifiability scenario with a real response measure.
    "A developer (source) must add a new payment provider (stimulus) to the checkout module (artifact) during normal development (environment); the change is made and deployed (response) touching only the payments module and taking ≤ 3 person-days, with no change to the order or ledger services (measure)."
  • Who should write the scenarios?
    Stakeholders — product, ops, security, support, and developers — facilitated by the architect. Their prioritisation votes are what give the utility tree authority; an architect-only list just re-encodes the architect's assumptions.

"Be healthy" is not a medical diagnosis. "Resting heart rate under 70 while climbing two flights of stairs" is — same person, but now there is a test you can pass or fail.

saying these in an interview costs you the question

  • Treating scenarios as use cases (function) instead of quality-under-conditions
  • Response measures like "fast", "highly available", "easy to change"
  • Only writing performance scenarios and skipping modifiability/availability/security
  • Collecting 40 scenarios and trying to analyse all of them equally
  • Assuming scenarios are written once and never revised

context

open as a page

What is ATAM (the Architecture Tradeoff Analysis Method), and what does running one actually produce?

level: middleimportance: must knowfreq 38%

basics

~20 s

ATAM is a structured workshop where stakeholders write measurable quality scenarios, the architect explains how the design handles the top ones, and the group records risks, non-risks, sensitivity points and trade-off points — not a pass/fail score.

open as a page

Your organisation cannot run a multi-day formal evaluation for every significant decision. How do you design a lightweight, continuous architecture evaluation process (RFCs, decision records, an architecture review board or advice process) that catches real risk without becoming a bottleneck?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Only review decisions that are expensive to reverse. Use a written proposal template (context, options, trade-offs, chosen option), require the author to consult affected experts, timebox feedback, record the outcome as a decision record, and automate recurring checks instead of re-reviewing them.

open as a page

Architecture reviews distinguish four kinds of finding: risks, non-risks, sensitivity points and trade-off points. Define each and explain why recording non-risks is worth the effort.

level: middleimportance: should knowfreq 28%

basics

~20 s

A risk is a decision that may stop a goal being met; a non-risk is a decision that's fine given stated assumptions; a sensitivity point is a decision that strongly swings one quality attribute; a trade-off point swings two the opposite way. Non-risks record the assumptions that could later expire.

open as a page

An architecture evaluation surfaced 30 risks and you can fund maybe four. How do you prioritise, and how do you present the result so the business acts on it?

level: principalimportance: should knowfreq 22%

basics

~20 s

Group related risks into a few themes, tie each theme to the business goal it endangers, estimate the value of fixing it and the cost, and present ranked themes with expected loss and cost — not a list of thirty technical items.

open as a page

How does SAAM (Software Architecture Analysis Method) differ from ATAM (Architecture Tradeoff Analysis Method), and when would you choose the simpler one?

level: seniorimportance: nice to knowfreq 15%

basics

~20 s

SAAM came first and is simpler: stakeholders write change scenarios, you check which components each one touches, and compare candidate designs — mainly for modifiability. ATAM extends it to many quality attributes at once and adds explicit trade-off and sensitivity analysis.

open as a page