skip to content

In architecture evaluation, what is a "quality attribute scenario", and why is a goal like "the system must be scalable" not usable as an evaluation criterion?

level: juniorimportance: must knowfreq 40%

answer

  1. six parts: source, stimulus, artifact, environment, response, measure
  2. adjective → number
  3. utility tree: importance × difficulty
  4. growth + exploratory scenarios
  5. modifiability measured in effort, not ms

basics

~20 s

A quality attribute scenario is a concrete, testable sentence: some source sends a stimulus to the system in a given environment, and the system responds with a measurable result. "Scalable" has no measure, so nobody can agree whether a design meets it.

solid answer

~50 s

Evaluation compares a design against required quality attributes (performance, availability, modifiability, security). Bare adjectives like "scalable" are untestable: two reviewers can stare at the same design and disagree forever. A quality attribute scenario replaces the adjective with six parts: source of stimulus, stimulus, artifact, environment, response, response measure. Example: "During a peak-sales window (environment) 10,000 concurrent shoppers (source) submit orders (stimulus) to the order service (artifact); orders are accepted and persisted (response) with p99 latency under 500 ms and zero lost orders (response measure)." Now the design can be argued against evidence: which components sit on that path, what queues/caches/replicas exist, what measurements or models say. Scenarios are elicited from stakeholders, organised into a utility tree grouped by quality attribute, and prioritised by business importance and technical difficulty, because you can only analyse a handful deeply.

go deeper

for a junior

Name the six parts and give one concrete example with a number in it.

for a middle

Add the utility tree, the importance × difficulty prioritisation, and distinguish use-case / growth / exploratory scenarios.

for a senior

Show scenarios driving analysis: which components lie on the path, which decisions are sensitive, and how a measure maps onto an SLO or a load test you'd actually run.

for a principal

Tie scenarios to business drivers and money, keep them alive as fitness functions/SLOs, and use conflicting scenarios to force explicit trade-off decisions with owners.

## Why adjectives fail **Architecture evaluation** = judging whether a system's structure (its components, the connections between them, and the decisions behind them) will deliver the properties the business needs. Those properties are **quality attributes** (also "non-functional requirements"): performance, availability, modifiability, security, testability, usability, cost. Stated as adjectives they are useless for review: - unmeasurable: what number makes a system "scalable"? - context-free: scalable in read traffic? in data volume? in number of tenants? - unfalsifiable: no design can be shown to fail them, so a review produces opinions, not findings. ## The six-part scenario The standard form (from the SEI's work on ATAM/quality attribute workshops) has six parts: | Part | Meaning | Example | |---|---|---| | **Source of stimulus** | who/what triggers it | 10,000 shoppers; a developer; an attacker; a failing disk | | **Stimulus** | the event arriving | submit orders; add a payment provider; SQL injection attempt; node dies | | **Artifact** | what part is stimulated | order service; whole system; auth module | | **Environment** | the state at that moment | peak sale; normal load; degraded mode; during a deploy | | **Response** | what the system does | accepts and persists; change is made and shipped; request rejected and logged | | **Response measure** | the number that decides pass/fail | p99 < 500 ms, 0 lost orders; ≤ 3 person-days, no change to other modules; failover < 30 s | A scenario is *not* a use case: use cases capture function; scenarios capture how well, under what conditions. ## Kinds of scenarios - **Use-case scenarios** — normal expected operation. - **Growth scenarios** — anticipated change ("traffic 5x in a year", "add a second payment provider"). - **Exploratory / stress scenarios** — deliberately extreme, used to find where the design breaks ("the primary region disappears", "the schema owner leaves"). Modifiability scenarios are measured in change effort (files touched, modules affected, days, whether a deploy of other services is needed), not milliseconds — a common blind spot. ## The utility tree Scenarios are collected in a **utility tree**: root "utility", children = quality attributes, then refinements (performance → latency, throughput), leaves = concrete scenarios. Each leaf is scored twice, usually High/Medium/Low: **business importance** and **technical difficulty/risk**. (H,H) leaves are analysed first. This is the mechanism that keeps evaluation finite — a workshop may collect 40 scenarios and deeply analyse 6–10. ## Where the measures come from Good measures are ones you could later automate: SLO dashboards, load-test results, chaos-experiment outcomes, or a tracked metric like "time to add a new report type". If nobody can say how the number would be obtained, the scenario is still too vague. ## Edge cases and traps - Two scenarios can conflict (encrypt everything vs p99 latency) — that conflict *is* the finding; it marks a trade-off point. - A response measure of "as fast as possible" is not a measure. - Scenarios owned only by architects are worthless; stakeholder authorship is what makes the priorities real. - Scenarios should be revisited: last year's peak is this year's Tuesday.

  • Give a modifiability scenario with a real response measure.
    "A developer (source) must add a new payment provider (stimulus) to the checkout module (artifact) during normal development (environment); the change is made and deployed (response) touching only the payments module and taking ≤ 3 person-days, with no change to the order or ledger services (measure)."
  • Who should write the scenarios?
    Stakeholders — product, ops, security, support, and developers — facilitated by the architect. Their prioritisation votes are what give the utility tree authority; an architect-only list just re-encodes the architect's assumptions.

"Be healthy" is not a medical diagnosis. "Resting heart rate under 70 while climbing two flights of stairs" is — same person, but now there is a test you can pass or fail.

saying these in an interview costs you the question

  • Treating scenarios as use cases (function) instead of quality-under-conditions
  • Response measures like "fast", "highly available", "easy to change"
  • Only writing performance scenarios and skipping modifiability/availability/security
  • Collecting 40 scenarios and trying to analyse all of them equally
  • Assuming scenarios are written once and never revised

context