skip to content

A stakeholder writes the requirement "the system must be scalable". How do you turn that into something an architect can design and test against?

level: middleimportance: must knowfreq 55%

answer

  1. Six parts: source, stimulus, artifact, environment, response, measure
  2. Response measure = the pass/fail number
  3. Percentiles, not averages
  4. Growth + stress scenarios, not just normal
  5. Scenario → SLO / load test / fitness function

basics

~20 s

Rewrite it as a concrete scenario with numbers: who or what triggers it, under what conditions, what the system should do, and the measurable limit. E.g. "When traffic triples in 10 minutes, p99 latency stays under 800 ms with no dropped requests."

solid answer

~60 s

"Scalable" is unfalsifiable — nobody can pass or fail it. The standard fix is a **quality-attribute scenario** with six parts: **source** of stimulus (who/what), **stimulus** (the event), **artifact** (which part of the system), **environment** (the operating conditions — normal, peak, degraded, during deployment), **response** (what the system does), and **response measure** (the number that decides pass/fail). So: "During the Black Friday window (environment), incoming checkout traffic from real users (source) rises from 2k to 6k requests/second over 10 minutes (stimulus) against the checkout service (artifact); the system scales out and serves all requests (response) with p99 latency ≤ 800 ms, zero dropped orders, and cost ≤ $X/hour (measure)." This does three things: it becomes testable (load test / game day / fitness function), it exposes which structural decisions matter (statelessness, queueing, sharding key, autoscaling lag), and it forces the trade-off conversation — the same load target with 10 ms latency would demand a completely different design. I also collect scenarios for growth ×10 and for failure modes, not only happy-path peaks.

code

text · 9 lines
text
SCENARIO  QA-07  Peak checkout load          priority: business=HIGH, risk=HIGH
source      real end users
stimulus    checkout rate 2,000 -> 6,000 req/s over 10 minutes
artifact    checkout service + order store
environment Black Friday peak, no infrastructure faults
response    scale out horizontally; accept and persist every order
measure     p99 <= 800 ms; errors <= 0.1%; 0 lost orders;
            scale-out completes <= 3 min; marginal cost <= $400/h
verified by nightly load test (k6) + production SLO + Nov game day

go deeper

for a junior

Say that "scalable" needs a number and rewrite it with concrete load, latency, and conditions. Naming even three of the six parts (stimulus, response, measure) is a good junior answer.

for a middle

Produce the full six-part scenario fluently, use percentiles at a stated load, and mention that the scenario becomes a load test or SLO. Note that "scalable" must name its dimension (rate, data, tenants, connections).

for a senior

Add prioritization (business value × technical risk), the three scenario families (normal, growth, stress), and how each scenario maps to a structural decision and to an ongoing check (SLO, chaos test, fitness function).

for a principal

Own the elicitation process: run a quality-attribute workshop with the right stakeholders, tie scenarios to business goals and error budgets, keep a maintained scenario set as the architecture's regression suite, and use it to force explicit trade-off decisions rather than implicit ones.

## The problem with one-word qualities "Must be scalable / fast / secure / maintainable / highly available" are **unfalsifiable**. Two engineers can build opposite systems and both claim compliance, and no test can settle it. Worse, they hide the real question: scalable *in what dimension*, *to what limit*, *at what latency*, *at what cost*, *by when*? The repair used across the industry (formalized in the SEI's Attribute-Driven Design / ATAM literature, but you will meet it as "write it as a scenario" everywhere) is the **six-part quality-attribute scenario**. ## The six parts | Part | Question it answers | Example | |---|---|---| | **Source of stimulus** | Who or what causes the event? | Real end users; an internal batch job; an attacker; a developer; a failing disk | | **Stimulus** | What exactly happens? | Traffic rises 3× in 10 minutes; a node dies; a new payment provider must be added; 10k concurrent WebSocket connections open | | **Artifact** | Which part is stimulated? | The checkout service; the primary database; the whole system; the deployment pipeline | | **Environment** | Under what conditions? | Normal operation; peak season; degraded mode with one AZ down; during a rolling deploy; at 3 a.m. with one on-call engineer | | **Response** | What should the system do? | Autoscale and serve all requests; fail over within N seconds; shed non-essential load; log and alert; reject with a typed error | | **Response measure** | How do we know it worked? | p99 ≤ 800 ms; zero lost orders; failover < 30 s; RPO ≤ 5 s; change implemented in ≤ 3 person-days touching ≤ 2 modules; cost ≤ $X/hour | The **response measure** is the part people skip, and it is the only part that makes the scenario a requirement rather than a wish. ## Worked rewrites **"Must be scalable" →** > *Source:* real users. *Stimulus:* checkout request rate rises from 2,000 to 6,000 req/s over 10 minutes. *Artifact:* checkout service + order database. *Environment:* Black Friday peak, normal (non-degraded) infrastructure. *Response:* the service scales horizontally and continues to accept all orders. *Measure:* p99 end-to-end latency ≤ 800 ms, error rate ≤ 0.1%, no order lost, scale-out completes within 3 minutes of the threshold breach, marginal cost ≤ $400/hour. **"Must be highly available" →** > *Source:* infrastructure fault. *Stimulus:* the primary database node becomes unreachable. *Artifact:* order-write path. *Environment:* normal production, business hours. *Response:* automatic failover to the standby; in-flight writes are retried idempotently. *Measure:* writes resume within 30 s; RPO ≤ 5 s of committed data; no duplicate orders created; users see at most a retriable error. **"Must be maintainable" →** (a development-time scenario — note the source is a person, not traffic) > *Source:* a developer. *Stimulus:* a new payment provider must be integrated. *Artifact:* the payments module. *Environment:* normal development, current team. *Response:* the provider is added behind the existing provider interface. *Measure:* ≤ 3 person-days, changes confined to the payments module plus configuration, no change to callers, existing provider tests still pass unmodified. **"Must be secure" →** > *Source:* an unauthenticated external actor. *Stimulus:* credential-stuffing attempts at 1,000 logins/minute from rotating IPs. *Artifact:* the login endpoint. *Environment:* normal production. *Response:* rate-limit per account and per IP, require step-up challenge, alert. *Measure:* < 0.01% of attempts reach the password check after the first 5 per account per minute; detection within 2 minutes; no legitimate-user error rate above 0.5%. ## Which scenarios to collect 1. **Use-case / normal scenarios** — the everyday load and the everyday change. 2. **Growth scenarios** — plausible near-future change: 10× users, a second region, a new tenant model. These test whether the architecture has headroom. 3. **Exploratory / stress scenarios** — deliberately extreme: the whole region dies, a 100× traffic spike, the schema must change with zero downtime. These are not commitments; they find brittleness and hidden coupling. Collect them in a workshop with stakeholders (the SEI name is a **Quality Attribute Workshop**), prioritize each on two axes — *business importance* and *technical difficulty/risk* — and design against the high-high ones first. Ten to fifteen prioritized scenarios usually pin down an architecture. ## Turning scenarios into ongoing checks A scenario is only worth writing if something later re-checks it: - **Runtime** → SLOs and alerts (p99 latency, error budget), load tests in CI, chaos/game-day exercises for the failure scenarios. - **Development-time** → "fitness functions": architecture tests that assert module boundaries, dependency rules, bundle-size or build-time budgets — the automated stand-in for a modifiability measure. ## Common mistakes - **No measure**, or a measure with no environment ("p99 < 200 ms" — at what load? during a deploy? with a cold cache?). - **Average instead of percentile.** Averages hide the tail that users actually feel. - **Scenarios only for the happy path**, so nothing drives redundancy or degraded-mode design. - **Every scenario marked critical.** If everything is priority one, the scenario set has stopped being a prioritization tool. - **Confusing scale dimensions.** "Scalable" can mean request rate, data volume, tenant count, concurrent connections, or team size — each drives a different structure (stateless replicas vs sharding vs multi-tenancy isolation vs modularization). Name the dimension. ## Interview framing When an interviewer hands you a vague NFR, the strongest move is not to guess a number but to *ask the six questions* aloud and then propose a defensible starting number with its rationale ("today we peak at 2k/s; the marketing plan implies 3×; so 6k/s with p99 800 ms, because checkout conversion drops measurably past a second"). That demonstrates both the technique and the business grounding.

  • Your stakeholder cannot give you a number. What do you do?
    Derive one and get it confirmed. Anchor on current measurements (today's peak, today's p99), on business plans (marketing forecast, contracted SLA, seasonality), and on user-impact evidence (conversion drop past a second). Then propose "6k req/s at p99 800 ms" as a working target with its rationale written down. A wrong-but-explicit number is revisable; silence is not. You can also express it as a range with a stretch scenario.
  • Why insist on percentiles rather than an average latency target?
    Averages are dominated by the many fast requests and hide the slow tail. A service with a 60 ms average can still time out for 1 in 100 users. Users experience the tail, and in fan-out architectures a request touching 50 backends is very likely to hit at least one p99 event, so tail latency compounds. Specify p50 for the typical experience plus p99/p99.9 for the worst case, always paired with the load level.
  • How do you write a measurable scenario for modifiability, where there is no runtime metric?
    Make the stimulus a specific anticipated change and the measure a cost of change: person-days, number of modules/files touched, whether callers change, whether existing tests need editing. Then automate an approximation with fitness functions — dependency rules, layering tests, module-boundary tests — so drift is caught continuously rather than discovered at the next change.

"Be scalable" is like a fire code that says "the building should be safe". A real code says: with 400 occupants on floor 12 at night, everyone must reach the street within 6 minutes with one stairwell blocked. Only the second version tells the architect where to put stairwells — and only the second one can be drilled.

saying these in an interview costs you the question

  • Accepting "must be scalable/fast/secure" as a requirement and starting to design without asking for a measure.
  • Writing a target as an average latency instead of a percentile at a stated load.
  • Omitting the environment, so "p99 < 200 ms" is untestable (cold cache? during deploy? one AZ down?).
  • Specifying a solution ("must be stateless and horizontally scalable") instead of the target it must achieve.
  • Marking every scenario as critical priority, which destroys the prioritization the exercise exists for.
  • Writing only happy-path scenarios, so nothing in the requirement set drives redundancy or degraded modes.
  • Saying "scalable" without naming the dimension — request rate, data volume, tenants, connections, and team size need different structures.

context