skip to content

An architecture evaluation surfaced 30 risks and you can fund maybe four. How do you prioritise, and how do you present the result so the business acts on it?

level: principalimportance: should knowfreq 22%

answer

  1. 30 risks → 3-5 themes
  2. theme → endangered business driver + exposure
  3. CBAM: utility vs cost, rank by ratio
  4. tail risk and deadlines override ratios
  5. close loop: ADRs, owners, SLOs, fitness functions

basics

~20 s

Group related risks into a few themes, tie each theme to the business goal it endangers, estimate the value of fixing it and the cost, and present ranked themes with expected loss and cost — not a list of thirty technical items.

solid answer

~50 s

Thirty individual risks are unreadable and unfundable. First, roll them into three to five **risk themes** ("no failure isolation in the data tier", "no operational visibility", "tenant isolation depends on application code"), because themes reveal systemic causes and a single fix often clears several risks. Second, map each theme to the **business driver** it endangers — revenue during peak, compliance certification, time-to-market, cost per transaction — and quantify exposure roughly: probability × impact, or expected downtime minutes × revenue per minute. Third, apply CBAM-style thinking: for each candidate response, estimate the utility gained (how much closer the response measure gets to the desired one) against implementation and operational cost, then rank by benefit/cost, sanity-checking against risk appetite — an existential compliance risk outranks a better ratio. Present a one-page ranked table with owners, then convert the accepted items into ADRs, backlog epics with dates, and monitored SLOs/fitness functions so progress is visible without another review.

go deeper

for a junior

Say you group risks, connect each to what the business cares about, and fix the highest-impact ones first.

for a middle

Add explicit exposure estimation (probability × impact), cost estimates, and owners with dates.

for a senior

Introduce CBAM-style utility-versus-cost ranking, non-linear utility curves, and converting accepted measures into SLOs and automated checks.

for a principal

Frame the whole economics: risk themes tied to drivers, tail risk and option value overriding ratios, explicit recorded acceptance, sequencing against capacity and deadlines, and recurring cross-project themes escalated as platform investment.

## Why raw risk lists fail A review that ends with thirty bullet points fails predictably: the audience cannot tell a fatal issue from a nit, the items are phrased in engineering terms with no link to money or obligation, no single item is big enough to fund, and nothing has an owner. Prioritisation is not an afterthought of evaluation — it is the step that makes evaluation worth doing. ## Step 1 — Roll risks into themes Cluster by **underlying cause**, not by component: - "No failure isolation in the data tier" may absorb five risks (shared database, no connection-pool partitioning, no bulkheads, no read-replica routing, no circuit breakers). - "No operational visibility" absorbs missing tracing, no SLOs, alerting on symptoms not on burn rate, no error budget. - "Tenant isolation enforced only in application code" absorbs several security risks at once. Themes matter because (a) one architectural change often clears many risks, (b) leadership can hold a theme in their head, and (c) themes expose systemic weaknesses — recurring themes across projects point at a platform or capability gap, not a project mistake. ## Step 2 — Attach each theme to a business driver For every theme, complete: *"If we do nothing, X happens, which endangers Y."* > "Any single tenant's runaway query can exhaust the shared pool and stall checkout for everyone. Historical incidents suggest two such events per year, ~40 minutes each, at roughly € X per minute of lost checkout — plus the SLA credits in the top-20 customer contracts." Drivers are usually: revenue protection, cost, compliance/certification obligations, time-to-market, customer retention, and key-person/operational risk. Rough numbers beat none — an order of magnitude with stated assumptions is decision-grade; false precision is not. ## Step 3 — Value the responses (CBAM-style) **CBAM** (Cost Benefit Analysis Method) is ATAM's economic follow-on. For each theme, list candidate architectural responses (do nothing; tactical mitigation; structural fix) and for each estimate: - the **response measure achieved** (e.g. failover from 20 min to 60 s), - **utility** — how much stakeholders value moving from the current measure to that one (often elicited on a 0–100 scale per scenario, weighted by scenario priority), - **cost** — build effort plus ongoing operational cost, - then rank by benefit/cost ratio. Utility is rarely linear: going from 20 minutes to 5 minutes of failover may be worth far more than from 5 to 1, and past a knee the extra investment buys little. That knee is where you should stop. ## Step 4 — Adjust for things ratios don't capture - **Risk appetite / tail risk** — a low-probability, company-ending outcome (data breach, losing certification) outranks its ratio. - **Option value and sequencing** — some fixes unlock others cheaply; some cannot be retrofitted later at any price (schema/tenancy decisions). - **Reversibility** — prefer cheap reversible mitigations while information is scarce. - **Team capacity and cognitive load** — four concurrent structural changes may be less deliverable than two. - **Deadline coupling** — a fix that must land before a certification audit has a hard date, not a ranking. ## Step 5 — Present so decisions actually happen One page, ranked, in the business's language: | Theme | Endangered driver | Exposure if we do nothing | Proposed response | Cost | Owner / date | |---|---|---|---|---|---| Rules that make it land: lead with exposure, not technology; show the assumptions behind every number; always include a "do nothing" row so acceptance is a visible choice; name an accountable owner and a date per accepted item; and state explicitly which risks are being **consciously accepted** — recorded acceptance is a legitimate, and often correct, outcome. ## Step 6 — Institutionalise, don't re-review - Accepted decisions → **ADRs** with rationale and consequences. - Accepted work → backlog epics with dates and owners, visible on the same board as feature work. - Response measures → **SLOs** and **fitness functions** (automated architectural tests, performance budgets, dependency rules) so progress and regression are observed continuously instead of via the next review. - Accepted risks and non-risk assumptions → a register with a periodic re-validation date, because assumptions expire. ## Failure modes Presenting all thirty items ranked by engineering severity; quantifying nothing so everything looks equal; letting the highest-ratio item win when the tail risk is existential; funding work with no owner; and running the next review from scratch because nothing from the last one was tracked.

  • How do you quantify exposure when you have no data?
    Use ranges and state assumptions: incident frequency from past postmortems, minutes of impact from previous outages, revenue per minute from finance. Present it as an order of magnitude with the assumptions visible so stakeholders can challenge the inputs rather than the conclusion — a defensible estimate beats an unquantified adjective.
  • Is deciding not to fix a risk a failure of the review?
    No. Conscious, recorded acceptance with a named owner and a re-validation date is a legitimate outcome and often the right one. The failure mode is unconscious acceptance — a risk nobody decided on, nobody owns, and nobody will revisit.
  • How do you stop the next review from re-discovering the same risks?
    Convert accepted response measures into SLOs and automated fitness functions, record decisions as ADRs, keep an owned risk register with re-validation dates, and fold recurring rules into platform defaults so the paved road makes the failure mode unreachable.

A doctor doesn't hand you thirty lab values — they give you two diagnoses, what each costs you if untreated, the treatment options with side effects, and which one to start Monday.

saying these in an interview costs you the question

  • Presenting a flat list of thirty technical risks to leadership
  • Ranking purely by engineering severity with no business driver attached
  • Refusing to estimate because the numbers are uncertain
  • Letting benefit/cost ratios override an existential tail risk
  • Treating conscious risk acceptance as a failure rather than a recorded decision
  • Funding items with no named owner, date, or measure of done

context