skip to content

Evaluating and Evolving Architecture

Architecture is not finished when it ships, so this covers evaluating it, measuring it, and letting it change safely: review methods, fitness functions, trade-off analysis, coupling metrics, and managing architectural risk and debt.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

29

In Robert C. Martin's component/package metrics, what do afferent coupling (Ca), efferent coupling (Ce), and instability I = Ce / (Ca + Ce) measure, and what do I = 0 and I = 1 say about a component?

level: juniorimportance: must knowfreq 55%

answer

  1. A = Arriving in, E = Exiting out
  2. I = Ce / (Ca + Ce), range 0..1
  3. I=0 stable = hard to change
  4. Depend toward lower I (SDP)
  5. Ratio hides magnitude — read Ca, Ce too

basics

~20 s

Ca counts things outside the component that depend on it (incoming arrows). Ce counts things it depends on (outgoing arrows). I = Ce/(Ca+Ce) runs 0 to 1: 0 = maximally stable (everyone depends on it, it depends on nobody), 1 = maximally unstable.

solid answer

~50 s

Ca (afferent, "arriving") is the number of external elements that depend on the component; Ce (efferent, "exiting") is the number of external elements it depends on. Instability I = Ce / (Ca + Ce) normalizes the ratio to [0, 1]. I = 0 means only incoming dependencies: nothing can force this component to change from outside, but changing it forces many others to change — it is stable, hence expensive to modify. I = 1 means only outgoing dependencies: nobody depends on it, so it is cheap to change — it is unstable/responsible-to-change. Neither extreme is inherently good. The Stable Dependencies Principle says dependencies should point in the direction of increasing stability (decreasing I): a component should only depend on components at least as stable as itself. Volatile leaf components (UI, main/composition root) should sit high in I; shared kernels and core policies low.

code

text · 8 lines
text
Dependency graph:            Ca  Ce   I = Ce/(Ca+Ce)
  Web  -> Billing            0   2    1.00   (volatile leaf)
  Jobs -> Billing            0   1    1.00
  Billing -> Money           2   1    0.33
  Money  (depends on nothing)2   0    0.00   (stable core)

Arrows run 1.00 -> 0.33 -> 0.00 : decreasing I, so SDP holds.
An arrow Money -> Web (0.00 -> 1.00) would violate it.

go deeper

for a junior

Define Ca and Ce correctly (in vs out), state the formula, and say what the two extremes mean.

for a middle

Add the Stable Dependencies Principle and a concrete example of an arrow that violates it, plus dependency inversion as the fix.

for a senior

Discuss counting conventions, the ratio-hides-magnitude trap, cycles, and where static metrics are blind (runtime/data coupling).

for a principal

Frame it as a governance tool: encode the rule as an automated fitness function in CI, track trends rather than absolute values, and be explicit that the metric is a proxy for change cost, not the goal.

### The vocabulary A **component** here means a deployable/releasable grouping of code — a package, module, assembly, library, or service. A **dependency** means "element X refers to element Y" (calls it, imports it, subclasses it, uses its type). Draw the components as nodes and the dependencies as arrows. - **Afferent coupling (Ca)** — *afferent = carrying toward*. The number of elements **outside** the component that depend on classes **inside** it. Arrows pointing **in**. Sometimes called "fan-in" or "incoming dependencies". - **Efferent coupling (Ce)** — *efferent = carrying away*. The number of elements **outside** the component that classes inside it depend on. Arrows pointing **out**. "Fan-out", "outgoing dependencies". Mnemonic: **A** for **A**rriving, **E** for **E**xiting. ### Instability ``` I = Ce / (Ca + Ce) (defined as 0 by convention when Ca = Ce = 0) ``` It is a ratio, so it is dimensionless and lives in [0, 1]. - **I = 0** — purely afferent. Many depend on it; it depends on nothing. Changing it ripples outward into everything, so in practice you *don't* change it. That is what "stable" means here: **hard to change**, not "bug-free" and not "good". - **I = 1** — purely efferent. It depends on others; nobody depends on it. You can rewrite it freely — nothing breaks downstream. "Unstable" = **easy to change**, not "flaky". - **I ≈ 0.5** — balanced; the component both serves and consumes. Worked example: component `Billing` is imported by `Web`, `Reporting`, `Jobs` (Ca = 3) and imports `Persistence`, `Money` (Ce = 2). I = 2 / (3 + 2) = 0.4 — moderately stable. ### Why anyone cares: the Stable Dependencies Principle (SDP) > *Depend in the direction of stability.* A component should only depend on components that are **more stable than itself** (lower I). The reason is change propagation. If a volatile component (high I) is depended on by a stable one (low I), every churn in the volatile piece forces changes into the piece that many others rely on — instability is "transmitted" into the core. Concretely: a UI module (I near 1) depending on a domain model (I near 0) is healthy. A domain model importing a UI helper is an SDP violation and shows up as an arrow pointing from low I to high I. When you must depend on something volatile, insert an **abstraction you own** (an interface in your stable component that the volatile component implements) — dependency inversion turns the arrow around. ### Counting rules and caveats - Counting granularity varies by tool: some count distinct **classes/types**, others distinct **packages/components**. Comparisons are only meaningful within one tool and one counting convention. - **I is a ratio, so it hides magnitude.** Ca=1/Ce=1 and Ca=100/Ce=100 both give I = 0.5. Always read Ca and Ce alongside I. - The metric is **static**: it sees compile-time references. It cannot see runtime coupling created by reflection, dependency injection by string name, plugins loaded dynamically, shared database tables, message topics, or shared file formats. Two services with zero static coupling can be tightly coupled through a shared schema. - I says nothing about **whether the dependencies are appropriate** — a component with Ce = 40 has a design smell (too many responsibilities) even if I looks fine. - Metrics are **signals, not verdicts**. Their best use is as a trend or as an automated architecture fitness function ("no arrow may point from I ≤ 0.2 to I ≥ 0.8") rather than a score to optimize. ### How this is used in practice Tools (JDepend, NDepend, Structure101, ArchUnit-style rules, `jdeps`, dependency-cruiser, import-linter) compute Ca/Ce/I per package and let you fail the build on violations. The typical findings: a "utils" grab-bag with huge Ca *and* huge Ce (everything depends on it and it depends on everything — a change amplifier), and cycles, where instability becomes meaningless because every component in the cycle must be released together.

  • Is a component with I = 1 a problem?
    No. I = 1 is exactly right for a leaf that nothing should depend on — the main/composition root, the UI shell, a CLI entry point. It is only a problem if something stable starts depending on it.
  • What do you do when a stable component genuinely needs a volatile one?
    Apply dependency inversion: define the abstraction inside the stable component and have the volatile component implement it. The source-code arrow now points from volatile to stable while the runtime call still flows outward.
  • Why does instability become meaningless inside a dependency cycle?
    Components in a cycle must be built and released together, so none of them is independently changeable; the ratio no longer predicts change cost. Break the cycle first (extract a shared abstraction or apply dependency inversion), then read the metric.

Think of a building. The foundation has huge afferent coupling — every floor rests on it — and no efferent coupling; it is 'stable' in the sense that you cannot renovate it without evacuating the building. The paint on a top-floor wall is 'unstable': nothing rests on it, so you can repaint any weekend. You want the load path to run downward, from paint to foundation, never the reverse.

saying these in an interview costs you the question

  • Reading 'stable' as 'reliable/bug-free' and 'unstable' as 'flaky' — the terms only mean hard/easy to change
  • Swapping the definitions: afferent is incoming, efferent is outgoing
  • Treating I = 0 as the goal for every component; a system of all-stable components cannot evolve
  • Optimizing I without looking at raw Ca/Ce — a ratio can't distinguish 1-and-1 from 100-and-100
  • Assuming low static coupling means low coupling, ignoring shared databases, message schemas, and reflection

context

open as a page

In architecture evaluation, what is a "quality attribute scenario", and why is a goal like "the system must be scalable" not usable as an evaluation criterion?

level: juniorimportance: must knowfreq 40%

basics

~20 s

A quality attribute scenario is a concrete, testable sentence: some source sends a stimulus to the system in a given environment, and the system responds with a measurable result. "Scalable" has no measure, so nobody can agree whether a design meets it.

open as a page

In evolutionary architecture, what is an architectural fitness function, and how does it differ from an ordinary unit test of business logic?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A fitness function is an automated, objective check that the system still meets a required architectural quality — response time, allowed dependencies, security rules. A unit test checks business behaviour; a fitness function guards a structural or quality property.

open as a page

What is technical debt, and what do the terms "principal" and "interest" mean when applied to it?

level: juniorimportance: must knowfreq 82%

basics

~20 s

Technical debt is a shortcut in design or code that makes today faster but future changes slower. The principal is the work needed to fix the shortcut; the interest is the extra effort every change costs while it stays unfixed.

open as a page

Architects often say "there are no right answers in architecture, only trade-offs". What does that mean in practice, and can you name two quality attributes that pull against each other?

level: juniorimportance: must knowfreq 72%

basics

~10 s

It means improving one desirable property usually costs another. Example: caching copies of data makes reads fast but copies can be stale; adding redundant servers raises availability but raises cost and operational complexity.

open as a page

Define abstractness (A) for a component, and explain the "main sequence" A + I = 1, the distance metric D = |A + I - 1|, and what the zone of pain and zone of uselessness are.

level: middleimportance: must knowfreq 35%

basics

~20 s

A = abstract types / total types in a component (0 = all concrete, 1 = all interfaces). The main sequence is the line A + I = 1: stable components should be abstract, unstable ones concrete. D = |A + I - 1| is how far you sit from that line — smaller is better.

open as a page

What is ATAM (the Architecture Tradeoff Analysis Method), and what does running one actually produce?

level: middleimportance: must knowfreq 38%

basics

~20 s

ATAM is a structured workshop where stakeholders write measurable quality scenarios, the architect explains how the design handles the top ones, and the group records risks, non-risks, sensitivity points and trade-off points — not a pass/fail score.

open as a page

Fitness functions in evolutionary architecture are commonly classified as atomic vs holistic and triggered vs continuous. Explain each of these four categories with an example, and why the distinction matters.

level: middleimportance: must knowfreq 45%

basics

~20 s

Atomic checks one characteristic in isolation (a layering rule); holistic checks several interacting at once (security plus performance under load). Triggered runs on demand — a build or deploy; continuous runs constantly against the live system, e.g. monitoring latency.

open as a page

How would you enforce architectural structure rules — for example 'the domain layer must not depend on the persistence layer' and 'there are no cyclic dependencies between modules' — as automated checks in a build, and what are the practical pitfalls of doing so?

level: middleimportance: must knowfreq 50%

basics

~20 s

Add a test that reads the code's dependency graph — from source, bytecode, or import statements — and asserts rules like 'domain must not reference persistence' and 'no cycles'. Run it in the build so violations fail like any other test.

open as a page

In a distributed data store, what does the CAP theorem actually force you to choose between, and how does PACELC extend that framing?

level: middleimportance: must knowfreq 68%

basics

~20 s

CAP says that when the network splits (a partition), a distributed store must choose: keep answering with possibly-stale or conflicting data (availability), or refuse some requests to stay correct (consistency). PACELC adds: even with no partition, you still trade latency against consistency.

open as a page

Your organisation cannot run a multi-day formal evaluation for every significant decision. How do you design a lightweight, continuous architecture evaluation process (RFCs, decision records, an architecture review board or advice process) that catches real risk without becoming a bottleneck?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Only review decisions that are expensive to reverse. Use a written proposal template (context, options, trade-offs, chosen option), require the author to consult affected experts, timebox feedback, record the outcome as a decision record, and automate recurring checks instead of re-reviewing them.

open as a page

Evolutionary architecture is defined as supporting guided, incremental change across multiple dimensions. Which concrete practices let a team change architecture incrementally while keeping confidence that nothing broke?

level: seniorimportance: must knowfreq 35%

basics

~20 s

Make small reversible changes behind a deployment pipeline that runs fitness functions on every commit; change contracts additively (add new, migrate, remove old); replace systems piece by piece rather than all at once; and use feature flags so releasing is separate from deploying.

open as a page

Distinguish architectural erosion from architectural drift, explain how a system becomes a "big ball of mud", and describe concrete mechanisms for preventing both.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Erosion is when code violates the intended architecture — for example a layer calling something it is not allowed to call. Drift is when things are added that the architecture never covered, so the design loses coherence without any explicit rule being broken. Both, left unchecked, end in a structureless "big ball of mud".

open as a page

How do you decide what architectural debt to remediate and when, so that it competes fairly with feature work rather than being permanently deferred?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Prioritise by how much the debt actually costs: fix the places you change often and that hurt every time, not the ugliest code. Estimate the fix cost and the ongoing cost of leaving it, tie remediation to upcoming features, and put it on the same backlog as features so it is compared, not deferred.

open as a page

What are second-order consequences of an architectural decision, and how do you surface them before you commit?

level: seniorimportance: must knowfreq 48%

basics

~20 s

First-order effects are the intended ones; second-order are what those effects then cause — often later and elsewhere. Splitting a service speeds deploys (first order) but creates distributed transactions, on-call load and cross-team coordination (second order).

open as a page

What is connascence in software design, and what distinguishes the static forms (name, type, meaning, algorithm, position) from the dynamic forms (execution order, timing, value, identity)?

level: middleimportance: should knowfreq 28%

basics

~20 s

Connascence (Meilir Page-Jones) means two pieces of code are 'born together': if one changes, the other must change too, or the system breaks. Static forms are visible by reading the code (names, types, order of parameters); dynamic forms only appear at run time (call order, timing, values, object identity).

open as a page

How is cohesion measured in practice — what does the LCOM (Lack of Cohesion of Methods) family of metrics compute, and how does it map onto the classic cohesion spectrum from coincidental to functional?

level: middleimportance: should knowfreq 20%

basics

~20 s

LCOM looks at which methods of a class touch which fields. If methods share fields, the class is cohesive; if they form separate groups touching disjoint fields, the class is really several classes glued together. High LCOM = low cohesion, a hint to split.

open as a page

Architecture reviews distinguish four kinds of finding: risks, non-risks, sensitivity points and trade-off points. Define each and explain why recording non-risks is worth the effort.

level: middleimportance: should knowfreq 28%

basics

~20 s

A risk is a decision that may stop a goal being met; a non-risk is a decision that's fine given stated assumptions; a sensitivity point is a decision that strongly swings one quality attribute; a trade-off point swings two the opposite way. Non-risks record the assumptions that could later expire.

open as a page

Explain Martin Fowler's technical debt quadrant (deliberate vs inadvertent, prudent vs reckless) and how the classification changes your response.

level: middleimportance: should knowfreq 58%

basics

~20 s

Fowler classifies debt on two axes: was it taken on purpose (deliberate) or by accident (inadvertent), and was the decision sensible (prudent) or careless (reckless). The four combinations call for different responses — from scheduled repayment to coaching or changing how the team works.

open as a page

What is risk-storming as an architecture practice, how is a session actually run, and why is the first step done silently and individually?

level: middleimportance: should knowfreq 42%

basics

~20 s

Risk-storming is a group technique where people mark risks directly on the architecture diagrams. Everyone first writes risks alone and silently, then all notes are placed on the diagram, discussed, and ranked by probability and impact so the riskiest spots are visible and get mitigation owners.

open as a page

How would you build a weighted decision matrix to choose between competing architecture options, and what are its main failure modes?

level: middleimportance: should knowfreq 42%

basics

~20 s

List the options as columns and the criteria (quality attributes) as rows, give each criterion a weight agreed with stakeholders, score every option per criterion, then sum weight × score. Its value is the discussion it forces, not the number it prints.

open as a page

Connascence is evaluated along three dimensions — strength, locality, and degree. Define each, and state the refactoring rules that follow from them.

level: seniorimportance: should knowfreq 22%

basics

~20 s

Strength = how hard the coupling is to detect and change (name is weak, timing/identity strong). Locality = how far apart the coupled parts are (same function vs different services). Degree = how many places share it. Rules: reduce strength, keep strong forms local, and shrink degree.

open as a page

An architecture evaluation surfaced 30 risks and you can fund maybe four. How do you prioritise, and how do you present the result so the business acts on it?

level: principalimportance: should knowfreq 22%

basics

~20 s

Group related risks into a few themes, tie each theme to the business goal it endangers, estimate the value of fixing it and the cost, and present ranked themes with expected loss and cost — not a list of thirty technical items.

open as a page

As the architect for a large system, how would you derive a fitness-function suite from the system's architectural characteristics, and how do you govern it over time — including what to do when two fitness functions conflict?

level: principalimportance: should knowfreq 20%

basics

~20 s

Pick the few characteristics that really drive the design, make each measurable with a threshold, and automate it in the pipeline. Review the suite as business needs change. When two conflict, make the trade-off explicit and decide by business priority — do not silently weaken one.

open as a page

Name several architecture smells and the structural metrics used to detect them, and explain the limits of managing architecture by such metrics.

level: principalimportance: should knowfreq 38%

basics

~20 s

Architecture smells are structural warning signs above the code level: dependency cycles between modules, a god component everything depends on, a shared 'utils' hub, and features scattered across many modules. Metrics such as coupling counts, instability and change coupling help spot them, but they are indicators, not verdicts.

open as a page

How do you avoid over-optimising an architecture for a single quality attribute — for example designing for hyperscale, or for maximum flexibility, on day one?

level: principalimportance: should knowfreq 38%

basics

~20 s

Tie every design choice to a measured or agreed requirement. Ask what evidence says this attribute matters at this scale, what it costs the other attributes, and whether the decision is easy to reverse later. Optimise last-responsibly, not first.

open as a page

How does SAAM (Software Architecture Analysis Method) differ from ATAM (Architecture Tradeoff Analysis Method), and when would you choose the simpler one?

level: seniorimportance: nice to knowfreq 15%

basics

~20 s

SAAM came first and is simpler: stakeholders write change scenarios, you check which components each one touches, and compare candidate designs — mainly for modifiability. ATAM extends it to many quality attributes at once and adds explicit trade-off and sensitivity analysis.

open as a page

In architecture evaluation, what is the difference between a 'sensitivity point' and a 'trade-off point', and how do you find them in a design?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

A sensitivity point is a design choice where one quality attribute changes sharply if you tweak it. A trade-off point is a choice that is a sensitivity point for two or more attributes at once, so tuning it helps one and hurts another.

open as a page

Static coupling metrics like Ca/Ce/I/A/D report a healthy architecture, yet a change to one service still forces coordinated releases of four others. What kinds of coupling do these metrics miss, and how would you govern coupling in a way that actually catches this?

level: principalimportance: nice to knowfreq 12%

basics

~20 s

Static metrics only see code references. They miss coupling through shared databases, shared message schemas, implicit data meanings, required call ordering, timing assumptions, and shared libraries. Measure it instead by what actually co-changes and co-deploys, and guard boundaries with automated contract tests and fitness functions.

open as a page