skip to content

What are the limitations of using distance from the main sequence (D = |A + I − 1|) as an architectural health metric, and how would you govern it responsibly across a large codebase?

level: principalimportance: nice to knowfreq 12%

answer

  1. D covers SAP only — not cohesion, cycles, correctness
  2. No volatility term: frozen concrete high-fan-in is fine
  3. A is a gameable type count (Goodhart)
  4. Boundaries determine the numbers — cross-repo comparison is meaningless
  5. Rank + trend as a fitness function; enforce direction with arch tests

basics

~20 s

D only measures the balance between abstractness and how depended-upon a component is. It ignores volatility, cohesion, cycles, and whether abstractions are really used. Use it to rank and to watch trends, never as a hard pass/fail target.

solid answer

~50 s

D is a proxy with well-known blind spots. It ignores **volatility** — a frozen concrete component with high fan-in scores terribly and costs nothing; ignores **cohesion and cycles** (REP/CCP/CRP and the Acyclic Dependencies Principle are separate); is **gameable**, since A is a raw count of abstract types, so interfaces implemented once inside the same component raise A without inverting anything; is **boundary-dependent**, because merging or splitting modules changes Ce, Ca, and Nc at once; is **noisy for small components**; and is **language-sensitive** (Go's implicit interfaces, duck-typed languages, generated stubs). It also drops the sign, so (0,0) and (1,1) look identical. Govern it as an architectural fitness function: report the signed distance, weight it by VCS churn, exclude generated/test/published-API modules, review the top-N ranked offenders rather than gating on a threshold, and alert on trend regressions. Pair it with cycle detection and a dependency-direction rule enforced by architecture tests, which is what actually prevents decay.

go deeper

for a junior

Say that D only checks whether abstractness matches how depended-upon a component is, and that a low score isn't automatically a bug.

for a middle

List the concrete blind spots — volatility, cohesion, cycles — and note the absolute value hides which zone you're in.

for a senior

Add gameability (A is a raw type count), boundary-dependence, small-module noise, and language sensitivity; propose ranking plus churn weighting over thresholds.

for a principal

Design the governance: signed distance × churn × fan-in ranking, explicit exclusions, trend alerting, and binary enforced dependency rules in the build for the properties that actually matter — plus the framing that this is a 1990s compiled-monolith heuristic that under-serves distributed and dynamically-typed systems.

### What D actually claims `D = |A + I − 1|` (Martin's original divides by √2) measures one thing: **is a component's Abstractness proportional to its Instability?** — i.e. compliance with the Stable Abstractions Principle. Everything else about architectural health is outside its scope. Treating it as a general quality score is the root of most misuse. ### The limitations, one by one **1. It has no notion of volatility.** A component only causes pain when it must change. A frozen, concrete, universally-used component (a platform string library, a versioned wire format, generated protobuf stubs) sits at (0,0) with `D ≈ 1` and is entirely fine — Martin says so explicitly. Conversely a component with a modest D that churns weekly against ten teams may be your worst problem. **Weight D by change frequency** from version control; `D × churn × fan-in` is a far better priority signal than D alone. **2. The absolute value discards direction.** (0,0) — stable and concrete, the *zone of pain*, rigidity and ripple — and (1,1) — unstable and abstract, the *zone of uselessness*, dead weight — both give `D = 1`, yet they need opposite remedies (extract/invert vs. delete/wire-up) and have wildly different severity. Always report the **signed** `A + I − 1`. **3. A is a crude, gameable count.** Abstractness is `Na/Nc`: abstract types over total types, unweighted by size, usage, or whether anything depends *through* them. You can raise A by adding a one-implementation interface right beside its implementation in the same component — the metric improves, the coupling is unchanged. Any metric adopted as a target invites this (Goodhart's law). The thing that matters — **do dependents point at abstractions rather than concretions?** — is a dependency-direction property, and it needs a dependency rule, not a ratio. **4. It is entirely a function of where you drew the boundaries.** Merge two components and inter-component edges become internal ones: Ce and Ca both fall, A changes as Nc changes. Split one and the numbers move again. So D partly measures your packaging taxonomy. This cuts both ways — repackaging is a legitimate *fix*, but it also means cross-project D comparisons are meaningless unless boundaries are drawn on comparable principles. **5. It is silent on cohesion and cycles.** Component design has three cohesion principles (REP — Reuse/Release Equivalence, CCP — Common Closure, CRP — Common Reuse) and three coupling ones (ADP — Acyclic Dependencies, SDP — Stable Dependencies, SAP — Stable Abstractions). D covers **SAP only**. A component can sit perfectly on the line while being an incoherent grab-bag inside a dependency cycle. **6. Small components make it noisy.** With Nc = 2, A can only be 0, 0.5, or 1; with Ce+Ca = 1, I is 0 or 1. Micro-modules swing across the graph on a single type. Set a minimum size before reporting, or aggregate. **7. Language and toolchain sensitivity.** `Na` is well-defined in nominally-typed languages (Java, C#, Kotlin, C++, TypeScript) and murky elsewhere: Go's implicit interfaces, Python protocols/ABCs, Ruby/JavaScript duck typing, Rust traits with blanket impls. Generated code, annotation processors, and macro-expanded types distort both counts. Multi-language repos cannot be compared on one scale. **8. Structural, not behavioural.** D says nothing about correctness, performance, test coverage, security, operational coupling, data coupling (two services sharing a table have huge real coupling and no static edges), or runtime coupling via message contracts, reflection, and DI. Distributed systems in particular hide most of their coupling from a static analyser. ### Governing it responsibly **Treat it as an architectural fitness function, not a gate.** 1. **Report signed distance and the components' coordinates**, not a single scalar. A scatter plot of all components on the A/I square with the main sequence drawn is far more actionable than a table of D values. 2. **Rank, then review top-N.** Sort by `|signed D| × churn × fan-in` and put the top handful into the architecture review or refactoring backlog each quarter. Absolute thresholds ("all modules < 0.3") produce ceremony. 3. **Scope-exclude by design:** generated stubs/DTOs, test fixtures, published-API modules (their consumers live outside the build, so Ca reads 0 falsely), and modules deliberately declared frozen. Make the allowlist explicit and reviewed, not hidden in a config. 4. **Alert on trend, not level.** A component whose signed distance drifts −0.2 → −0.7 over three releases is decaying; that delta is the signal. Store history and chart it. 5. **Enforce the thing you actually care about with rules.** ArchUnit / dependency-cruiser / Spring Modulith verification / NDepend rules can assert *dependency direction* ("no consumer may reference `*-impl`"), *acyclicity*, and *layer boundaries*. These are binary, non-gameable, and prevent regression — the metric only *detects* it. 6. **Combine with independent signals:** cycle count, change-coupling (files that change together across commits), team ownership overlap per component, and build fan-out time. Together these tell an architectural story D cannot. 7. **Keep the human judgement.** Publish the interpretation with the number: which zone, whether it's volatile, what the remedy class is (extract/invert, split by CCP/CRP, delete, freeze, duplicate). Numbers without that framing get either ignored or blindly optimised. ### The honest summary The main sequence is a **rule of thumb from the era of compiled, statically-linked, monolithic builds**, and it remains a genuinely useful lens: it makes "is this thing hard to change and does it have a seam?" visible at a glance. But it is a heuristic derived from two ratios of type counts. Use it to *start conversations and rank work*; use enforced dependency rules and change data to *make decisions*.

  • If D is so easy to game, why compute it at all?
    Because it is nearly free, it renders the whole component graph as one picture, and the outliers it surfaces are usually real. Its value is as a screening and conversation tool — 'why is this module stable, concrete, and touched every sprint?' — not as a score. Gaming only becomes a problem when you make it a target, which is exactly the failure mode governance is meant to avoid.
  • What would you enforce as a hard build failure instead of a D threshold?
    Binary, non-gameable dependency facts: no dependency cycles between components (Acyclic Dependencies Principle), no consumer referencing an implementation component or another module's internal package, allowed-dependency allowlists per module, and layer-direction rules. Tools like ArchUnit, dependency-cruiser, Spring Modulith verification, or Java module boundaries make these deterministic. Metrics go on a dashboard; rules go in the build.
  • How does this metric hold up for a distributed system of services rather than a monolith's packages?
    Poorly, on its own. Most coupling between services is runtime and data coupling — shared databases, message schemas, synchronous call graphs, deployment ordering — none of which appears as a static type reference. You can apply A/I per service's codebase, but the meaningful analogues are API-contract versioning discipline, consumer-driven contract tests, call-graph fan-in, and change-coupling across repos.

It's like BMI for architecture: cheap to compute, genuinely correlated with a real condition at population scale, and badly wrong on individuals — the frozen utility is the athlete whose numbers look alarming and whose health is fine. You screen with it; you don't diagnose with it.

saying these in an interview costs you the question

  • Turning D into a hard CI threshold, which incentivises interface-for-interface's-sake and mechanical mirror packages
  • Comparing D across projects or languages as if the scale were universal, when boundaries and abstract-type semantics differ
  • Reporting only unsigned D, so the zone of pain and the zone of uselessness are indistinguishable
  • Believing a component on the main sequence is architecturally healthy — cohesion, cycles, and coupling direction are untested by it
  • Applying it unchanged to distributed systems, where most coupling is runtime and data coupling invisible to static analysis

context