You own a widely used shared library whose consumers each pull it in for only one of five unrelated feature areas. Walk through how you would apply the Common Reuse Principle to split it, and what new costs the split introduces.
answer
- Evidence first: usage matrix + irrelevant-release rate
- Cluster by joint usage, not taxonomy
- Extract shared kernel → core, most stable, no cycles
- Aggregator artifact + dated deprecation, never big-bang
- New costs: N pipelines, compat matrix, diamonds, ownership
basics
~20 sMeasure which classes each consumer actually uses, cluster by observed joint usage, split into one artifact per cluster, and keep the old artifact as a thin aggregator that depends on the new ones so nothing breaks. New costs: more versions, releases, pipelines and coordination.
solid answer
~50 sStart with evidence, not taxonomy: scan consumers' imports and build graphs to see which classes are used together, and check which of the last N releases were irrelevant to which consumer. Cluster classes by observed joint usage; keep tightly bound structures (an interface with its main implementation, a type and its parameter types, a builder and its product) inside one cluster. Extract any genuinely shared kernel into its own small, very stable artifact so the new components don't end up depending on each other cyclically. Publish one artifact per cluster, then re-publish the original coordinates as a thin aggregator depending on all of them so existing consumers keep compiling; deprecate it with a stated window and migrate consumers to the specific artifacts. New costs: N versioning streams and changelogs, N pipelines, cross-artifact compatibility matrices, the risk of lockstep versioning if the seams are wrong, more discovery friction, and a larger dependency graph that must stay acyclic.
code
text · 14 linesBEFORE
acme-lib 4.7.0 -- pdf, events, csv, auth, imaging (+ all their transitives)
AFTER (physical split)
acme-core 1.0.0 (shared value types, errors, config; no deps)
acme-pdf 1.0.0 -> acme-core
acme-events 1.0.0 -> acme-core
acme-csv 1.0.0 -> acme-core
acme-auth 1.0.0 -> acme-core
acme-imaging 1.0.0 -> acme-core
COMPATIBILITY BRIDGE (keeps existing consumers compiling)
acme-lib 5.0.0 (no code) -> all of the above, re-exported
deprecated, EOL announced two releases outgo deeper
Say you'd look at which classes consumers actually use, group those, split into separate published packages, and keep the old package working for a while.
Add the shared-kernel extraction, the aggregator and deprecation migration path, and two or three concrete new costs.
Lead with measurement (usage matrix, irrelevant-release rate, transitive attribution), justify granularity against an artifact budget, handle cycles and stability, and pick a versioning policy with a BOM.
Frame as a platform-portfolio decision: ownership model, artifact-count budget, deprecation policy, org-wide migration tooling, and a follow-up measurement to prove the split delivered.
## Step 0 — Confirm the problem is real Don't split on aesthetics. Collect three numbers: 1. **Usage matrix.** For each consumer, which public types does it actually reference? Get this from import scanning, build-graph analysis, bytecode/AST references, or published-API telemetry. The output is a consumer × class matrix. 2. **Irrelevant-release rate.** Over the last N releases, what fraction was driven by code a given consumer never uses? This is the direct cost of the CRP violation. 3. **Transitive attribution.** Which third-party dependencies exist in the tree *only* because of a particular feature area? That is the closure a split would remove for everyone else. If the matrix is dense (most consumers use most of it), you do not have a CRP problem and splitting will only add cost. ## Step 1 — Cluster by observed joint usage Group classes so that within a group, classes tend to be used by the same consumers; across groups, usage is disjoint. Practical rules: - **Keep structurally inseparable things together.** An interface with its primary implementation; a type and the types in its method signatures; a builder/factory and its product; an exception hierarchy and the API that throws it. Splitting these produces components that are useless alone — the opposite CRP failure, since a consumer must then take all the fragments anyway. - **Beware the "shared kernel".** Usually a handful of types (config, error types, core value objects) are used by all five areas. Do not duplicate them and do not let the five components depend on each other to get them. Extract them into a **core** artifact that everything depends on and that depends on nothing. It should be the most stable component in the set (Stable Dependencies Principle: depend in the direction of stability). - **Check for cycles.** If two proposed components would depend on each other, the seam is wrong; either merge them or extract the shared abstraction into core. Cycles across published artifacts violate the Acyclic Dependencies Principle and are far more painful than in-process cycles because they force lockstep releases. ## Step 2 — Choose the granularity deliberately Five feature areas does not necessarily mean five artifacts. Merge areas that are always used together or always change together (the CCP pull). Ask what an artifact-count budget can bear: every artifact carries fixed overhead — pipeline, ownership, changelog, docs, deprecation policy, security review. Prefer the smallest number of artifacts that eliminates the observed irrelevant-release traffic. ## Step 3 — Migrate without breaking anyone The safe pattern: 1. Publish the new fine-grained artifacts (`lib-core`, `lib-pdf`, `lib-events`, ...) at 1.0. 2. Republish the **original coordinates** as a thin **aggregator**: same group/name, new version, containing no code, depending on all the new artifacts and re-exporting their API. Existing consumers upgrade and keep compiling with zero source changes. 3. Mark the aggregator deprecated, publish a mapping table ("class X now lives in lib-pdf"), and give a dated end-of-life. 4. Help consumers migrate (codemods, dependency-report tooling, a migration guide). 5. Retire the aggregator. Avoid the two common mistakes: a big-bang split with no aggregator (breaks everyone at once), and an aggregator with no end-date (nothing ever migrates and you now maintain N+1 artifacts forever). Also decide the **versioning policy**: independent semantic versions per artifact (maximum CRP benefit; requires a documented compatibility matrix), versus a single train version applied to all (simpler to reason about but reintroduces lockstep and thus much of the original problem). If the seams are good, independent versioning is the right answer; if you find you must always bump all five together, your seams follow change lines that never diverge and the split bought you less than you hoped. ## Step 4 — Verify the benefit Re-measure after a quarter: irrelevant-release rate per consumer, transitive-tree size, CVE count, build/bundle size. If the numbers didn't move, the split was cosmetic. ## The new costs, stated plainly - **Release and versioning overhead** times N: changelogs, release notes, compatibility matrix, CI pipelines, publishing credentials. - **Coordination cost** for cross-cutting changes (a change touching core plus three areas now needs an ordered, multi-artifact release — the CCP penalty). - **Diamond risk:** consumers may end up with two artifacts pinning different versions of core. - **Discovery friction:** newcomers must learn which artifact holds what; mitigate with docs and a BOM/platform manifest that pins a compatible set. - **Ownership dilution:** more artifacts than the team can meaningfully own leads to unmaintained ones. - **Refactoring friction:** moving a class between artifacts is now a breaking, semver-visible change rather than an internal edit. A BOM (bill of materials) or version-catalog artifact that pins a known-good combination is the standard mitigation for the diamond and discovery costs.
- What do you do about the handful of types used by all five feature areas?Extract them into a small, highly stable core artifact that everything depends on and that depends on nothing else. Duplicating them creates type-identity conflicts; letting the feature components depend on each other to reach them creates cycles and forces lockstep releases.
- Independent semantic versions per new artifact, or one shared train version?Independent versions capture most of the CRP benefit, since a consumer only sees releases of what it uses; the price is a documented compatibility matrix, usually published as a BOM or version-catalog artifact. A single train version is simpler but reintroduces lockstep releases and negates much of the split.
- How would you know afterwards whether the split was worth it?Re-measure the metrics that justified it: irrelevant releases absorbed per consumer, transitive dependency count and size attributable to unused code, open CVEs in that subtree, build and bundle time. If those didn't improve, the seams were cosmetic rather than usage-based.
Splitting a general store into specialist shops. Customers stop wading through unrelated aisles, but now you run five leases, five inventories and five sets of opening hours — and you need a directory at the entrance so people can still find things.
saying these in an interview costs you the question
- Splitting along a conceptual taxonomy ("models", "services", "utils") instead of measured joint usage — this usually makes CRP worse, since every consumer needs one of each.
- Big-bang split with no aggregator or deprecation window, breaking all consumers simultaneously.
- Duplicating shared types into each new artifact instead of extracting a core, causing type-identity and version conflicts.
- Allowing the new artifacts to depend on one another circularly, forcing permanent lockstep releases.
- Assuming splitting is free; ignoring pipeline, versioning, ownership and discovery overhead.
- Keeping the aggregator forever with no end-of-life, so you maintain N+1 artifacts and nobody migrates.