skip to content

When picking the first capability to peel off a monolith with a strangler-fig migration, what makes a good 'seam,' and what characteristics of a module make it a poor first choice?

level: middleimportance: should knowfreq 50%

answer

  1. low coupling / high cohesion boundary
  2. bounded context (DDD)
  3. code churn / co-change analysis
  4. shared-table = bad seam
  5. low blast radius for first cut

basics

~20 s

A good first piece to split off is one that talks to the rest of the system through few, well-defined connections and doesn't share much data — so pulling it out doesn't require touching everything else.

solid answer

~40 s

A seam (Michael Feathers' term) is a place where you can alter behavior without editing the surrounding code — in migration terms, a boundary with low coupling and high cohesion where a new implementation can be inserted with minimal ripple effects. Good first candidates: self-contained bounded contexts with few inbound/outbound dependencies, well-defined APIs already, owned end-to-end data (not scattered across tables shared with unrelated features), low change-coupling with other modules, and lower business risk so mistakes are cheap while the team builds migration muscle. Poor first choices: modules at the center of the dependency graph, modules with heavy synchronous chatter to many other parts, modules sharing database tables with several unrelated features, and anything in the critical path with zero tolerance for downtime before the team has proven its tooling.

go deeper

for a junior

Should understand intuitively that some parts of code are more tangled than others and tangled parts are riskier to pull out first.

for a middle

Should describe concrete signals — few dependencies, owns its own data, clear API — and name at least one analysis technique (git churn or dependency graph).

for a senior

Should discuss the data-coupling trap (looks isolated in code, coupled via shared tables) and the strategic sequencing argument (easy first to build tooling, not necessarily highest value first).

for a principal

Should frame seam selection as a portfolio/risk decision across the whole migration roadmap, balancing blast radius, dependency graph centrality, and where early wins build organizational trust to fund the multi-year effort.

## What a seam is A seam, in **Michael Feathers'** original usage from legacy-code work, is a place in a system where you can alter behavior without editing the code at that exact point — a joint along which the system can be safely pulled apart. Applied to monolith decomposition, a good seam is a module boundary with **low coupling** to the rest of the system and **high internal cohesion**, such that extracting it into a separate service touches as little of the surrounding codebase and data model as possible. Finding good seams is arguably as important a skill in a strangler-fig migration as executing the extraction itself, because a wrong choice can turn what should be a confidence-building early win into the migration's first and most damaging failure. ## Two sources of signal Concretely, useful signals for identifying seams come from two sources: the code and the data. - **On the code side**, teams look at dependency/call-graph analysis (how many other modules call into this one, and how many does it call out to — low fan-in and fan-out is favorable), and at git co-change or churn analysis (which files tend to be edited together in the same commits; files that rarely co-change with the rest of the system are more independent than the call graph alone might suggest). - **On the data side** — often the more decisive factor — teams look at which database tables a candidate module reads and writes, and whether those tables are exclusive to it or shared with other, unrelated features; a module that owns its own tables outright is a far better seam than one that looks structurally isolated in code but reaches into a dozen tables other features also depend on. Domain-driven design's **bounded-context mapping** is frequently used as a complementary technique: talking through the domain with subject-matter experts to find natural conceptual boundaries often lines up with, and helps validate, what the coupling analysis finds mechanically. ## Why the first choice carries so much weight The reason this matters so much is that strangler-fig migrations are typically sequenced from easier to harder, and the first extraction sets the tone. It's where the team builds and proves: - its routing infrastructure; - its rollback procedures; - its data-migration tooling; - and its confidence with stakeholders that the whole multi-year effort is worth funding. A poorly chosen first seam — one with hidden coupling that only becomes apparent mid-extraction — can turn a planned two-week project into a stalled six-month one, souring organizational appetite for the rest of the migration. Good candidates therefore combine technical isolation with modest, not necessarily maximal, business value: valuable enough that finishing it matters, but not so central or so risk-sensitive (core billing, core auth) that a mistake is catastrophic before the team's tooling is proven. ## What a poor seam looks like Poor seams share recognizable characteristics: - they sit at the center of the dependency graph with high fan-in from many other modules, so almost everything depends on their current in-process behavior and must be touched or at least tested against the new remote version; - they share tables with several unrelated features, so no clean data boundary exists without either giving the new service broad cross-schema database access (defeating the purpose of the split) or building many new data integrations at once; - or they participate in synchronous, chatty interaction patterns that would become slow, unreliable network calls immediately, a classic sign that the module is not actually a separate bounded context but an artificial slice of one. ## The trap of one-sided isolation A commonly underestimated trap is a module that looks isolated by one signal but not the other — the canonical example being a **reporting or analytics module** that almost nothing else calls (very low inbound code coupling, which looks great on paper) but which itself reads directly from thirty tables owned by ten different features (very high outbound data coupling). Judged only by "who calls this," it looks like an ideal first seam; judged by total coupling including data dependencies, it's one of the worst, because extracting it either requires it to keep reaching across service boundaries into other services' private data (undermining the whole point of decomposition) or requires building data-replication or API integrations with ten different teams simultaneously before the extraction can even begin. Effective seam identification always checks both directions and both code and data before committing.

  • What's a concrete technique for finding low-coupling seams in a large legacy codebase before committing to one?
    Run a git co-change analysis (which files/modules tend to be edited in the same commits) alongside static call-graph/dependency analysis to find clusters with few edges to the rest of the system, and cross-check against the database schema to see which tables the candidate module reads/writes exclusively versus shares. Combining code-level coupling with data-ownership analysis catches the common trap of a module that looks structurally isolated in code but is deeply coupled through shared tables.
  • Why might a team deliberately avoid extracting the highest-value module first, even though it would deliver the most business impact?
    The highest-value modules are usually also the most central, most coupled, and most risk-sensitive, so extracting them first means learning the migration playbook — routing, dual-write handling, rollback — under maximum stakes. Teams typically extract one or two lower-risk seams first to validate tooling, observability, and rollback procedures, then apply that proven process to the higher-value, higher-risk modules.
  • A reporting module looks like an easy first seam because almost nothing calls into it, but it reads directly from 30 tables owned by other features. Is it a good first choice?
    No — despite low inbound coupling, it has extremely high outbound data coupling, so extracting it either requires giving the new service broad direct database access (violating service data ownership) or building 30 different data-access integrations. A better signal is combined inbound and outbound coupling, not just whether other code calls into it.

Like choosing which branch to prune first when pollarding a tree: you start with a branch that has few intertwined limbs so you can cut cleanly, not the thick central branch that everything else grows out of.

saying these in an interview costs you the question

  • Picks the module purely because it 'seems small' without checking data/table coupling
  • Doesn't mention checking database table ownership as part of seam identification
  • Suggests starting with the most business-critical or most central module for maximum impact
  • Conflates 'seam' with 'this team owns it' rather than actual technical coupling
  • No mention of gradually building confidence/tooling via an easier extraction first

context