Is a distributed monolith always a mistake to fix immediately? Discuss when it can be a rational, temporary state, and how Conway's Law and team topology influence whether it persists.
answer
- transitional vs. stuck state
- strangler fig passes through this look deliberately
- Conway's Law: architecture mirrors org chart
- Inverse Conway Maneuver
- team boundary must cross service boundary to create pressure
basics
~20 sNot always - during a migration, having some temporary shared coupling can be the safer path. It becomes a real problem when it's permanent and when the teams responsible for the coupled services aren't structured to fix it.
solid answer
~40 sA distributed monolith is often a legitimate transitional state during an incremental strangler-fig migration - you extract a service's code before fully splitting its data, accepting temporary shared-database coupling because doing everything at once is riskier. The failure mode isn't the temporary state; it's when the org never finishes the migration because the pain is tolerable enough that nobody prioritizes fixing it, or because Conway's Law works against you: if a single team, or tightly coordinated teams, own both sides of the coupling, there's no organizational pressure pushing toward autonomy, since the humans coordinate informally anyway. Real, lasting fixes usually require redrawing team ownership to match the target service boundaries, not just refactoring code, because the architecture tends to converge back to whatever shape the organization's communication structure takes.
go deeper
Not expected to reason about org design; should just recognize that 'in progress' and 'permanently stuck' are different situations.
Should articulate that migrations pass through temporary coupling states without it being alarming, given a plan.
Should invoke Conway's Law by name and explain why team ownership overlap lets coupling persist unnoticed.
Should discuss the Inverse Conway Maneuver, the risk of reorganizing teams without technical remediation, and give criteria such as trend, ownership, and timeline for distinguishing rational transitional coupling from a stuck anti-pattern.
## Transitional against stuck The anti-pattern label properly applies to the stuck, permanent state, not to every instance of temporary coupling. A team executing a well-planned **strangler-fig migration** will, by design, pass through states that look exactly like a distributed monolith — a shared database mid-split, some synchronous calls not yet converted to async — and that's expected and fine as long as it's actively moving toward the target, with a plan and shrinking coupling over time. It becomes the true anti-pattern when it is the resting, steady state: no plan, no shrinking trend, coupling accepted indefinitely as simply how the system is. ## When accepting it is rational There are rational reasons to accept it temporarily or even in limited permanent scope. - **Premature decomposition is a real risk.** Splitting data and call paths before the domain model is well understood often produces the wrong boundaries, forcing an expensive redo, so a temporary shared database can be safer than guessing. - **The migration has a real cost.** Full data-ownership migration, with dual writes, event sourcing, and contract tests, is genuinely expensive engineering time that competes with feature work, so an organization may rationally sequence it behind higher-value work while explicitly tracking it as debt with an owner and a rough timeline, rather than deferring it indefinitely with no plan. - **Scale also matters.** A small team running several 'microservices' that are really a distributed monolith may be entirely rational, since the coordination cost of a shared database is low when one team already touches everything, and the organization hasn't yet grown enough independent teams to benefit from true service autonomy. ## Why Conway's Law keeps it alive **Conway's Law**, Melvin Conway's 1968 observation that organizations design systems mirroring their own communication structure, explains why this persists. - **One team owns both.** If one team, or a small, tightly synced set of teams, is responsible for both 'Service A' and 'Service B,' the coordination overhead of a cross-boundary change is near zero for that team — they simply talk to themselves. There's no organizational pain driving investment in decoupling the code, so shared-database and lockstep-deploy patterns persist indefinitely because the team absorbs the coupling cost internally without much friction. - **The team boundary crosses the service boundary.** Conversely, when the team boundary actually crosses the service boundary — two separate teams, on different roadmaps, with different priorities, owning A and B respectively — the coordination cost of the existing coupling becomes visible and painful: blocked pull requests, scheduling conflicts, and cross-team meetings to plan every shared release. That visible organizational pain is typically what forces genuine investment in autonomy. ## The implication for practice This has a direct implication for practice: the fix for a chronically stuck distributed monolith is often organizational before it's technical. If genuine independence between two services is the goal, put genuinely independent teams in charge of them — the **'Inverse Conway Maneuver,'** popularized in the Team Topologies literature, deliberately reorganizes team boundaries toward the target architecture and lets the resulting communication friction pressure the code to follow. The reverse mistake is also real: restructuring teams around service boundaries without doing the corresponding technical remediation work produces two teams fighting over one shared, still-coupled codebase, which can be worse in the short term, with constant cross-team blocking, than one team quietly tolerating self-inflicted coupling. ## A worked scenario A worked scenario illustrates the mechanism: a company splits its 'Orders' monolith module into a nominal 'Orders Service' and 'Fulfillment Service' during a reorg, assigning each to a different team, but doesn't yet split the underlying database or the synchronous call chain between them. Cross-team friction spikes immediately — schema-change pull requests blocked on cross-team review, deploy-coordination meetings — which is unpleasant short term, but is exactly the pressure that, within a couple of quarters, drives the teams to actually invest in single-writer ownership and event-driven decoupling, because the pain is now felt by the people with the authority and incentive to fix it, rather than absorbed silently by one team that never prioritized it. The principal-level insight this reveals is that technical remediation plans are necessary but not sufficient on their own: they only receive sustained investment, and don't silently regress back into coupling, when the team topology is aligned so that autonomy is the path of least resistance for the people doing the day-to-day work.
- How would you tell, as an outsider evaluating a system, whether its distributed-monolith coupling is a rational transitional state or a permanently stuck anti-pattern?Ask for the migration plan and check for a trend: is the amount of shared-database or synchronous coupling measurably shrinking release over release, with an owner and a rough timeline, or has it looked the same for a year or more with no active work against it? A stuck state typically has no ticket, no owner, and only a vague 'we know, it's on the roadmap' with no actual roadmap.
- What is the 'Inverse Conway Maneuver,' and why might it help fix a chronically stuck distributed monolith?It's the deliberate practice of restructuring team boundaries to match the target architecture before, or while, doing the technical work, on the theory that Conway's Law will happen regardless, so you might as well point it at the architecture you want. Splitting the teams first introduces real coordination friction across the service boundary, which creates organizational pressure to actually finish decoupling the code, rather than letting technical remediation stall indefinitely inside a single team's backlog.
- Can reorganizing teams to match service boundaries make things worse?Yes - if you split team ownership before doing any technical decoupling, you get two teams that now have to coordinate every change to a still-shared database and still-synchronous call chain, which is often more painful short-term than one team quietly absorbing the coupling itself. It's generally best paired with, or quickly followed by, the technical remediation work, not left as a reorg alone.
Like two housemates who technically have separate bedrooms but share one bank account - as long as it's the same two people managing money together, the shared account causes no friction; the friction only appears once they try to live as financially independent adults, which is exactly when they're forced to actually split the account.
saying these in an interview costs you the question
- Treats every instance of temporary coupling during a migration as automatically an anti-pattern requiring immediate fix
- Never mentions Conway's Law or team topology when discussing why coupling persists
- Assumes reorganizing teams alone, without technical remediation, fixes coupling
- Can't distinguish a migration that's actively shrinking coupling from one that's permanently stuck
- Thinks small startups should always fully decouple every service regardless of team size