For a platform with 30+ engineers and a domain that is still not well understood, what makes a modular monolith the wrong default choice, and what long-term organizational failure mode most often erodes a modular monolith's boundaries even after they were enforced correctly at launch?
answer
- premature boundaries = guessed wrong, costly to redraw
- different scaling/runtime/compliance needs beat shared-process constraint
- boundary is policy enforced by tooling, not physically inescapable like a network
- exemption creep: allowed-dependency entries and suppressions pile up over years
- erosion is invisible until extraction/audit reveals it
basics
~20 sIt's the wrong pick mainly when a domain is still changing shape fast, so boundaries would be guessed wrong and redrawn constantly, or when different parts genuinely need very different scaling/tech - and even a well-built one can rot years later if governance around the enforcement rules quietly weakens.
solid answer
~1 minA modular monolith is a poor default in two situations: when the domain is genuinely not understood yet, so any module boundary drawn today is a guess likely to be wrong, and the cost of moving code between modules inside a monolith, while cheaper than moving it between services, still isn't free, versus deferring the boundary decision and staying in a single undivided codebase until the domain settles; and when parts of the system have fundamentally different non-functional needs from day one - wildly different scaling profiles, different compliance/isolation requirements, or a genuine need for different runtimes - where the shared-process constraint of a monolith (everyone scales together, everyone shares a runtime, one bug can take everyone down) is itself the problem, not a temporary inconvenience. The long-term organizational failure mode is 'boundary erosion through exemption creep': the verification tool is real at launch, but under deadline pressure a team adds a suppression, an allowed-dependency entry, or a 'temporary' shared utility class that both sides depend on, each individually reasonable, until years later the boundary has enough sanctioned holes in it that it's effectively decorative - the code has quietly become a big ball of mud again, just with tests that still pass.
go deeper
Should have a basic notion that if a team doesn't know its domain very well, guessing at module boundaries too early can be wasteful, without necessarily discussing erosion or exemption dynamics.
Should be able to name at least one concrete scenario where a modular monolith isn't the right default, such as very different scaling needs across parts of the system.
Should distinguish between the two distinct causes of failure - premature/wrong boundary choice versus long-term erosion of a correctly chosen boundary - and describe how each shows up.
Should discuss governance mechanisms for preventing erosion over years across a growing team, weigh modular monolith against monolith-first and microservices-first strategies based on domain maturity and non-functional diversity, and reason about the organizational incentives that drive erosion.
## Two situations where it is the wrong default There are two genuinely distinct situations where a modular monolith is the wrong default, and they're worth separating because they call for different alternatives. ## The first: the domain is not yet understood The first is when the domain is not yet understood: a brand-new product, a company entering a market it has never operated in, or a team still discovering what the real seams of the business even are. - Any module boundary drawn under those conditions is a guess, and the guess is often wrong, because the natural decomposition of a domain usually only becomes clear after real usage and iteration reveal where change actually clusters. - Redrawing a wrong boundary inside a monolith is cheaper than redrawing it across separately deployed microservices, but it still isn't free: other code has already been written against the wrong module's public API and DTOs, and migrating those callers, deprecating the old interface, and moving entities between modules' internal packages is real, deliberate work. In this situation, many teams deliberately choose to start with a single, undivided codebase — sometimes called 'monolith first' — and defer imposing enforced boundaries until usage patterns make the real seams obvious, rather than paying repeatedly to redraw premature ones. ## The second: different needs from day one The second situation is when parts of the system have fundamentally different non-functional needs from the outset — not eventually, but from day one. A part of the system that needs to scale independently by orders of magnitude, that has a genuine compliance requirement for physical or network-level isolation such as a payment-card-scope boundary, or that is a much better fit for a different runtime or language, is not well served by a modular monolith no matter how cleanly its module is bounded in code, because the shared-process constraint — everyone in the same JVM, everyone scaling and deploying together — is itself the mismatch. A code-level boundary, however well enforced, cannot give one module its own scaling profile or its own separately audited network perimeter; that requires an actual process and deployment boundary, which is what a microservice or a genuinely separate deployable provides. ## The slow failure: boundary erosion through exemption creep Even when neither of those conditions applies at launch, and the modular monolith is initially the right call with genuinely well-enforced boundaries, a slower and more insidious failure mode tends to erode it over years: boundary erosion through exemption creep. Unlike a microservice's network boundary, which is physically inescapable — you literally cannot call another service's private in-memory state without going over the wire — a modular monolith's boundary is a policy enforced entirely by tooling running inside one process, and policy can be weakened from the inside. A critical production bug needs a fix in twenty minutes, and the 'correct' fix of adding a new API method and waiting for review and deploy is slower than importing the internal class directly and suppressing the check just this once with a documented TODO. Multiply that by dozens of engineers over several years, and each individual exemption is defensible on its own, but the aggregate quietly reintroduces exactly the everything-can-reach-everything coupling the enforcement was built to prevent. ## How the erosion shows up How this shows up in practice is telling: - Reviewers stop treating a boundary-tool failure as automatically blocking because there are already 'too many exceptions to bother'. - A new engineer can't tell from the code which dependencies are load-bearing architecture and which are legacy escape hatches nobody ever cleaned up. - When the team finally attempts a real service extraction years later, they discover the boundary is far leakier than the passing CI check implied, because the exceptions were granted once and never revisited. ## The ongoing tax, not the one-time setup Teams that have publicly discussed maintaining large modular monoliths over many years, including well-known cases of Ruby-on-Rails codebases adopting boundary-enforcement tooling at scale, describe this as an ongoing tax, not a one-time setup. It requires: - periodic audits of the exception list; - a policy that every suppression or allowed-dependency entry carries an owner and either an expiry date or a linked ticket rather than being permanent; - treating the trend in the number of exemptions — not just whether the verification test currently passes — as a first-class health metric the team actively tries to shrink, not merely tolerate.
- Why might a team deliberately choose to start with a single undivided codebase rather than a modular monolith, if they eventually expect to need module boundaries?If the domain is still being discovered and the team doesn't yet know where the real seams are, drawing boundaries too early means guessing, and the guess is often wrong; moving code and callers across a wrong boundary still costs real time even inside one monolith, since other code has already been written against the wrong module's public API. Deferring the split until real usage patterns reveal the natural seams can be cheaper overall than repeatedly redrawing premature boundaries.
- What's a concrete, low-ceremony governance practice that helps prevent boundary-exemption creep?Requiring every suppression or allowed-dependency exception to carry an owner and either an expiry date or a linked ticket to remove it, rather than letting it become a permanent, silent entry, combined with periodically reviewing and reporting the total count of exemptions as a trend so growth in exceptions is visible to the team, not just to whoever added the tenth one.
- How is boundary erosion in a modular monolith fundamentally different from the equivalent risk in microservices?In microservices, the network is a physical boundary - you cannot call another service's private in-memory state without going over the wire, so the boundary can't be silently weakened by an internal shortcut, though microservices have their own coupling risks such as shared databases or chatty synchronous chains. In a modular monolith, the boundary is enforced entirely by tooling policy inside one process, which means a determined engineer or a deadline can add an exception that quietly weakens it, and nothing but discipline and review stops that from accumulating.
Like a gated community whose gate code gets shared 'just this once' with a dozen different contractors over the years - each exception individually reasonable, but eventually so many people have the code that the gate isn't actually restricting anyone anymore, even though it's still standing there looking secure.
saying these in an interview costs you the question
- treats 'draw the boundaries as early as possible' as always correct, with no mention of the cost of getting them wrong in an unclear domain
- assumes an enforced boundary, once set up, stays enforced forever with no further effort
- can't name any scenario where a modular monolith is the wrong choice
- believes boundary erosion can't happen because a verification tool exists, without considering exemptions/suppressions
- conflates the 'boundary can be weakened from inside' risk with something that also applies equally to microservices' network boundary