skip to content

A company has split its monolith into 15 separately-deployed services, but every release still requires deploying most of them together in a fixed order, and a single slow service brings down request chains everywhere. What is this failure mode called, and what specific decomposition mistakes typically cause it?

level: seniorimportance: must knowfreq 75%

answer

  1. separate deploys, locked-step releases
  2. shared DB or shared domain library = smoking gun
  3. long synchronous call chains cascade failures
  4. worst of both worlds: ops cost + monolith coupling
  5. term popularized by Sam Newman

basics

~10 s

This is a distributed monolith: separate services still tightly coupled (shared code, chatty synchronous calls, or a shared database), so you can't deploy them independently — all of microservices' overhead, none of the benefit.

solid answer

~50 s

This is the distributed-monolith trap: services are deployed as separate processes but remain coupled tightly enough that they can't be built, tested, deployed, or scaled independently — the worst of both worlds. Typical causes include: boundaries drawn along technical layers instead of business capabilities, so one feature spans many services; a shared database or shared schema library that couples every service's data model together; synchronous call chains where service A calls B calls C calls D, so a slow or failing D degrades everyone upstream and versioning one breaks callers; and shared code libraries for business logic, not just utilities, that force lockstep releases. The fix is usually the same medicine as good decomposition in general: realign boundaries to business capabilities and bounded contexts, give each service its own datastore, replace synchronous fan-out chains with async events or aggregation at the edge, and enforce contract/versioning discipline so services can evolve independently.

go deeper

for a junior

Should recognize the term and give one example of what coupling looks like, such as services that must be deployed together.

for a middle

Should be able to list two or three concrete causes, such as shared DB, sync chains, or shared code, when shown symptoms.

for a senior

Should diagnose which specific coupling is responsible in a given scenario and propose a targeted fix, such as replacing a sync chain with an event or splitting a shared schema, rather than a generic 'redesign everything.'

for a principal

Should assess the organizational root cause, such as deploy-order runbooks or team structure not matching services, estimate the cost/risk of remediation versus living with it, and sequence a re-decomposition that doesn't require a full stop-the-world rewrite.

## What a distributed monolith is A **distributed monolith** is a system that has been physically split into multiple independently-deployed services but has not shed the coupling that made it a monolith in the first place — the result is that you pay every operational cost of distribution (network latency, serialization, more infrastructure to run and monitor, more places for partial failure) while getting none of microservices' actual payoff, which is independent development, testing, and deployment per service. It typically forms through a handful of specific, identifiable mistakes made during decomposition rather than as some unavoidable fate of splitting a system. ## The four decomposition mistakes behind it 1. **The first and most common cause is drawing boundaries along technical layers instead of business capabilities** — separate services for presentation, business rules, and data access — so that any real business change fans out across most of the services and has to be deployed together. 2. **The second is a shared database or shared schema**: if two services read or write the same tables, or even just share a library that defines the canonical shape of a core entity, then a schema change in one forces a coordinated migration and deploy in the other, and neither can evolve its data model independently even though they run in separate processes. 3. **The third is deep synchronous call chains**: when handling one request requires service A to call B, which calls C, which calls D, the whole chain's availability is the product of every link's availability, and the whole chain's latency is at least the sum of every link's latency — a slowdown or outage anywhere propagates upward to every caller, with no isolation the way there would be with async, event-driven integration. 4. **The fourth is shared business-logic libraries**, as opposed to genuinely generic utility code, baked into every service: if a validation rule or pricing calculation lives in a shared package that every service links against, changing that rule still requires rebuilding and redeploying every consumer in lockstep, exactly as it would in a monolith's shared internal module. ## Why it is so expensive to fix The trade-off in fixing this is significant, which is part of why organizations stay stuck in it: - **Untangling a shared database** means picking apart which service really owns which tables and migrating data ownership, a multi-step, carefully sequenced project, not a config change. - **Replacing synchronous call chains with asynchronous events** means accepting eventual consistency somewhere you previously had a synchronous guarantee, and reworking client code to no longer expect an immediate, always-fresh answer. - **Splitting a shared code library into per-service copies**, accepting some duplication, trades an easy central update point for genuine independence. None of these fixes are free, which is why distributed monoliths, once formed, tend to persist for a long time — the remediation looks a lot like doing the original decomposition work properly, except now on a live system with sunk cost in the existing shape. ## Failure modes The failure modes show up operationally rather than in a design document: - release runbooks that specify a fixed order across several 'independent' services because contract changes silently break downstream consumers; - a single slow or degraded service, or its external dependency, causing timeouts and error-rate spikes across several unrelated-looking services because they're all synchronously chained together; - integration tests that require spinning up the full set of services to pass because none of them can be meaningfully verified in isolation; - and on-call engineers who can't confidently say what breaks if they deploy just one service, so they default to deploying several together 'to be safe,' which further entrenches the coupling. ## Where it shows up This term and the underlying anti-pattern are extensively discussed by **Sam Newman** in his writing on microservices, as one of the primary risks of decomposing a system's deployment topology without also decomposing its actual coupling — data ownership, call topology, and contract stability. A common concrete instance is a checkout flow where Cart, Pricing, Tax, and Payment services all call each other synchronously and share a customers table maintained by none of them cleanly; a slow third-party tax-rate lookup inside the Tax service then degrades checkout latency across the board, and nobody can deploy a Pricing change without first checking whether it breaks Tax's assumptions about the shape of a price object — **the system is fifteen containers wide and one monolith deep**.

  • Why do organizations end up here even when everyone explicitly set out to build microservices?
    Most often because the technical split into separate deployables happened first, while the underlying design — data model, call patterns, team ownership — never actually changed to match. It's easy to carve a monolith's modules into separate repos and processes; it's much harder to also rethink the boundaries, own data per-service, and replace synchronous coupling with async integration, so teams do the easy part and stop.
  • What's a fast diagnostic to check whether a set of services is a distributed monolith?
    Ask 'can I deploy this one service alone, right now, without coordinating with any other team or deployment?' If the honest answer involves a fixed deploy order, a shared migration script, or 'we always release these three together,' that's a distributed monolith regardless of how many separate repos or containers exist. Another quick check: does a single service's outage or slowdown cascade through several others via synchronous calls?
  • Once you're in this state, is the fix to merge everything back into a monolith?
    Rarely — merging back gives up the deployment and ownership boundaries you may still want, and the underlying coupling problem, shared model and sync chains, would just move back in-process. The more common fix is to re-decompose along business-capability lines, give each service its own datastore, and replace synchronous fan-out with asynchronous events, even though this is a slower and more invasive fix than the original split was.

Like cutting a single circuit board into 15 physical pieces connected by wires that all have to be soldered and powered on together in the right sequence — you now have the packaging overhead of 15 separate boards, but none of the independence, because the wiring still forces them to act as one circuit.

saying these in an interview costs you the question

  • Says having many separate services automatically means the system isn't a monolith
  • Can't name the shared-database or shared-library smell as a cause of distributed monolith
  • Proposes 'just add more retries and timeouts' as the fix for cascading synchronous coupling instead of addressing call topology
  • Assumes independent deployability is guaranteed by using containers or Kubernetes rather than by the actual coupling between services

context