skip to content

A team has decomposed its product into a dozen microservices, but every release still requires deploying all twelve together, and the checkout service is effectively down whenever any of three specific other services is down. What anti-pattern does this describe, and what does it tell you about the original decomposition?

level: seniorimportance: should knowfreq 65%

answer

  1. deploy topology vs actual coupling
  2. coordinated deploys = no real independence
  3. correlated availability = synchronous chains
  4. shared DB recouples 'separate' services
  5. wrong split axis = boundaries not bounded contexts

basics

~20 s

This is a 'distributed monolith': many separate services that are still forced to move and fail together like one big app, except now with all the extra network overhead. It means the services were split in the wrong places — they're not actually independent.

solid answer

~50 s

This is the 'distributed monolith' anti-pattern: the system has been physically decomposed into many independently-deployable processes, but those processes remain logically coupled — they must be released together and cannot tolerate each other's failures — so the team pays the full cost of distribution while getting none of microservices' actual benefits. It typically results from drawing service boundaries around technical layers or arbitrary team assignments rather than around real bounded contexts with stable, low-frequency interfaces; from synchronous call chains with no fallback or timeout so one service's downtime propagates through several hops; or from a shared database that recouples services that look separate in the deployment diagram. The fix is either to redraw the boundaries around the actual domain seams (possibly merging some services back together) or to decouple the calls that shouldn't be synchronous, rather than adding more services or more infrastructure.

go deeper

for a junior

Should recognize that having separate deployments doesn't automatically mean the services are actually independent, even without using the term 'distributed monolith.'

for a middle

Should name at least one concrete cause (synchronous chains with no fallback, or a shared database) and describe the coordinated-deploy symptom.

for a senior

Should be able to diagnose the anti-pattern from external signals (correlated deploys/outages) and propose both the boundary-redraw and the async-decoupling fixes.

for a principal

Should be able to make the organizational call — when to merge services back together versus invest in decoupling — and cite real precedent to support that recommendation to stakeholders.

## What a distributed monolith is A distributed monolith is a system that has the **deployment topology of microservices** — many separately built and separately deployed processes, usually communicating over HTTP or a message bus — but retains the **coupling behavior of a monolith**: the services can't actually be released independently, and the failure of one routinely takes down others that shouldn't logically depend on it. It's called a distributed monolith rather than just "bad microservices" because the diagnosis matters: the team has paid every one of the operational costs of distribution — network latency and failure handling, service discovery, per-service CI/CD pipelines, distributed tracing, on-call rotations per service — without receiving the benefits those costs are supposed to buy, namely independent deployability, fault isolation, and independent scaling. ## How to recognize it Two concrete symptoms describe how to recognize it. 1. **First, coordinated deploys.** If shipping any single feature routinely requires deploying several services together, in a specific order, because their APIs or data models are too tightly interlocked to evolve one without the others, the team has not actually achieved independent deployability — they've just moved the monolith's single release train into a more expensive form. 2. **Second, correlated availability.** If checkout becomes unavailable whenever any of three other specific services is down, checkout has a synchronous, unmitigated dependency on those services with no fallback or asynchronous decoupling — so its actual availability is the product of everyone's availability multiplied together, which is mathematically worse than a single monolith's availability. ## Why it goes wrong This goes wrong for a specific and common reason: services were split along the wrong axis. - **The classic mistake** is drawing boundaries around technical layers (a "validation service," a "formatting service") or around org-chart convenience rather than around real business capabilities with naturally low-frequency, stable interfaces — what Domain-Driven Design calls bounded contexts. When two services are split along a line that doesn't correspond to how the business domain actually decomposes, every real user-facing feature ends up needing changes on both sides of the line, so the two "independent" services are in practice always changing together. - **A second common cause** is defaulting every inter-service call to synchronous request/response with no fallback: if service A calls B calls C synchronously to answer one user request, A's availability is now bounded above by B's and C's combined, and a slow or down C takes A down too — except now there's network latency and failure-handling code at every hop that a real monolith wouldn't have needed at all. - **A third is a shared database.** If two "separate" services both read and write the same physical tables, they are coupled at the schema level regardless of how cleanly separated their deployment pipelines look. ## The worst of both worlds In production, the failure modes are the worst of both worlds: a distributed monolith is typically worse than a plain monolith, because it adds network unreliability and operational complexity on top of coupling that never actually went away. Teams in this state often respond by adding more infrastructure — a service mesh, more elaborate CI orchestration to sequence the coordinated deploys — which treats the symptom rather than the cause and increases the sunk cost of staying decomposed. ## The fix The actual fix is architectural, not operational: - **redraw the service boundaries** around real bounded contexts with genuinely independent data and interfaces — which sometimes means merging services back together, as Segment publicly did when it consolidated over 100 services after finding most of them had highly correlated load and interdependence - or, where the boundary is basically right but the communication pattern is wrong, **replace synchronous call chains** with asynchronous, event-driven communication and local fallbacks/caches so that one service being down degrades functionality gracefully instead of cascading "Distributed monolith" is the standard term precisely to warn teams against declaring victory just because their architecture diagram shows many boxes — the diagram is not the architecture; the actual coupling of deploys and failures is.

  • How would you diagnose whether a set of services is a distributed monolith versus healthy microservices, from the outside, without reading the code?
    Look at the deploy history: do releases of different services cluster together in the same window, or do they ship independently on their own cadence? Also look at incident correlation: when one service has an outage, do unrelated services' error rates spike at the same time? Frequent joint deploys and correlated outages are strong external signals of a distributed monolith.
  • If merging services back together isn't politically or technically feasible, what's the next-best fix?
    Replace synchronous, unmitigated call chains with asynchronous communication (events, message queues) where the business logic tolerates eventual consistency, and add caching, fallbacks, or default values for calls that must stay synchronous, so that one dependency's outage degrades a caller rather than taking it fully down.
  • Does having a service mesh or Kubernetes prevent a system from becoming a distributed monolith?
    No — a service mesh gives you better observability and traffic control over inter-service calls, but it doesn't fix boundaries drawn in the wrong place or synchronous call chains with no fallback. It can make the symptoms easier to see without addressing the underlying coupling.

It's like renting twelve separate offices for one company but still requiring every employee to show up and leave at the exact same time and be unable to work if any other office loses power — you now pay twelve leases for zero extra flexibility.

saying these in an interview costs you the question

  • Thinks having many separately deployed services automatically means the system has real microservices benefits
  • Can't explain why correlated availability across services is worse than a single monolith's availability
  • Proposes adding more infrastructure (mesh, orchestration) as the fix rather than questioning the boundaries
  • Doesn't consider that a shared database can recouple services that look separate in a deployment diagram
  • Assumes the fix is always 'split further' rather than possibly merging services back

context