skip to content

Monolith-to-Microservices Migration

Extracting services from a monolith incrementally: find a seam, put a facade in front, route traffic gradually, and decompose the database last. Parallel-run and branch-by-abstraction techniques let you cut over without a big-bang release, which is what interviewers want to hear.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a monolith-to-microservices migration, what is the strangler fig pattern and why do teams use it instead of a full rewrite?

level: juniorimportance: must knowfreq 70%

answer

  1. vine strangling host tree
  2. routing facade / proxy layer
  3. incremental not big-bang
  4. shrink monolith leaf by leaf
  5. reversible per-slice cutover

basics

~10 s

You build new features/services alongside the old monolith and slowly route traffic to them, until the old system does nothing and can be deleted. It avoids one risky big-bang rewrite.

solid answer

~50 s

The strangler fig pattern, named after a vine that grows around a host tree and eventually replaces it, is an incremental migration strategy: instead of rewriting a monolith in one big-bang release, you place a routing layer (proxy, facade, or gateway) in front of it, then extract one capability at a time into a new service, redirecting matching traffic to the new service while everything else still flows to the monolith. Each extraction is small, independently deployable, and reversible by re-routing back. Over months or years, the monolith shrinks as more traffic is 'strangled' away, until it either disappears or remains as a thin legacy core. Teams prefer it over a full rewrite because it de-risks delivery — you ship value continuously, get production feedback on each slice, and avoid the classic 'rewrite that never ships' trap where old and new systems must be kept at feature parity for years.

go deeper

for a junior

Should know the one-sentence idea — incremental replacement with a routing layer, safer than a rewrite — and be able to say why a big-bang rewrite is risky.

for a middle

Should be able to describe the routing facade concretely (URL path matching, feature flags) and walk through extracting one endpoint end to end.

for a senior

Should discuss the failure mode of stalled migrations, decommissioning discipline, and how to sequence extraction by coupling/risk.

for a principal

Should connect the pattern to organizational risk management — funding multi-year migrations, setting decommission SLAs/exit criteria, and knowing when strangler-fig is the wrong tool entirely.

## Where the name comes from The strangler fig pattern is named after tropical strangler fig vines that germinate high in a host tree's canopy, send roots down to the ground, and gradually envelop the host with their own growing trunk; over years the vine becomes self-supporting and the original tree, deprived of light and space, dies and rots away, leaving the vine standing in its shape. **Martin Fowler** applied this image in 2004 to describe an incremental system-replacement strategy, and it has since become one of the standard playbooks for moving a monolith toward a microservices (or any new) architecture without a single risky cutover. ## The mechanics Mechanically, the pattern starts by placing a **routing layer** — a reverse proxy, API gateway, or an internal facade inside the application — in front of both the existing monolith and (eventually) the new services being built. Initially, 100% of traffic flows through this layer straight to the monolith unchanged. The team then picks one capability (a URL path, a set of endpoints, a bounded business capability) and builds a new implementation of it as an independent service. The routing layer is updated to send requests matching that capability to the new service instead, while every other request continues to the monolith exactly as before. This is repeated capability by capability: 1. pick a **seam**; 2. build the replacement; 3. redirect matching traffic; 4. validate; 5. and — critically — **delete** the now-dead code path in the monolith. Over an extended period, sometimes years, the monolith is "strangled": its surface area shrinks as ever more traffic is diverted away, until it either disappears entirely or remains only as a thin core of capabilities that genuinely belong together. ## Why not a full rewrite The reason teams reach for this instead of a full rewrite is **risk management**. A full rewrite requires the team to build a replacement system in parallel with an unchanging feature target while the old system keeps evolving under production demands — the "second system" has to hit a moving target, feature parity slips further away the longer the rewrite takes, and the eventual cutover is an all-or-nothing event where any missed edge case becomes a production incident with no easy partial rollback. The strangler fig pattern avoids all three problems: - each extraction is **small** enough to review and test properly; - each **ships independently**; - each **can be rolled back** by simply pointing the router back at the monolith. It also delivers value continuously — the business gets improvements throughout the migration instead of nothing until a distant "big bang" release — and lets the team learn from real production behavior on each slice before tackling the next, harder one. ## The trade-off The trade-off is that this incrementality is bought with sustained complexity and duration. For the length of the migration, the organization runs two systems side by side: - two deployment pipelines; - two sets of dashboards and alerts; - and a routing layer that is itself a new piece of critical infrastructure requiring its own reliability engineering. ## Failure modes 1. **Skipping the delete step.** Teams must also maintain discipline about actually deleting old code paths after each extraction stabilizes; skipping that step is the single most common failure mode in practice — the "monolith that never dies," where new services accumulate on top of a monolith that keeps all its old code running just in case, so total system complexity grows rather than shrinks. 2. **Silent behavioral divergence.** A second common failure: the new service's implementation subtly differs from the old one (different rounding, different validation rules, different error handling), and because both paths are live simultaneously, the difference shows up as inconsistent behavior depending on which system happened to handle a given request. 3. **A seam that is more coupled than expected.** A third failure mode is picking a seam that turns out to be more tightly coupled to the rest of the monolith than expected, resulting in a new service that must make many synchronous, chatty calls back into the monolith — trading one large deployable for a "distributed monolith" that has all the coordination problems of microservices with none of the independence benefits. ## Where it shows up In practice, the pattern is most associated with large, long-lived systems where downtime is unacceptable and the business can't simply pause feature development for a rewrite — e-commerce platforms, banking systems, and other systems with continuous uptime requirements are classic candidates. **Shopify** has publicly discussed modularizing its core Rails monolith along bounded-context lines as groundwork before extracting select components as separately deployed services, an approach very much in the strangler-fig spirit: prove the internal seams first, then peel pieces off once confidence and tooling exist.

  • What component actually performs the traffic routing in a strangler-fig migration, and where does it typically live?
    Usually a reverse proxy or API gateway (nginx, Envoy, an internal facade service, or even a router class inside the monolith itself) sitting in front of both systems, routing by URL path, header, or feature flag. Early on it can be as simple as a config-driven proxy; later, teams often adopt a proper API gateway or service mesh so routing rules, canary weights, and rollback are centrally managed rather than hand-rolled.
  • What's the most common reason a strangler-fig migration stalls indefinitely?
    Teams extract the easy, low-risk seams first and leave the hard, deeply coupled core (shared database, complex domain logic) for 'later,' and later never comes because there's no product pressure to finish an invisible infra project. The result is a permanent dual-system tax: two deployment pipelines, two on-call surfaces, and a monolith that's now harder to reason about because half its logic silently proxies elsewhere.
  • How do you decide the order in which to strangle capabilities out of the monolith?
    Prioritize by a mix of business value and technical risk: extract seams that are loosely coupled to the rest of the domain first to build confidence and tooling, then move toward higher-value or higher-change-rate areas once the pattern is proven. Avoid starting with the most tangled, highest-risk module — that's usually saved for once the team has route-and-rollback machinery battle-tested.

Like a strangler fig vine in the rainforest that germinates in the canopy and sends roots down around a host tree — over years the vine's own trunk forms and eventually the original tree inside dies and rots away, leaving the vine's shape behind. The application does the same: new service 'tissue' grows around monolith functionality until nothing of the original remains.

saying these in an interview costs you the question

  • Describes it as 'just rewrite the monolith piece by piece' without mentioning the routing/facade layer
  • Doesn't mention decommissioning old code — thinks migration is 'done' once new services exist alongside monolith
  • Assumes it requires microservices from day one rather than any incremental replacement strategy
  • Can't explain what routes traffic between old and new
  • Thinks strangler fig applies only to greenfield systems, not legacy ones

context

open as a page

What is branch-by-abstraction, and how does it let a team extract a module into a separate service without a long-lived feature branch or a risky flag-day cutover?

level: middleimportance: must knowfreq 65%

basics

~20 s

You put a stable interface in front of the old code, build the new implementation behind that same interface, switch callers over gradually, then delete the old code — all on the main branch, no long branches.

open as a page

When extracting a service from a monolith that shares one relational database, what are the main strategies for giving that service its own data, and what makes this the hardest part of a strangler-fig migration?

level: seniorimportance: must knowfreq 65%

basics

~20 s

You give the new service its own tables or database instead of sharing the monolith's, using techniques like splitting tables, dual writes, or event-based sync. It's hard because data has to stay consistent while both old and new code touch it.

open as a page

When picking the first capability to peel off a monolith with a strangler-fig migration, what makes a good 'seam,' and what characteristics of a module make it a poor first choice?

level: middleimportance: should knowfreq 50%

basics

~20 s

A good first piece to split off is one that talks to the rest of the system through few, well-defined connections and doesn't share much data — so pulling it out doesn't require touching everything else.

open as a page

How does a parallel run (shadow traffic / dark launch) work when cutting over from a monolith's old code path to a newly extracted service, and what specifically does it validate that a staging-environment test cannot?

level: seniorimportance: should knowfreq 55%

basics

~20 s

You send real production traffic to both the old and new systems at once, compare their outputs, but only use the old system's answer — so you catch bugs on real data without risking real users.

open as a page

Under what circumstances would you advise against using a strangler-fig migration to break up a monolith, and what would you recommend instead?

level: principalimportance: should knowfreq 35%

basics

~20 s

If the system is small, short-lived, or nobody actually needs it running while being rewritten, a slow incremental split can cost more than it saves — sometimes a full rewrite or just improving the monolith is smarter.

open as a page