skip to content

questions

18

In a monolith-to-microservices migration, what is the strangler fig pattern and why do teams use it instead of a full rewrite?

level: juniorimportance: must knowfreq 70%

answer

  1. vine strangling host tree
  2. routing facade / proxy layer
  3. incremental not big-bang
  4. shrink monolith leaf by leaf
  5. reversible per-slice cutover

basics

~10 s

You build new features/services alongside the old monolith and slowly route traffic to them, until the old system does nothing and can be deleted. It avoids one risky big-bang rewrite.

solid answer

~50 s

The strangler fig pattern, named after a vine that grows around a host tree and eventually replaces it, is an incremental migration strategy: instead of rewriting a monolith in one big-bang release, you place a routing layer (proxy, facade, or gateway) in front of it, then extract one capability at a time into a new service, redirecting matching traffic to the new service while everything else still flows to the monolith. Each extraction is small, independently deployable, and reversible by re-routing back. Over months or years, the monolith shrinks as more traffic is 'strangled' away, until it either disappears or remains as a thin legacy core. Teams prefer it over a full rewrite because it de-risks delivery — you ship value continuously, get production feedback on each slice, and avoid the classic 'rewrite that never ships' trap where old and new systems must be kept at feature parity for years.

go deeper

for a junior

Should know the one-sentence idea — incremental replacement with a routing layer, safer than a rewrite — and be able to say why a big-bang rewrite is risky.

for a middle

Should be able to describe the routing facade concretely (URL path matching, feature flags) and walk through extracting one endpoint end to end.

for a senior

Should discuss the failure mode of stalled migrations, decommissioning discipline, and how to sequence extraction by coupling/risk.

for a principal

Should connect the pattern to organizational risk management — funding multi-year migrations, setting decommission SLAs/exit criteria, and knowing when strangler-fig is the wrong tool entirely.

## Where the name comes from The strangler fig pattern is named after tropical strangler fig vines that germinate high in a host tree's canopy, send roots down to the ground, and gradually envelop the host with their own growing trunk; over years the vine becomes self-supporting and the original tree, deprived of light and space, dies and rots away, leaving the vine standing in its shape. **Martin Fowler** applied this image in 2004 to describe an incremental system-replacement strategy, and it has since become one of the standard playbooks for moving a monolith toward a microservices (or any new) architecture without a single risky cutover. ## The mechanics Mechanically, the pattern starts by placing a **routing layer** — a reverse proxy, API gateway, or an internal facade inside the application — in front of both the existing monolith and (eventually) the new services being built. Initially, 100% of traffic flows through this layer straight to the monolith unchanged. The team then picks one capability (a URL path, a set of endpoints, a bounded business capability) and builds a new implementation of it as an independent service. The routing layer is updated to send requests matching that capability to the new service instead, while every other request continues to the monolith exactly as before. This is repeated capability by capability: 1. pick a **seam**; 2. build the replacement; 3. redirect matching traffic; 4. validate; 5. and — critically — **delete** the now-dead code path in the monolith. Over an extended period, sometimes years, the monolith is "strangled": its surface area shrinks as ever more traffic is diverted away, until it either disappears entirely or remains only as a thin core of capabilities that genuinely belong together. ## Why not a full rewrite The reason teams reach for this instead of a full rewrite is **risk management**. A full rewrite requires the team to build a replacement system in parallel with an unchanging feature target while the old system keeps evolving under production demands — the "second system" has to hit a moving target, feature parity slips further away the longer the rewrite takes, and the eventual cutover is an all-or-nothing event where any missed edge case becomes a production incident with no easy partial rollback. The strangler fig pattern avoids all three problems: - each extraction is **small** enough to review and test properly; - each **ships independently**; - each **can be rolled back** by simply pointing the router back at the monolith. It also delivers value continuously — the business gets improvements throughout the migration instead of nothing until a distant "big bang" release — and lets the team learn from real production behavior on each slice before tackling the next, harder one. ## The trade-off The trade-off is that this incrementality is bought with sustained complexity and duration. For the length of the migration, the organization runs two systems side by side: - two deployment pipelines; - two sets of dashboards and alerts; - and a routing layer that is itself a new piece of critical infrastructure requiring its own reliability engineering. ## Failure modes 1. **Skipping the delete step.** Teams must also maintain discipline about actually deleting old code paths after each extraction stabilizes; skipping that step is the single most common failure mode in practice — the "monolith that never dies," where new services accumulate on top of a monolith that keeps all its old code running just in case, so total system complexity grows rather than shrinks. 2. **Silent behavioral divergence.** A second common failure: the new service's implementation subtly differs from the old one (different rounding, different validation rules, different error handling), and because both paths are live simultaneously, the difference shows up as inconsistent behavior depending on which system happened to handle a given request. 3. **A seam that is more coupled than expected.** A third failure mode is picking a seam that turns out to be more tightly coupled to the rest of the monolith than expected, resulting in a new service that must make many synchronous, chatty calls back into the monolith — trading one large deployable for a "distributed monolith" that has all the coordination problems of microservices with none of the independence benefits. ## Where it shows up In practice, the pattern is most associated with large, long-lived systems where downtime is unacceptable and the business can't simply pause feature development for a rewrite — e-commerce platforms, banking systems, and other systems with continuous uptime requirements are classic candidates. **Shopify** has publicly discussed modularizing its core Rails monolith along bounded-context lines as groundwork before extracting select components as separately deployed services, an approach very much in the strangler-fig spirit: prove the internal seams first, then peel pieces off once confidence and tooling exist.

  • What component actually performs the traffic routing in a strangler-fig migration, and where does it typically live?
    Usually a reverse proxy or API gateway (nginx, Envoy, an internal facade service, or even a router class inside the monolith itself) sitting in front of both systems, routing by URL path, header, or feature flag. Early on it can be as simple as a config-driven proxy; later, teams often adopt a proper API gateway or service mesh so routing rules, canary weights, and rollback are centrally managed rather than hand-rolled.
  • What's the most common reason a strangler-fig migration stalls indefinitely?
    Teams extract the easy, low-risk seams first and leave the hard, deeply coupled core (shared database, complex domain logic) for 'later,' and later never comes because there's no product pressure to finish an invisible infra project. The result is a permanent dual-system tax: two deployment pipelines, two on-call surfaces, and a monolith that's now harder to reason about because half its logic silently proxies elsewhere.
  • How do you decide the order in which to strangle capabilities out of the monolith?
    Prioritize by a mix of business value and technical risk: extract seams that are loosely coupled to the rest of the domain first to build confidence and tooling, then move toward higher-value or higher-change-rate areas once the pattern is proven. Avoid starting with the most tangled, highest-risk module — that's usually saved for once the team has route-and-rollback machinery battle-tested.

Like a strangler fig vine in the rainforest that germinates in the canopy and sends roots down around a host tree — over years the vine's own trunk forms and eventually the original tree inside dies and rots away, leaving the vine's shape behind. The application does the same: new service 'tissue' grows around monolith functionality until nothing of the original remains.

saying these in an interview costs you the question

  • Describes it as 'just rewrite the monolith piece by piece' without mentioning the routing/facade layer
  • Doesn't mention decommissioning old code — thinks migration is 'done' once new services exist alongside monolith
  • Assumes it requires microservices from day one rather than any incremental replacement strategy
  • Can't explain what routes traffic between old and new
  • Thinks strangler fig applies only to greenfield systems, not legacy ones

context

open as a page

What is Conway's Law, and why does it matter when a company splits a monolith into microservices?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Conway's Law says the software you build ends up shaped like the teams that built it. If teams don't talk much, the code splits along those same lines. So when moving to microservices, you have to think about team structure, not just code structure.

open as a page

What do you gain and what do you give up when you split a single application into multiple independently deployable microservices, versus keeping it as one monolith?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A monolith is one big program deployed as a single unit; microservices split it into many small programs that talk over the network. Microservices let teams work and scale independently, but add network calls, more infrastructure, and coordination overhead.

open as a page

What is branch-by-abstraction, and how does it let a team extract a module into a separate service without a long-lived feature branch or a risky flag-day cutover?

level: middleimportance: must knowfreq 65%

basics

~20 s

You put a stable interface in front of the old code, build the new implementation behind that same interface, switch callers over gradually, then delete the old code — all on the main branch, no long branches.

open as a page

What is the 'inverse Conway maneuver', and what concrete steps would a company take to apply it before a microservices migration?

level: middleimportance: must knowfreq 55%

basics

~20 s

It's deliberately reorganizing your teams to match the architecture you want, instead of letting the architecture accidentally end up looking like whatever teams you already have. You design the org chart on purpose so the software naturally comes out the way you planned.

open as a page

What is a 'modular monolith,' and how does it try to capture the team-autonomy benefits of microservices while avoiding the distributed-system costs of splitting services over the network?

level: middleimportance: must knowfreq 75%

basics

~20 s

A modular monolith is one deployable app internally split into strict modules with clear boundaries, so teams work independently like in microservices, but everything still runs and deploys as one process, so there's no network overhead or distributed failure modes.

open as a page

What is the 'monolith-first' heuristic often attributed to Martin Fowler, and what's the reasoning for starting a new system as a monolith even when the team expects it will eventually need to scale like microservices?

level: middleimportance: must knowfreq 70%

basics

~20 s

Start new projects as one simple app (a monolith) instead of many small services, because early on you don't yet know where the real boundaries between parts of your system should be — splitting too soon locks in guesses that are often wrong.

open as a page

When extracting a service from a monolith that shares one relational database, what are the main strategies for giving that service its own data, and what makes this the hardest part of a strangler-fig migration?

level: seniorimportance: must knowfreq 65%

basics

~20 s

You give the new service its own tables or database instead of sharing the monolith's, using techniques like splitting tables, dual writes, or event-based sync. It's hard because data has to stay consistent while both old and new code touch it.

open as a page

What does the 'you build it, you run it' operating model mean, and what conditions have to be in place for it to actually improve reliability rather than just burning out engineers?

level: seniorimportance: must knowfreq 65%

basics

~20 s

The team that writes a service's code is also the team that gets paged when it breaks in production, instead of handing it off to a separate ops team. The idea is that owning the pain of your own bugs makes you write more reliable code. It only works well if that team also has good tools, reasonable on-call load, and real authority to fix root causes.

open as a page

What does 'single-team ownership' of a microservice mean in practice, and what breaks when two teams share ownership of the same service?

level: seniorimportance: must knowfreq 60%

basics

~20 s

One team should be fully responsible for a service: writing its code, deciding its roadmap, and getting paged when it breaks. If two teams share a service, decisions get slow, nobody's fully accountable, and the code quality suffers because there's no single owner enforcing consistency.

open as a page

Concretely, what does the 'distributed system tax' mean when a team moves from in-process function calls inside a monolith to network calls between microservices? Walk through the specific new failure modes and operational mechanisms this forces onto the team.

level: seniorimportance: must knowfreq 80%

basics

~20 s

When code calls another service over the network instead of directly in the same program, that call can now be slow, drop partway through, or never come back — things that can't happen with a normal function call. Handling that safely requires extra machinery like timeouts, retries, and tracing.

open as a page

When picking the first capability to peel off a monolith with a strangler-fig migration, what makes a good 'seam,' and what characteristics of a module make it a poor first choice?

level: middleimportance: should knowfreq 50%

basics

~20 s

A good first piece to split off is one that talks to the rest of the system through few, well-defined connections and doesn't share much data — so pulling it out doesn't require touching everything else.

open as a page

In the Team Topologies model, what is the difference between a stream-aligned team and a platform team, and why do most microservices orgs need both?

level: middleimportance: should knowfreq 45%

basics

~20 s

A stream-aligned team builds and runs features for one part of the business, end to end. A platform team builds internal tools (like deployment or infra) that other teams use, so those teams don't all have to solve the same problem separately. Most companies need both so feature teams can move fast without reinventing plumbing.

open as a page

How does a parallel run (shadow traffic / dark launch) work when cutting over from a monolith's old code path to a newly extracted service, and what specifically does it validate that a staging-environment test cannot?

level: seniorimportance: should knowfreq 55%

basics

~20 s

You send real production traffic to both the old and new systems at once, compare their outputs, but only use the old system's answer — so you catch bugs on real data without risking real users.

open as a page

A team has decomposed its product into a dozen microservices, but every release still requires deploying all twelve together, and the checkout service is effectively down whenever any of three specific other services is down. What anti-pattern does this describe, and what does it tell you about the original decomposition?

level: seniorimportance: should knowfreq 65%

basics

~20 s

This is a 'distributed monolith': many separate services that are still forced to move and fail together like one big app, except now with all the extra network overhead. It means the services were split in the wrong places — they're not actually independent.

open as a page

Under what circumstances would you advise against using a strangler-fig migration to break up a monolith, and what would you recommend instead?

level: principalimportance: should knowfreq 35%

basics

~20 s

If the system is small, short-lived, or nobody actually needs it running while being rewritten, a slow incremental split can cost more than it saves — sometimes a full rewrite or just improving the monolith is smarter.

open as a page

A CTO proposes reorganizing every team around microservices boundaries company-wide within one quarter to 'fix' Conway's Law problems. As a principal engineer, what would make you push back on the timing or scope of that plan?

level: principalimportance: should knowfreq 30%

basics

~20 s

Reorganizing everyone at once, fast, is risky: people lose their teammates and context all at the same time, work slows down company-wide, and if the target architecture guess turns out wrong, you've paid the cost twice. It's usually better to reorganize in stages, starting where the pain is worst.

open as a page

How does Conway's Law interact with the decision to split a system into microservices, and under what organizational conditions does that split actually deliver the promised autonomy and scaling benefits versus just adding cost?

level: principalimportance: should knowfreq 55%

basics

~20 s

Conway's Law says a company's software ends up mirroring how its teams are organized. Splitting into microservices only pays off if the team structure already matches the split — separate, mostly-independent teams each fully owning one service end to end. If teams are shared across services or the split doesn't match reporting lines, you just add network overhead on top of the same coordination problems.

open as a page