Walk through how you'd use the strangler fig pattern to extract one capability — say, Inventory Management — out of a monolith into its own service, without a big-bang rewrite. What are the concrete steps, and how is traffic routed during the transition?
answer
- proxy/facade routes old vs new
- dual-write or CDC syncs data mid-transition
- migrate slice by slice, verify, expand
- delete dead code only after full cutover
- named after the strangler fig vine, coined by Martin Fowler
basics
~20 sBuild the new Inventory service beside the old monolith, redirect traffic bit by bit through a router. Once the old code sees no traffic, delete it — like a strangler fig vine replacing host tree.
solid answer
~50 sThe strangler fig pattern extracts a capability incrementally rather than rewriting the whole monolith at once. First, put a routing layer (reverse proxy, API gateway, or feature-flagged facade) in front of the monolith so all calls for the target capability pass through it. Build the new service to handle that capability, backed initially by the same data (via replication, dual-write, or CDC) so both old and new paths stay consistent. Then migrate call-by-call or endpoint-by-endpoint: route a slice of traffic (one endpoint, one tenant, one percentage) to the new service, verify correctness and monitor for divergence, and expand the slice. Once 100% of traffic for that capability goes to the new service and the monolith's old code path is unused, delete the dead code and cut over full data ownership. The monolith 'shrinks' one capability at a time instead of being replaced outright.
go deeper
Should understand the core idea — replace incrementally, not all at once — and be able to explain the fig-vine analogy.
Should be able to name the routing layer and data-sync approaches (dual-write/CDC) and sketch a rough step order.
Should reason about consistency risk during dual-write, rollback strategy for a partially migrated slice, and how to decide the cutover order.
Should plan a strangler-fig migration as a multi-quarter program: sequencing across teams, setting a hard decommission milestone to avoid indefinite dual-running, and defining success metrics beyond 'the new service exists.'
## The pattern and where its name comes from The **strangler fig pattern**, named by Martin Fowler after the strangler fig vine that germinates in a host tree's canopy and gradually grows roots down around the trunk until it can stand alone once the original tree decays, is a way to replace a piece of a monolith with a new service incrementally, keeping the system releasable and rollback-able at every step, rather than freezing feature work for a rewrite. ## The mechanism, step by step The mechanism has a repeatable shape. 1. **First**, you introduce a routing layer in front of the capability you're extracting — a reverse proxy, an API gateway, or a facade module inside the monolith itself — so that every call for that capability passes through a single point that can direct traffic to either the old code path or the new service. 2. **Second**, you build the new service (say, Inventory Management) to handle the capability, and you solve the data problem: since the monolith's database still holds the authoritative inventory data at the start, the new service needs a way to see consistent data, typically either dual-writes (the monolith writes to its own table and calls the new service, or vice versa) or change-data-capture tooling like `Debezium` that streams row-level changes out of the monolith's database into the new service asynchronously. 3. **Third**, you migrate traffic incrementally — one endpoint, one tenant, or one percentage slice at a time — verifying at each step that the new path produces correct results and doesn't regress latency or error rates, then widening the slice. 4. **Finally**, once all traffic for that capability flows through the new service and the old code path sees zero traffic, you delete the dead code in the monolith and cut the new service over to full, sole ownership of its data. ## Why not a big-bang rewrite This exists because the alternative — a big-bang rewrite where you build the whole new system in parallel and cut over on one release date — is one of the riskiest moves in software delivery. It compounds the risk of the rewrite itself (subtle behavior differences, missed edge cases) with the risk of an all-or-nothing release, and it typically freezes meaningful feature work on the old system for months while the rewrite catches up, sometimes called the 'second-system effect.' The strangler fig pattern converts that one large bet into many small, individually reversible bets: each traffic slice you migrate can be rolled back by flipping the router, without touching anything else, and the team keeps shipping unrelated features on the untouched parts of the monolith throughout. ## The trade-off The trade-off is that you run two implementations of the same capability simultaneously for the duration of the migration, which is real, ongoing cost: - the routing layer itself needs maintenance and observability; - the data-sync mechanism (dual-write or CDC) is new infrastructure with its own failure modes; - and engineers have to reason about which path a given request took when debugging. This period can last weeks to many months for a nontrivial capability, and every day it continues is a day of running effectively double the operational surface for that one piece of functionality. ## Failure modes - **Data drift** is the most common failure mode: dual-writes are not atomic across two independent systems, so a write that succeeds in the monolith's database but fails on the call to the new service leaves the two stores silently inconsistent, and without reconciliation tooling this compounds over time into hard-to-debug data-integrity bugs. - **The migration stalling permanently at 80-90% complete** is a second common failure: the easy, high-traffic, high-value slices get migrated first, and the long tail — legacy admin screens, obscure batch jobs, an old partner integration — never gets prioritized, so the 'temporary' routing facade and dual-write logic become permanent fixtures, and the team ends up maintaining two systems forever instead of one. ## Where it shows up A well-known real-world pattern of this is large e-commerce and SaaS platforms extracting a single high-traffic capability, such as pricing or checkout, from a legacy monolith by putting an edge proxy in front of the relevant routes and shifting traffic to a newly built service over a period of controlled canary releases, with **CDC** keeping the new service's read model in sync with the monolith's system-of-record tables until the new service is promoted to be the source of truth itself and the old code is deleted.
- How do you keep data consistent between the monolith's database and the new service's database while both are live during the migration?Common approaches are dual-writes (the monolith writes to both its own table and the new service, often via an event), or change-data-capture (CDC) tooling like Debezium that streams the monolith's database changes into the new service asynchronously. Dual-write is simpler to reason about but risks partial-failure inconsistency; CDC decouples the two systems better but adds operational complexity and eventual-consistency lag you must design around.
- What's the risk of routing traffic to the new service by percentage rather than by a clean boundary like one endpoint or one tenant?Percentage-based routing mixes old and new code paths for the same feature simultaneously, which is great for canary-style risk reduction but means both paths must stay correct and compatible with the same shared data during the overlap — any schema or business-rule drift between them causes different users to see inconsistent behavior. It also complicates rollback, since some requests already committed state through the new path.
- What causes a strangler fig migration to stall indefinitely instead of finishing?Usually the last 10-20% of traffic touches the least-used, most tangled edge cases (legacy admin tools, obscure report generators, batch jobs) that nobody wants to prioritize once the visible user-facing win is captured, so the team declares victory early and leaves a permanent facade layer routing between two systems. Without an explicit decommission milestone and code-deletion step tracked as part of the work, the 'temporary' proxy and dual-write logic become permanent technical debt.
Named after the strangler fig vine, which germinates in a host tree's canopy and grows roots down around the trunk, gradually taking over structural support until the original tree can rot away and the fig stands on its own — the new service grows around the monolith's functionality until the old code is no longer load-bearing and can be removed.
saying these in an interview costs you the question
- Describes 'build the entire new service first, then switch over on day one' as the strangler fig approach — that's big-bang, its opposite
- No mention of a routing or facade layer directing traffic between old and new implementations
- No plan for keeping data consistent between the two systems during the transition
- Treats the migration as done once the new service exists, without decommissioning the old code path