skip to content

When extracting a service from a monolith that shares one relational database, what are the main strategies for giving that service its own data, and what makes this the hardest part of a strangler-fig migration?

level: seniorimportance: must knowfreq 65%

answer

  1. shared schema -> owned schema per service
  2. dual writes vs CDC
  3. sagas replace cross-service transactions
  4. foreign keys become network calls
  5. hardest, most irreversible step

basics

~20 s

You give the new service its own tables or database instead of sharing the monolith's, using techniques like splitting tables, dual writes, or event-based sync. It's hard because data has to stay consistent while both old and new code touch it.

solid answer

~50 s

Database decomposition is widely considered the hardest step of a strangler-fig migration because data has to remain consistent under concurrent access from both the old and new systems during the transition, and once split it usually can't be trivially undone. Common strategies: (1) identify table ownership and physically split schemas along bounded-context lines so the new service owns its tables outright; (2) for a transition period, use dual writes — the monolith writes to both its own tables and the new service, accepting a window of eventual consistency; (3) use change-data-capture (e.g., Debezium) to replicate data changes into the new service's store without dual-write code; (4) for cross-service transactions that used to be a single DB transaction, replace them with sagas or eventual consistency plus compensating actions. It's hard because foreign keys, joins, and ACID transactions that were 'free' in a shared schema must be redesigned as network calls or asynchronous events, and any mistake risks data corruption or lost updates.

go deeper

for a junior

Should understand at a high level that the new service needs its own data rather than reaching into the old database directly.

for a middle

Should describe dual writes and CDC as two approaches to keeping two data stores in sync during transition, and know why sharing one DB long-term is an anti-pattern.

for a senior

Should explain saga-based replacement of cross-service transactions, the partial-failure risk of application-level dual writes vs CDC, and why this step is harder to reverse than code routing.

for a principal

Should discuss sequencing the whole migration's data strategy, tooling investment (CDC pipelines, reconciliation jobs), and organizational decisions like accepting temporary eventual consistency as a product-level trade-off, not just an engineering detail.

## Why this is the hardest step Splitting a shared relational database is widely regarded as the hardest part of a strangler-fig migration, harder than the routing and code-extraction work, because: - data — unlike code — must remain correct and consistent under concurrent access from both the old and new systems throughout the transition; - a completed data split is much harder to reverse than a routing change. Before any split, a monolith and its new service typically share one schema; the new service's tables are entangled with everyone else's via foreign keys, joins, and transactions that the relational database enforces "for free." None of that comes for free anymore once ownership is split across two independently deployed systems with two separate databases. ## Establishing table ownership The first step is establishing clear **table ownership**: for each table, decide which service is now the single owner responsible for writing to it, following the bounded-context lines identified during seam analysis. Tables owned by the newly extracted service are physically moved into a schema (or a wholly separate database instance) that only that service can write to; any other service that needs that data must now go through the owning service's API rather than querying its tables directly — a discipline that's easy to state but requires real vigilance to enforce, since a straight-line query is always tempting during the transition. ## Keeping two stores consistent The central technical challenge is keeping the old and new stores consistent while both are live. Two main approaches exist. | Approach | How it works | How it behaves under failure | |---|---|---| | **Application-level dual writes** | have the application code explicitly write to both the monolith's original tables and the new service on every relevant operation | this is simple to reason about but fragile, because the two writes aren't atomic — if the application crashes or the network fails between the first write succeeding and the second completing, the two stores silently diverge with no automatic recovery | | **Change-data-capture (CDC)**, using a tool such as Debezium that reads a database's write-ahead log or binlog | instead streams every already-committed change as an event without any dual-write code in the application at all | because it reads from the single authoritative source of truth after commit, there's no window where a crash can leave the two stores in different states relative to what the application intended — it can at worst lag, not diverge | For this reason CDC is generally preferred over hand-rolled dual writes wherever it's available, though it adds its own operational component (a CDC pipeline) that must be run and monitored. ## When one transaction becomes a saga A second major challenge is what happens to operations that used to be a single **ACID transaction** spanning tables that are now owned by different services — for example, canceling an order and issuing a refund, which the monolith could do atomically in one commit. Once orders and payments are separate services with separate databases, that atomicity is no longer available; distributed two-phase-commit transactions across services are technically possible but rarely used in practice because they couple services tightly and reduce availability (all participants must be up and responsive). The standard replacement is a **saga**: a sequence of local transactions in each service, with an explicit compensating action defined for each step so the overall operation can recover from partial failure without a distributed lock. This trades strong, immediate consistency for eventual consistency, meaning the system genuinely can be observed in an intermediate state — the design has to account for and expose that rather than pretend it can't happen. ## Failure modes The database-split step's failure modes are more consequential in production than most other migration steps because they involve real user data. - **Dual writes that lack reconciliation.** Application-level dual writes that lack any reconciliation job routinely produce silent divergence that isn't discovered until a customer notices a missing order or a wrong balance, sometimes weeks later. - **Distributed transactions under load.** Attempting a distributed transaction across newly split services under load tends to surface as cascading timeouts and reduced availability, since the transaction can only complete as fast as its slowest participant and blocks resources on all of them meanwhile. - **Rollback after both systems have accepted writes.** And unlike a routing flag that can be flipped back instantly if a code-level extraction goes wrong, once two systems have independently accepted writes to the same logical record for any period of time, reconciling those writes on rollback may involve genuine, unrecoverable conflicts — which is why teams invest heavily in comparison/verification tooling before finalizing a split, treating it as the point of no return in the migration.

  • What is change-data-capture (CDC) and why is it often preferred over application-level dual writes when splitting a database?
    CDC (e.g., via a tool like Debezium reading a database's write-ahead/binlog) captures every committed change to a table and streams it as an event, without requiring the application to explicitly write to two places. It's preferred over hand-written dual writes because dual writes are prone to partial failure — the app can crash or the network can fail after writing to the first store but before the second, causing silent data divergence, whereas CDC reads from the authoritative committed log so it can't miss a change the primary database itself doesn't have.
  • What replaces a single ACID database transaction when a business operation now spans two services with separate databases?
    A saga — a sequence of local transactions in each service, each with a corresponding compensating action that undoes it if a later step fails, coordinated either by orchestration (a central saga coordinator) or choreography (each service reacts to the previous service's event). It trades strong consistency for eventual consistency, so the system may briefly be in an intermediate state, which the design must expose or tolerate rather than hide.
  • Why is database decomposition usually the point of no return in a strangler-fig migration, unlike most code-level extraction steps?
    Code-level routing changes can typically be reverted instantly by flipping a flag back, but once data has been split, backfilled, and both systems have been independently writing to their own stores for a period, reverting means reconciling divergent data, which may involve genuinely unrecoverable conflicts. That asymmetric reversibility is why teams invest heavily in dual-write/CDC verification and dry-run comparisons before finalizing a database split.

Like separating utilities in a duplex that used to share one water meter and one electrical panel: you can't just declare two houses — you have to physically run new pipes and wires, keep both households supplied while the work happens, and only disconnect the shared meter once you're sure the new lines actually work.

saying these in an interview costs you the question

  • Suggests the new service can just keep querying the monolith's database directly as a long-term solution
  • Doesn't mention the risk of dual-write failure (partial write to one store but not the other)
  • Assumes cross-service transactions can just use a distributed 2PC transaction with no downsides
  • No mention of sagas or eventual consistency replacing former single-DB transactions
  • Treats database splitting as equally reversible/low-risk as a routing change

context