skip to content

You inherit a 'microservices' system that's really a distributed monolith - shared database, synchronous call chains, lockstep deploys. Walk through a concrete plan to migrate it toward true service autonomy.

level: seniorimportance: should knowfreq 65%

answer

  1. strangler fig, not big-bang
  2. single-writer first, then split DB
  3. events as CQRS across services
  4. contract tests replace the FK safety net
  5. lockstep-deploy rate = real success metric

basics

~20 s

Figure out which service should really own each piece of data and behavior, move that ownership over step by step, replace direct cross-service database reads with API calls or events, and only split each piece of data into its own database once nothing else touches it directly anymore.

solid answer

~50 s

Treat it as an incremental strangler-fig migration, not a rewrite. First, map data and behavior to bounded contexts to decide who should own what. Second, for each shared table, migrate to a single writer: make the owning service the only one issuing writes while adding a public API or event stream for what other services need. Third, replace other services' direct reads with calls to that API or with event-driven local read models, removing SQL dependencies one table at a time. Fourth, once no cross-service SQL remains against a table, physically split it into the owner's own database. In parallel, replace critical synchronous chains with async messaging where the business doesn't require immediate consistency, and use contract tests to keep teams safely decoupled during the transition. Sequence the smallest, most-isolated tables and services first, and track the rate of lockstep-deploy incidents as your success metric.

go deeper

for a junior

Should have a rough sense that you fix it gradually, one piece at a time, rather than all at once.

for a middle

Should describe single-writer-then-split-database as the sequence and mention replacing direct DB reads with an API or events.

for a senior

Should give the full incremental plan with concrete mechanics such as single writer, CQRS-style read models, and contract tests, plus circuit breakers on remaining sync calls, and name risk-reduction tactics like feature flags or dual writes.

for a principal

Should tie the technical migration to Conway's Law and team-ownership realignment, define success metrics like lockstep-deploy rate, and reason about sequencing and prioritization across a large, multi-team system.

## The overall strategy The right overall strategy is an **incremental strangler-fig migration**, not a big-bang rewrite. A big-bang rewrite risks months of effort with no incremental value delivered, a large batch of untested architectural decisions, and a single high-stakes cutover where problems surface all at once with no easy rollback path. Start instead with **domain modeling**: use domain-driven design's bounded-context mapping to assign clear ownership of each entity or table to exactly one service. This determines the target end state before any code is touched, so subsequent steps have a destination rather than being ad hoc. ## Migrating one shared table For each table currently written by more than one service, pick a target owner and migrate in stages. 1. **Step one** is establishing a single writer: route all writes through the owner, so other services' writes go through the owner's API instead of raw SQL, while everyone still reads directly for now. This is comparatively low-risk because uncoordinated writes are what actually corrupts invariants, while shared reads against a still-shared physical database are comparatively safe during the transition. 2. **Step two** replaces direct reads by other services with either synchronous calls to the owner's read API for low-volume, real-time needs, or an event stream the owner publishes, such as `OrderPlaced` or `OrderCancelled` events, that other services consume to build their own local, denormalized read models — essentially **CQRS** applied across service boundaries, which also removes real-time coupling on the read side. 3. **Step three**, once no other service issues SQL against that table, physically moves it to a database only the owner can reach, typically by revoking other services' database credentials, completing the split for that entity. ## The synchronous call chains In parallel, address synchronous call chains. Identify which calls sit on the critical, user-facing path and genuinely must remain synchronous, such as a payment authorization, versus which were made synchronous only because that was the default when the module was extracted from the monolith. Convert the latter to asynchronous messaging so producer and consumer are decoupled in time. For any synchronous call that must remain during the migration, add timeouts, circuit breakers, and bulkheads to bound the blast radius while coupling is still partially present. ## Safety nets Safety nets matter throughout the transition. - **Consumer-driven contract tests**, such as those built with a tool like Pact, replace the safety a shared database's foreign keys and constraints used to provide silently, catching a breaking API or schema change before deploy rather than in production. - **Feature flags and dual-write-then-verify patterns** let a table be cut over to its new owner gradually, with the ability to roll back the read path if inconsistencies are detected, rather than forcing a hard, irreversible cutover. The best success metric is not a subjective sense of 'cleaner architecture' but the shrinking rate of **lockstep-deploy incidents** and cross-service on-call escalations — a falling rate is direct evidence that autonomy has genuinely increased, not merely that services were relabeled. ## Sequencing and the org chart Sequencing and organizational alignment round out the plan. Prioritize the highest-value, lowest-risk splits first: tables unambiguously owned by one team with few external readers are quick wins that build momentum and organizational trust in the approach, while the tangled, heavily-shared core tables, often something like a shared users or accounts table, are best tackled later once the team has practiced the pattern on smaller cases. This ties back to **Conway's Law**: real autonomy requires the team structure — who's on-call, who reviews pull requests, who owns the roadmap — to match the new service boundaries, not just the code. A common failure mode is completing the technical split while a single team still owns both resulting services, which quietly reintroduces coordination costs through shared humans instead of shared code. Large, well-documented decompositions at companies like Amazon followed roughly this incremental-ownership-migration shape — moving data ownership and call patterns step by step behind a stable interface — rather than rewriting the system from scratch.

  • Why migrate to single-writer before splitting the database, rather than splitting immediately?
    Splitting the database first, while multiple services still need to write to the same logical data, forces you to solve distributed transactions or sagas and data synchronization immediately and all at once, which is high risk. Establishing a single writer first is lower-risk: reads can still hit the same physical database temporarily, so you decouple write ownership, the harder invariant-protecting part, before tackling physical separation.
  • What's the risk of doing a 'big-bang' rewrite instead of this incremental approach?
    A big-bang rewrite means months without shipping business value, a large batch of untested architectural decisions, and a high-stakes cutover where problems surface all at once in production with no easy rollback. Incremental strangling keeps the system releasable throughout and lets you validate each ownership boundary against real production traffic before moving to the next one.
  • How do consumer-driven contract tests help during this migration specifically?
    Once you break the shared database's implicit referential-integrity and schema guarantees, nothing stops a service's API or event schema from silently drifting. Consumer-driven contracts let each consumer declare the exact shape of data or interaction it needs, and the provider's CI fails if a change would break any registered consumer, restoring at the API layer the safety a foreign key used to give at the database layer.
  • What organizational change usually needs to accompany this technical migration?
    Team ownership boundaries need to be redrawn to match the new service boundaries, ideally one team owning each service end-to-end across code, on-call, and roadmap. If a single team still owns both sides of a 'split' data boundary, or if two teams jointly own one service, you've reintroduced coordination costs through shared people even after the technical split, per Conway's Law.

Like renovating a house room by room while people still live in it - you don't demolish everything at once; you seal off one room, redo its plumbing so it no longer shares a pipe with the rest of the house, confirm it works standalone, then move to the next room.

saying these in an interview costs you the question

  • Proposes a full rewrite or big-bang cutover as the remediation strategy
  • Splits the physical database before establishing single-writer ownership of the data
  • Doesn't mention contract testing or any replacement for the safety the shared schema used to provide
  • Ignores the org/team-ownership dimension entirely
  • Treats 'add more services' or 'add a service mesh' as a fix without addressing data ownership

context