skip to content

Service Decomposition & Boundaries

The hardest part of microservices is drawing the lines: how big a service should be, which business capability it owns, and where the integration seams sit. Get this wrong and every other problem in the area gets worse.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

28

Your team's new 'Orders' microservice needs to read data from a 20-year-old legacy inventory system that uses cryptic status codes like 'S3' and denormalized flat records. Instead of letting the Orders service call the legacy API directly and pass those codes around internally, the team builds a small module that sits between them and converts everything into the Orders service's own clean model before anything else touches it. What is this module called, and what problem does it solve?

level: juniorimportance: must knowfreq 65%

answer

  1. translator at the seam
  2. bounded context protection
  3. one-time translation cost
  4. leaky ACL anti-pattern
  5. strangler fig companion

basics

~20 s

It's an anti-corruption layer - a translator sitting at the boundary that converts a foreign or messy system's data/language into your own clean model, so ugliness on the other side never leaks into your code.

solid answer

~50 s

An anti-corruption layer (ACL) is an isolating translation layer placed at an integration boundary. It exposes an interface in the consuming service's own domain language, and internally converts requests and responses to and from the upstream system's model - whether that's a legacy monolith, a third-party API, or another team's service built around a different bounded context. The goal isn't just data mapping; it's protecting the integrity of your domain model so foreign naming, status codes, or workflow assumptions never leak into your business logic. Without an ACL, every consumer of the legacy system becomes coupled to its quirks, and those quirks propagate outward - making the new service almost as hard to change as the old one it was meant to replace. The ACL absorbs that instability at a single, well-tested seam instead of scattering translation logic everywhere.

go deeper

for a junior

Should recognize the basic shape - a translation module at the boundary - and give a plausible reason (avoid messy legacy code spreading). Not expected to discuss bounded contexts by name or trade-offs in depth.

for a middle

Should articulate the domain-model-protection rationale explicitly, know where the ACL typically lives in a service's code, and mention that it needs its own tests.

for a senior

Should discuss the leaky-ACL failure mode from firsthand experience, know when to extract the ACL into its own service vs keep it embedded, and reason about maintenance cost/ownership.

for a principal

Should connect this to migration strategy (e.g., strangler fig), organizational ownership of the seam, and be able to argue for or against building one given the specific upstream's stability and expected lifespan.

## The two pieces it is built from Mechanically, an ACL is built from two pieces: - a **facade/interface** that speaks the consuming service's own domain language, and - a **translator** (sometimes split into an adapter for protocol/transport concerns and a translator for semantic/model mapping) that does the actual conversion to and from the upstream system's shapes. Say the Orders service defines a domain concept `StockLevel` with a clean enum `AVAILABLE/LOW/OUT_OF_STOCK`. The legacy inventory system returns flat records with a field `stat_cd` holding values like 'S1', 'S2', 'S3', plus a dozen other fields Orders doesn't care about (warehouse bay codes, a mainframe timestamp format, etc.). The ACL's translator is the only piece of code in the whole system that knows the mapping from 'S3' to `LOW`; everywhere else in Orders, code works exclusively with `StockLevel`. Calls into the legacy system, and only those calls, pass through this seam; the rest of the codebase never imports a legacy client library, never parses a legacy response body, and never reasons in legacy vocabulary. ## Why it exists This exists because **bounded contexts** (a core DDD idea) assume each service owns a model tailored to its own problem, and that model degrades the moment foreign concepts are allowed to seep in undigested. Integration is where model corruption happens fastest: a legacy system was built under different constraints, at a different time, often by a different team with different naming conventions, invariants, and failure semantics. If a new service adopts those shapes wholesale — passing legacy status codes through its APIs, storing legacy IDs as primary keys, or embedding legacy validation quirks into its own logic — it inherits legacy debt without inheriting legacy context. The ACL is the deliberate refusal to let that happen: it draws a hard line at the boundary and pays a one-time translation cost so everything inside stays coherent, testable, and independently evolvable. ## What it costs The cost is real and worth naming honestly. - Every field the upstream system can produce has to be **mapped**, which means writing and maintaining translation code that has no business value of its own — it doesn't ship features, it just insulates. - That code needs its **own tests** (often the most valuable tests in the integration, because they catch upstream schema drift before it reaches consumers). - There is also a **runtime cost**: an extra hop or transformation step adds latency, however small, and an extra failure surface — the translator can throw on unexpected values the legacy system emits that were never seen in testing. - And there's an **ownership cost**: someone has to keep the ACL current as both sides evolve, which is easy to underfund once the initial migration project is 'done' and the team moves on. ## Failure modes In production, ACL failures show up in a few recognizable shapes. 1. The first is the **leaky ACL**: a developer under deadline pressure passes an upstream field straight through 'just this once' because mapping it properly is inconvenient, and within a few sprints half the legacy vocabulary is back inside the domain model — the layer exists on paper but not in practice. 2. The second is **drift**: the legacy system adds a new status code or changes an enum's meaning, the ACL's mapping table doesn't cover it, and the translator either throws, silently defaults to a wrong value, or (worse) passes the unknown value through untranslated, corrupting downstream logic in a way that's hard to trace back to its source. 3. The third is the **ACL becoming its own legacy system**: over years it accretes special cases, becomes the only thing anyone still understands about how the two systems relate, and nobody dares touch it — exactly the brittleness it was built to prevent, just moved one layer over. ## Where it shows up A textbook real-world case is the **strangler fig** migration pattern: a team peeling functionality off a legacy monolith into new microservices puts an ACL in front of the monolith so each new service can be built against a clean model from day one, while the monolith itself is incrementally decommissioned behind that seam. Another common instance is integrating with an enterprise system like SAP or a mainframe order-management system, where the vendor's data model is fixed and verbose; teams write a dedicated adapter service (sometimes literally called an 'SAP ACL' or 'legacy facade') whose only job is exposing a small, well-named REST or event contract that hides the vendor's field names, unit conventions, and workflow states from every other service in the organization.

  • Where in the codebase should the ACL physically live - inside the consuming service, or as a separate deployable?
    Both are valid; it depends on how many consumers share the same upstream and how much translation logic there is. A single consumer with modest mapping usually embeds the ACL as an internal module/package to avoid an extra network hop. When multiple services need the same translated view of a legacy system, or the mapping logic is heavy, extracting it into its own small service avoids duplicating translation logic and centralizes the one place that has to change when the legacy system changes.
  • How would you test an ACL differently from the rest of the service?
    The translator itself deserves focused unit tests per mapping rule, including edge cases like unknown/unexpected upstream values, so schema drift is caught immediately rather than surfacing as a mystery bug three layers downstream. Contract tests against the upstream system's actual response shapes (or a recorded fixture of them) are also valuable, since the translator's correctness depends entirely on assumptions about a system you don't control.
  • What's the difference between an ACL and a plain DTO mapper?
    A DTO mapper is often just field-for-field data reshaping with no protective intent - it can still leak foreign semantics if the mapper reuses the source's enums or naming. An ACL is a deliberate architectural boundary: it enforces that the translated output is expressed entirely in the consumer's own domain vocabulary and invariants, and it's usually paired with an explicit decision about what NOT to expose from upstream at all.

Like a customs/immigration checkpoint at a border: goods and people from the other country get inspected, re-documented, and converted to local standards before they're allowed to move freely inside - nothing crosses untranslated.

saying these in an interview costs you the question

  • Says ACL and DTO mapper are the same thing
  • Can't explain why passing the legacy status code straight through is a problem
  • Assumes the ACL only needs to handle the response shapes seen during development
  • Thinks the ACL must always be a separate microservice
  • No mention of protecting the domain model, only 'converting formats'

context

open as a page

When breaking a monolithic application into microservices, what does it mean to decompose 'by business capability' (e.g., Order Management, Inventory, Billing) rather than by technical layer (e.g., a separate UI service, a business-logic service, and a database-access service)? Why is the capability-based split generally preferred?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Group services around business functions like 'Orders', each owning its own logic and data end to end. Splitting by technical layer (UI service, logic service, DB service) instead forces every feature to touch several services at once.

open as a page

In a microservices architecture, what is a 'distributed monolith', and what's usually the first symptom a team notices?

level: juniorimportance: must knowfreq 75%

basics

~20 s

It's when you've split code into separate services but they're still so tightly tied together that you can't deploy or change one without touching the others - you get all the complexity of microservices with none of the independence.

open as a page

In microservices design, why do teams typically draw each service's boundary around a Domain-Driven Design 'bounded context' rather than around a database table or an org-chart team?

level: juniorimportance: must knowfreq 75%

basics

~20 s

A bounded context is the area of the business where a word like "Order" has exactly one meaning. Building a service per bounded context keeps each service's data model and vocabulary consistent, instead of forcing every team to agree on one giant shared meaning for every term.

open as a page

What is a 'nanoservice' anti-pattern, and why is a service that's too small often worse than a monolith?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A nanoservice is a service split so small (like one endpoint) that it does barely any real work but still pays the full cost of being a separate service - network calls, its own database, deployment pipeline, on-call rotation. That overhead usually costs more than it saves.

open as a page

You're designing the internal structure of an anti-corruption layer that sits between a 'Billing' microservice and a third-party payment provider's SDK. Concretely, what components would you put inside it, and how do you decide where the boundary between 'translated domain model' and 'raw provider model' sits?

level: middleimportance: must knowfreq 55%

basics

~20 s

Split it into an adapter (talks the provider's protocol/SDK) and a translator (converts its data into your own domain types). The boundary sits exactly where the provider's names, codes, or shapes would otherwise leak into your business logic.

open as a page

How do bounded contexts (from Domain-Driven Design) inform where you draw microservice boundaries, and what concrete symptoms show up when a service's boundary crosses into a neighboring bounded context?

level: middleimportance: must knowfreq 70%

basics

~20 s

A bounded context is where one business term has one consistent meaning (e.g., 'Customer' differs between Billing and Support). Match each service to one context; crossing that boundary causes inconsistent meanings and constant translation code between services.

open as a page

Walk through how you'd use the strangler fig pattern to extract one capability — say, Inventory Management — out of a monolith into its own service, without a big-bang rewrite. What are the concrete steps, and how is traffic routed during the transition?

level: middleimportance: must knowfreq 65%

basics

~20 s

Build the new Inventory service beside the old monolith, redirect traffic bit by bit through a router. Once the old code sees no traffic, delete it — like a strangler fig vine replacing host tree.

open as a page

Why does having multiple 'microservices' read and write the same shared database (or even the same tables) recreate the coupling problems of a monolith?

level: middleimportance: must knowfreq 80%

basics

~10 s

The database schema becomes a hidden shared contract - if one service changes a table, every other service touching that table can break, so you can't change or deploy them independently anymore.

open as a page

When Service A calls Service B, which calls Service C, all over blocking synchronous HTTP, what is 'temporal coupling', and how does it turn a single slow dependency into a wider outage?

level: middleimportance: must knowfreq 78%

basics

~20 s

Temporal coupling means all three services must be up and responding fast at the same moment for the request to succeed - if C gets slow, B waits on it, then A waits on B, and the slowdown ripples backward through the whole chain.

open as a page

When mapping Domain-Driven Design aggregates onto microservice boundaries, why must a single aggregate always live entirely within one service, and what breaks if you split one across two services?

level: middleimportance: must knowfreq 65%

basics

~20 s

An aggregate is a group of objects that must stay consistent together, like an order and its line items. If you split it across two services, you can't guarantee both parts update together, so the data can end up in an invalid, inconsistent state.

open as a page

A team notices that a single request to their order-checkout API triggers 15 synchronous calls across 6 microservices before it can respond. What is this smell usually called, what causes it at the granularity level, and how would you address it?

level: middleimportance: must knowfreq 75%

basics

~20 s

This is called 'chatty' service communication - too many small back-and-forth network calls needed to do one piece of work. It usually means services were split too finely along lines that don't match how work actually flows, so simple operations require calling many of them in sequence. Fixes include combining related services, batching calls, or having one service own more of the workflow.

open as a page

What is the 'entity service' (a.k.a. CRUD-service) anti-pattern when decomposing a system into microservices, and why does it undermine the benefits microservices are supposed to deliver?

level: middleimportance: must knowfreq 65%

basics

~20 s

An entity service is a microservice built around one database table (like 'Customer' or 'Order') that just exposes create/read/update/delete endpoints for that table, with no real business logic. It looks like a clean split by 'noun,' but it pushes all the actual decision-making into whichever caller uses it, so business logic ends up duplicated or scattered across many services instead of owned by one.

open as a page

A company has split its monolith into 15 separately-deployed services, but every release still requires deploying most of them together in a fixed order, and a single slow service brings down request chains everywhere. What is this failure mode called, and what specific decomposition mistakes typically cause it?

level: seniorimportance: must knowfreq 75%

basics

~10 s

This is a distributed monolith: separate services still tightly coupled (shared code, chatty synchronous calls, or a shared database), so you can't deploy them independently — all of microservices' overhead, none of the benefit.

open as a page

Two microservices, each aligned to its own bounded context, need to integrate — one is the 'upstream' source of a concept, the other 'downstream' consumes it. What DDD context-mapping patterns describe how the downstream service should protect its own model from the upstream one, and when would you pick each?

level: seniorimportance: must knowfreq 55%

basics

~20 s

When one service depends on another's data, you can either just copy the upstream's model as-is (fast but fragile), or build a small translation layer that converts it into your own local model (more work but protects you from their changes). DDD calls these patterns things like "Conformist" and "Anti-Corruption Layer."

open as a page

How do cohesion and coupling act as opposing forces when you're deciding how large or small to make a microservice, and what heuristics help you find the right boundary?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Cohesion means things that change together and belong together should live in the same service. Coupling means how tangled services are with each other. You want high cohesion inside a service and low coupling between services. Getting the size right means grouping things that truly belong together, and not more.

open as a page

A platform team owns a 'Pricing' service used by a dozen other teams. Instead of leaving each consuming team to build its own anti-corruption layer against Pricing's internal model, the platform team publishes a versioned, documented API contract that Pricing commits to keeping stable, and consumers integrate against that contract instead of Pricing's internals. What pattern is the Pricing team applying, and how does it change - or reduce the need for - the ACLs each consumer would otherwise build?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The Pricing team is running an Open-Host Service with a Published Language: a stable, documented public contract instead of exposing its raw internal model. Consumers can then integrate against that contract directly, needing a much thinner ACL - or sometimes none - because the translation work already happened once, upstream.

open as a page

When you decompose a monolith's single database into per-service data stores, how do you decide which service owns a given piece of data, and when is it appropriate to replicate that data into another service versus having that service call the owner synchronously on every read?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Give each piece of data one owning service that's the source of truth for writes. Others call the owner for occasional fresh reads, or keep a replicated copy, updated via events, for frequent reads or to stay available if the owner is down.

open as a page

Without reading any code, what deployment and operational signals would tell you a 'microservices' system is actually a distributed monolith?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Look at the release calendar and incident history: if services almost always deploy together, or one service's outage always drags others down with it, that's the tell - you don't need to read the code to see the coupling.

open as a page

You inherit a 'microservices' system that's really a distributed monolith - shared database, synchronous call chains, lockstep deploys. Walk through a concrete plan to migrate it toward true service autonomy.

level: seniorimportance: should knowfreq 65%

basics

~20 s

Figure out which service should really own each piece of data and behavior, move that ownership over step by step, replace direct cross-service database reads with API calls or events, and only split each piece of data into its own database once nothing else touches it directly anymore.

open as a page

The usual advice is 'one bounded context maps to one service.' Under what circumstances is it reasonable to run several bounded contexts inside a single service, or conversely to split one bounded context across several services, and what risk does each deviation carry?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Sometimes it's fine to put two related but minor business areas in one service if they're small and rarely change, to avoid running lots of tiny services. But splitting one business area's logic across several services usually causes trouble, because it breaks up something that should stay together.

open as a page

What are the concrete costs of a microservice architecture that's too coarse-grained ('macroservices' or mini-monoliths) versus one that's too fine-grained, and how do you decide which risk to accept for a given team and system?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Too coarse means services are basically small monoliths - hard to scale or deploy independently, one team's bug blocks everyone. Too fine means way too many tiny services - huge operational overhead, slow chatty calls, hard to reason about the whole system. The right size depends on team size, how independently pieces need to change, and how much operational complexity the team can actually handle.

open as a page

You're the architect reviewing a proposal to build a dedicated anti-corruption layer service in front of a newly launched, actively-co-developed internal 'Catalog' service that another team on the same roadmap owns and is happy to evolve jointly with you. The proposing team argues 'ACLs are always good practice.' How would you push back, and what would make you decide an ACL is NOT worth building here?

level: principalimportance: should knowfreq 35%

basics

~20 s

An ACL isn't free - it's worth it when you can't influence or trust the upstream model, e.g. legacy, third-party, or unstable systems. If you can jointly evolve the API with a cooperative team, negotiating a shared contract is usually cheaper and better than building a translation layer against a moving target.

open as a page

A platform team has split a product catalog capability into 12 tiny services (ProductName, ProductPrice, ProductImages, ProductReviews-summary, and so on), each with its own repo, pipeline, and on-call rotation. What's this anti-pattern usually called, what does it cost in practice, and how would you decide the right granularity instead?

level: principalimportance: should knowfreq 45%

basics

~20 s

This is over-decomposition, or 'nanoservices' — splitting so finely that running each piece (deploys, monitoring, network calls) costs more than the benefit. Right size is roughly 'one team, one business capability,' not 'one service per field.'

open as a page

Two years after a DDD-based microservice decomposition, a Pricing service and a Promotions service constantly deploy together, share a hidden 'effective price' concept neither fully owns, and any change to one breaks the other's tests. As the architect asked to fix this, what diagnostic would you run, and how would you decide whether to merge them, redraw their boundary, or introduce a formal translation layer?

level: principalimportance: should knowfreq 40%

basics

~20 s

When two services always break together, it usually means they were never really separate business areas to begin with, or the line between them was drawn in the wrong place. You'd look at what concept they secretly share, then either merge them back into one service, move the shared idea fully into one of them, or add a clear translation step between them.

open as a page

How should service granularity evolve over the lifetime of a system, and what concrete signals tell an engineering organization it's time to split a service apart or merge services back together?

level: principalimportance: should knowfreq 55%

basics

~20 s

Granularity isn't a one-time decision - as a system and its teams grow, the 'right' size for a service changes. You split a service when a real, measurable reason shows up (different scaling needs, a new team owning part of it, slow deploys from unrelated changes). You merge services back when they always change together and the split is just adding overhead without any benefit.

open as a page

A new event-driven 'Shipping' service needs data from a legacy system that only exposes a synchronous SOAP API with no webhooks or change feed. The team builds an ACL that polls the legacy SOAP API on a schedule and publishes translated domain events onto the Shipping service's event bus whenever it detects a change. What is this ACL doing beyond data-shape translation, and what new failure modes does that introduce compared to a simple request/response ACL?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It's translating not just data shapes but integration style - turning a synchronous poll-based system into an event stream. This adds new failure risks: missed or duplicate events, detection lag between polls, and the ACL now holding state (what it last saw) instead of being stateless.

open as a page

Is a distributed monolith always a mistake to fix immediately? Discuss when it can be a rational, temporary state, and how Conway's Law and team topology influence whether it persists.

level: principalimportance: nice to knowfreq 40%

basics

~20 s

Not always - during a migration, having some temporary shared coupling can be the safer path. It becomes a real problem when it's permanent and when the teams responsible for the coupled services aren't structured to fix it.

open as a page