skip to content

Messaging & Integration

Patterns for asynchronous integration and for the gateway in front of your services: competing consumers, priority queue, pipes and filters, gateway aggregation, offloading and routing, anti-corruption layer and saga.

part ofResilience & cloud-native patternsoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

When a new service needs to integrate with a legacy system that has messy, inconsistent naming and data structures, what problem does putting an Anti-Corruption Layer between them solve?

level: juniorimportance: must knowfreq 70%

answer

  1. translation facade
  2. foreign model quarantine
  3. facade + adapters + translators
  4. prevents legacy leak
  5. paired with strangler fig

basics

~10 s

An Anti-Corruption Layer is a translator between two systems. It converts the old system's weird data and terms into clean data your new system understands, so the mess doesn't spread into your new code.

solid answer

~30 s

An Anti-Corruption Layer (ACL) is a translation facade at the boundary between a new system and a legacy or external one. It converts requests and responses between the two models so the new system's domain code never has to know the legacy system's field names, status codes, or quirks. Without it, those foreign concepts leak into the new codebase and every consumer becomes coupled to legacy shapes and bugs. The ACL isolates that translation in one place, so the new domain model stays clean while the legacy system keeps its own internal shape untouched.

go deeper

for a junior

Should be able to say, in plain terms, that an ACL is a translation boundary that keeps a legacy system's mess out of the new code, and give a rough real-world analogy.

for a middle

Should be able to name the concrete pieces (facade, adapter, translator/mapper) and explain, with a small example, how a request crosses the boundary and gets converted.

for a senior

Should discuss the maintenance and latency cost of the layer, and describe at least one production failure mode (leaky translation, translation drift) they'd guard against with tests or contract checks.

for a principal

Should place the ACL in a broader modernization strategy (e.g. strangler fig migration), discuss when the cost isn't justified, and describe how ownership and lifecycle of the layer is managed as the legacy system is eventually retired.

## Two mental models to reconcile Every integration between two systems has to reconcile two different mental models: the vocabulary, data shapes, and invariants each side uses to represent the same real-world concepts. When two systems are built by different teams, at different times, or joined through a merger or acquisition, their models rarely agree - one side might: - call a customer a 'party', - encode status as a three-letter legacy code instead of an enum, - represent money as an untyped string with an implicit currency. An **Anti-Corruption Layer (ACL)** is a dedicated translation boundary - typically implemented as a facade backed by adapters and mapper/translator components - placed between the 'clean' new system and the 'foreign' legacy or external one. Every inbound and outbound call crosses through this facade, which converts the foreign representation into the new system's own domain model (and back again), so that no legacy-specific concept, naming, or invariant ever appears inside the new codebase. ## The moving parts Mechanically, the ACL is usually built from three cooperating pieces. 1. A **facade** or gateway interface exposes operations phrased in the new system's own ubiquitous language (e.g. `getOrder(orderId): Order`), so consumers on the clean side never see the legacy contract directly. 2. Behind that facade sit **adapters** that know how to actually talk to the legacy system - a SOAP client, a raw JDBC connection to an old database, a screen-scraper, or a vendor SDK. 3. Between the facade and the adapters sit **translators** (mappers) that convert the legacy DTOs, field-by-field and rule-by-rule, into the new system's domain objects, resolving mismatches such as different enumerations, units, null semantics, or missing fields that must be defaulted or looked up elsewhere. In an asynchronous or event-driven setting, the same idea applies to message payloads: an ACL can subscribe to a legacy system's event stream, translate each message into the new system's event schema, and republish it, so downstream consumers only ever see the new format. ## Why the pattern exists The reason this pattern exists is to protect the integrity of the new system's domain model. If a team lets legacy concepts leak straight through - passing legacy status codes into business logic, embedding legacy field names in APIs, or copying legacy validation quirks - the corruption spreads: - every consumer of the new system now has to understand and defensively code around the legacy system's history, - any future change to the legacy system (or its eventual retirement) becomes a shotgun-surgery exercise touching dozens of call sites. Concentrating the translation logic in one boundary means the **blast radius** of a legacy quirk, bug, or format change is confined to the ACL's mappers, and the rest of the new system can be designed, tested, and reasoned about using only its own clean model. ## The trade-off The trade-off is real cost on the other side of that ledger. - **Extra code.** The ACL is extra code that has to be written, tested, and kept in sync with both models as they evolve - a genuine double-maintenance burden, since a schema change on either side can silently break a translator. - **An extra hop.** It also adds a network or in-process hop, which shows up as latency and, at scale, as another component that needs its own monitoring, retries, and failure handling. - **A second legacy system.** If a team is not disciplined, the ACL itself can accrete unrelated business logic over time and become a second legacy system - a shared, poorly-owned service that nobody wants to touch, defeating its own purpose. ## Production failure modes Production failure modes cluster around a few recurring shapes. 1. The most common is a **'leaky' ACL**, where under time pressure a legacy field, status code, or naming convention is passed through untranslated because 'we'll fix it later' - and it never gets fixed, so the corruption the layer was built to prevent happens anyway, just more slowly. 2. A second is **translation drift**: the legacy system's API or database schema changes (a new status value, a renamed column) and the translator silently maps it incorrectly instead of failing loudly, producing subtly wrong domain objects that surface as confusing bugs far downstream. 3. A third is **performance degradation** from naive translation - for example, calling the legacy system once per item instead of batching, turning a single logical operation into an N+1 storm across the ACL boundary. ## A worked example, and where the pattern comes from A concrete, well-documented example is a **strangler-fig-style migration**: a company migrating an old on-premises ERP to a new cloud-native order-management service builds an ACL as a small facade service that calls the ERP's SOAP endpoints, translates its idiosyncratic status codes and XML shapes into the new service's `Order` and `StockLevel` domain types, and exposes only that clean contract to the new microservices. This exact use case - protecting a new system from a legacy or third-party model during incremental modernization - is why cloud architecture guidance (such as Microsoft's Azure Architecture Center) catalogs Anti-Corruption Layer as one of its cloud design patterns, distinct in emphasis from (but historically rooted in) its original definition as a Domain-Driven Design pattern for managing relationships between bounded contexts.

  • Does the ACL have to be a separate deployed service, or can it be a library inside the consuming application?
    It can be either - the pattern only requires that translation logic be isolated behind a facade, not that it run in its own process. For a single consuming service, an in-process module or library with a facade interface and mapper classes is often enough and avoids an extra network hop. A standalone service makes sense when multiple independent consumers need the same translation, or when the legacy protocol (e.g. mainframe screen-scraping) needs specialized infrastructure that shouldn't be duplicated in every consumer.
  • How is an Anti-Corruption Layer different from a plain API gateway or a generic adapter pattern?
    A generic adapter typically just changes an interface's shape (e.g., method signatures) without necessarily protecting a whole domain model, and a gateway usually focuses on routing, auth, and cross-cutting concerns rather than semantic translation. An ACL is narrower in intent: its job is specifically to prevent a foreign domain model's concepts from leaking into a protected domain model, which is why it always includes deliberate mapping/translation logic, not just protocol or routing changes.
  • What happens if you skip the ACL and integrate directly against the legacy system's API?
    Every consumer ends up coupled to the legacy system's naming, status codes, and bugs, so a later legacy change or retirement requires touching every call site instead of one boundary. It also makes the new domain model harder to keep clean, since developers under deadline pressure will often just pass legacy shapes straight through rather than modeling them properly.

Like a diplomatic interpreter at a summit: each side keeps speaking its own language, and the interpreter translates in both directions so neither side's phrasing or idioms contaminate the other's understanding.

saying these in an interview costs you the question

  • Can't explain what leaks into the domain if there's no ACL
  • Thinks the ACL is only about changing method names/HTTP paths
  • Assumes ACL must always be a separate microservice
  • No mention of the extra latency/maintenance cost
  • Confuses this pattern with a generic API gateway with no translation logic

context

open as a page

A client calls a REST endpoint that kicks off a report that takes 2 minutes to generate. Instead of holding the connection open, the server responds immediately with HTTP 202 Accepted and a URL the client can check later. What is this approach called, and why use it instead of just blocking the connection until the report is ready?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The server says 'got it, working on it' right away instead of making the client wait. It hands back a link the client can check later to see if the work is done. This keeps the connection short and frees the client to do other things while it waits.

open as a page

In a messaging system, what is the Competing Consumers pattern, and why might a team run several consumer instances reading from the same queue instead of just one?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Several worker programs all watch the same queue and grab the next job when free. This lets jobs get done faster (more workers = more parallel work) and keeps things running if one worker crashes.

open as a page

A mobile app screen needs data from five different backend microservices to render fully. Instead of having the app call all five services directly over the network, what does putting a gateway aggregation pattern in front of them do, and why does that help?

level: juniorimportance: must knowfreq 75%

basics

~20 s

It puts one gateway in front of many services. The gateway makes all the backend calls itself and sends the client one combined response, so the client needs only one request instead of many slow round trips.

open as a page

In an API gateway architecture, why would a team terminate TLS at the gateway instead of having every backend service handle its own TLS handshake, and what does that decision cost them?

level: juniorimportance: must knowfreq 70%

basics

~10 s

The gateway holds the certificate and does the encryption handshake once, so backend services get plain, already-decrypted traffic instead of every service needing its own certificate and crypto work.

open as a page

In a system made of several backend services, what does a routing gateway do and what problem does it solve for the client?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A routing gateway is one front door for many backend services. It reads each request (like the URL path) and forwards it to the right service, so callers only need to know one address, not every internal service location.

open as a page

In a message-processing system, what is the Pipes and Filters pattern, and why would you split a data-transformation workload into several small filter stages connected by pipes instead of writing one big function that does everything?

level: juniorimportance: must knowfreq 55%

basics

~20 s

You break a big job into small steps, each done by its own worker (a filter). The workers pass data to each other through a queue (a pipe). Each worker only does one thing, so you can reuse, replace, or scale each step on its own.

open as a page

What is the Priority Queue pattern in message-based systems, and what problem does it solve?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It lets a system handle important messages before less important ones — usually by giving each priority level its own queue — so urgent requests aren't stuck waiting behind a big batch of routine work.

open as a page

An order-processing flow spans three separate services (inventory, payment, shipping), each with its own database. Why can't you just wrap the whole flow in one ACID database transaction, and what pattern is typically used instead?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Each service owns its own database, so one transaction can't span them all. A saga runs a sequence of local transactions instead, one per service; if a later step fails, it runs compensating actions to undo the already-completed earlier steps.

open as a page

Concretely, what are the moving parts you'd build when implementing an Anti-Corruption Layer between a new service and a legacy backend, and how does a request flow through them?

level: middleimportance: must knowfreq 60%

basics

~20 s

You build three parts: a front door that speaks your new system's language, a connector that knows how to call the old system, and a converter that maps data between the two. A request goes in the front door, gets converted, sent to the old system via the connector, and the response gets converted back.

open as a page

When you design the status-polling endpoint for an async request-reply API — the resource a client hits after receiving HTTP 202 — what should it return across the job's lifecycle, from just-submitted to complete, in terms of status codes, body shape, and headers?

level: middleimportance: must knowfreq 60%

basics

~20 s

The status endpoint should say 'still working' while the job runs, then either hand back the finished result or an error once it's done, and it can include a hint about how long to wait before checking again.

open as a page

Why do Competing Consumers setups typically guarantee at-least-once message delivery rather than exactly-once, and what must consumer code do to stay correct under that guarantee?

level: middleimportance: must knowfreq 75%

basics

~20 s

The system can't perfectly promise a message is handled exactly one time — crashes and retries mean it might get handled twice. So the code that processes a message has to be written so doing it twice causes no harm, like setting a value instead of adding to it.

open as a page

In queue-based worker pools (e.g., Amazon SQS), what is a visibility timeout, and what goes wrong in production if it's set too short or too long relative to how long a consumer takes to process a message?

level: middleimportance: must knowfreq 65%

basics

~20 s

It's a timer that hides a message from other workers while one worker is handling it, so two workers don't do the same job. If the timer's too short, someone else grabs it early and it gets done twice. Too long, and a crash leaves it stuck.

open as a page

A junior engineer implements an aggregation gateway endpoint that calls a pricing service, an inventory service, and a reviews service, then merges the three JSON responses into one payload. In their code, each call uses a blocking HTTP client and is invoked one after another inside a single method. What is wrong with this implementation, and what would fix it?

level: middleimportance: must knowfreq 70%

basics

~20 s

Calling each service one after another means the gateway waits for all of them added together, which is slow. Fixing it means starting all three calls at roughly the same time and only waiting once, so the total wait is about as long as the slowest single call.

open as a page

When an API gateway offloads authentication so backend services don't each implement login/token-checking logic, what typically gets validated at the gateway, and what does the gateway pass downstream so a service still knows who is calling it?

level: middleimportance: must knowfreq 75%

basics

~20 s

The gateway checks the caller's token or API key once, and if it's valid, forwards the request to the backend with a header saying who the caller is, so the backend trusts that header instead of re-checking credentials itself.

open as a page

How does gateway routing make blue-green deployments and canary releases possible, mechanically?

level: middleimportance: must knowfreq 60%

basics

~20 s

The gateway can send a slice of traffic, or all of it after a flip, to a new version instead of the old one, just by changing routing rules. Clients never need to change since they call the same address.

open as a page

Concretely, what request attributes can a routing gateway use to pick a backend, and what's a practical difference between routing by URL path versus routing by a request header?

level: middleimportance: must knowfreq 65%

basics

~20 s

A gateway can look at the URL path, hostname, a header, or a query parameter to pick a backend. Path routing is visible in the URL; header routing is hidden, so it needs the client to cooperate.

open as a page

In a Pipes and Filters pipeline where each filter reads from an input queue and writes to an output queue, one filter is consistently slower than its neighbors. What happens to the pipeline as a result, and how would you fix it operationally?

level: middleimportance: must knowfreq 60%

basics

~20 s

The slow step's input line piles up with waiting work, like a checkout line behind a slow cashier. Fix it by giving that step more workers, making it faster, or limiting how fast work is fed into it.

open as a page

What are the two main ways to implement the Priority Queue messaging pattern, and what does each cost you compared to the other?

level: middleimportance: must knowfreq 75%

basics

~20 s

You can either make several separate queues (one per priority level) and check the important ones first, or use one queue on a broker that supports priority natively and lets it sort messages itself. The first is simpler to set up anywhere; the second is neater but ties you to that broker's limits.

open as a page

You're implementing a saga for an e-commerce checkout that touches Order, Payment, and Shipping services. Compare a choreography-based saga (services react to each other's events) with an orchestration-based saga (a central coordinator issues commands) for this flow - how does each work mechanically, and what are the trade-offs as the number of participating services grows?

level: middleimportance: must knowfreq 85%

basics

~20 s

In choreography, each service listens for events and decides what to do next itself - no central boss. In orchestration, one coordinator service tells each step what to do and waits for the result, like a conductor directing an orchestra.

open as a page

A saga step for a Payment service is 'charge customer $50.' What makes a good compensating transaction for this step, and why can compensating transactions rarely be an exact mirror-image undo of the original action?

level: middleimportance: must knowfreq 80%

basics

~20 s

A good compensating action produces the opposite business effect (like a refund for a charge), not a literal undo, because by the time you need to compensate, other systems may have already seen and acted on the original change.

open as a page

In an async request-reply system, a client's initial POST times out on the network before the client receives the 202 response, so the client doesn't know whether the server actually accepted the job. What should the API design do to prevent this from silently creating duplicate jobs, and what other failure modes does a production implementation of this pattern need to guard against?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Let the client attach a unique 'idempotency key' to its request; if it retries after a timeout, the server recognizes the same key and returns the existing job instead of starting a new one. Beyond that, watch out for stuck jobs, crashed workers, and clients polling forever.

open as a page

A platform team wants to add a gateway aggregation layer in front of every microservice so that all clients only ever talk to one aggregation endpoint. What are the costs of this approach, and in what situations is gateway aggregation the wrong tool?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Aggregation adds an extra hop, an extra thing to build, run, and version, and makes the whole response only as fast as the slowest call inside it. If services are already fast and nearby, or clients need very different, fine-grained data, forcing everything through one big aggregator can do more harm than good.

open as a page

A team moves rate limiting from each backend service into the shared API gateway in front of them. What does the gateway need to track to enforce limits correctly, and what breaks if the gateway is horizontally scaled to multiple instances without care?

level: seniorimportance: must knowfreq 65%

basics

~20 s

The gateway counts how many requests each client has made in a time window and blocks them once they go over the limit. If there are several gateway machines each counting on their own, a client can sneak past the real limit by spreading requests across them.

open as a page

What does gateway routing cost a system in exchange for the topology decoupling it provides, and when would you deliberately avoid introducing it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Every request now takes an extra hop through the gateway, adding latency, and the gateway itself must stay up and correctly configured or nothing behind it is reachable. Not worth it for a tiny system with one or two services.

open as a page

In a Pipes and Filters pipeline built on message queues, a filter reads a message, starts processing it, and crashes before acknowledging completion to the queue. What failure modes does this create for the pipeline, and how do you design filters to survive it?

level: seniorimportance: must knowfreq 65%

basics

~20 s

The queue doesn't know the work finished, so it hands that message to someone else too — meaning it might get processed twice. Filters need to be written so doing the same work twice causes no harm (like setting a value, not adding to it).

open as a page

A service uses two separate queues — high-priority and low-priority — for order processing, both backed by the same consumer pool and the same downstream database. During an incident, a bug causes low-priority messages to be reprocessed in a tight retry loop, and soon high-priority order confirmations start timing out too. What went wrong with the isolation design, and how would you fix it?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Even though the messages were split into two queues, they still shared the same workers and the same database — so when the low-priority ones went haywire, they used up all the shared capacity and slowed down the important ones too. True isolation means separating the resources, not just the queues.

open as a page

How does a saga's approach to consistency across a distributed order-processing workflow differ from two-phase commit (2PC), and what specific trade-offs (availability, isolation, complexity) would push a team to choose one over the other?

level: seniorimportance: must knowfreq 65%

basics

~20 s

2PC locks all the databases involved and either commits everyone at once or nobody, giving true atomicity but requiring everything to stay locked and online together. Sagas skip the locking - each step commits on its own - trading that strict atomicity for services staying independent and available, using compensations instead of guaranteed rollback.

open as a page

A staff engineer is deciding whether to introduce an Anti-Corruption Layer for a new microservice's integration with an external partner API. What are the concrete costs versus benefits they should weigh before committing to it?

level: middleimportance: should knowfreq 55%

basics

~20 s

It costs extra code, extra testing, and a bit of speed to build and maintain the translator. It pays off by keeping your own system clean and easy to change even if the outside system is messy or changes later.

open as a page

An aggregation gateway composes a checkout confirmation page by calling an order-status service, which is required to render the page at all, and a recommendations service for a 'you might also like' section, which is a nice-to-have. If the recommendations call times out, should the gateway fail the whole request or return something else? Walk through how you'd design that behavior.

level: middleimportance: should knowfreq 65%

basics

~20 s

Only fail the whole response if a truly required piece of data is missing. For optional extras like recommendations, return the page without that section, or with a placeholder, rather than blocking the whole page on a non-critical service.

open as a page

showing 1–30 of 52