skip to content

Messaging & Integration

Patterns for asynchronous integration and for the gateway in front of your services: competing consumers, priority queue, pipes and filters, gateway aggregation, offloading and routing, anti-corruption layer and saga.

part ofResilience & cloud-native patternsoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

For filters connected by pipes in a message-processing pipeline to be truly reusable and composable — droppable into different pipelines without rewriting them — what do they need to agree on, and what breaks that reusability in practice?

level: middleimportance: should knowfreq 40%

basics

~20 s

All the workers need to speak the same language for the data passing between them — same fields, same format. If one worker expects data shaped differently than the last worker sends it, they can't be swapped or reused.

open as a page

In a Priority Queue setup with separate high- and low-priority queues, strict 'always serve high-priority first' consumer logic can starve the low-priority queue during sustained high-priority load. Name two concrete techniques to prevent this, and explain how each works.

level: middleimportance: should knowfreq 60%

basics

~20 s

Instead of always doing important work first no matter what, you can either give a small guaranteed share of time to the less-important queue (like 1 in 10 turns), or make old low-priority messages count as more important the longer they wait, so they eventually get processed.

open as a page

In a cloud migration where a new set of microservices is being built to gradually replace an on-premises monolith, how does an Anti-Corruption Layer typically get deployed and operated as a cloud integration pattern, and how does that differ from its original description as a Domain-Driven Design concept?

level: seniorimportance: should knowfreq 45%

basics

~20 s

In cloud migrations, the ACL is often its own small deployed service (or gateway) that talks to the old on-prem system and the new cloud services, translating between them while the old system is gradually replaced. Originally the same idea was described as code inside a team's own codebase for keeping domain models from mixing.

open as a page

In a production system, an Anti-Corruption Layer sits in front of a flaky third-party API and has been running fine for months. What are the most likely ways this kind of layer fails or degrades in production, and how would you design against each one?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The translator can quietly mis-map data when the outside system changes without telling you, the extra layer can slow things down or become a single point of failure, and it can turn into a dumping ground for random logic if nobody owns it. Guard against these with tests, monitoring, and clear ownership.

open as a page

An async request-reply API can notify completion either by having the client poll a status endpoint, or by having the client register a webhook URL that the server calls back with the result. What are the operational trade-offs between these two delivery mechanisms, and when would you choose one over the other?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Polling means the client keeps checking back for updates; a callback (webhook) means the server calls the client back when the work is done. Polling is simpler to set up on the client side; callbacks are faster and less wasteful but need the client to run something reachable.

open as a page

Why does scaling out a queue with more competing consumer instances typically break global message ordering, and how would you preserve ordering for related messages (e.g., all events for one customer) while still processing the queue in parallel?

level: seniorimportance: should knowfreq 55%

basics

~20 s

When many workers pull from one line, a later job can finish before an earlier one if it lands on a faster worker. To keep related jobs in order, you route them all to the same one worker instead of letting any worker grab them.

open as a page

In a Competing Consumers setup, a single message keeps causing a processing exception no matter which consumer handles it, and it keeps getting redelivered — starving the queue for other work. How do you design around this 'poison message' problem?

level: seniorimportance: should knowfreq 60%

basics

~20 s

One bad message that always fails can get retried forever and clog things up. The fix is to count how many times it's failed, and after a limit, move it out of the main queue into a separate 'failed' queue so a human can look at it later, instead of retrying forever.

open as a page

An aggregation gateway fans out to five downstream services using a single shared thread pool. One of those services starts responding very slowly, though not erroring outright. What can happen to the gateway as a whole, and how would you isolate the damage?

level: seniorimportance: should knowfreq 55%

basics

~20 s

If all five calls share the same limited pool of workers, the slow service can hog enough workers that the gateway can't process calls to the other four healthy services either, so one bad dependency drags down everything. Fix it by giving each downstream its own separate pool or limit and a circuit breaker, so one slow service can only exhaust its own slice.

open as a page

A gateway is configured to cache GET responses and gzip-compress payloads on behalf of every backend service behind it. What has to be true of a response for gateway-level caching to be safe, and what commonly goes wrong when a team turns it on without checking?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The gateway can save a copy of a response and reuse it for later identical requests, saving backend work, but only if the response is truly the same for everyone who asks; if the response depends on who's asking (like personal data), caching it for all users is a serious bug.

open as a page

In a production system, what are the specific ways a gateway's routing configuration goes wrong, and how does each one actually show up to an on-call engineer?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Routing rules can overlap and steal each other's traffic, point at a service that's gone, or fall out of sync across gateway instances. To on-call, this looks like plain errors with no obvious cause, since the backends are healthy.

open as a page

If you scale out a filter in a message pipeline to run multiple parallel instances for higher throughput, what happens to the order in which messages are processed and re-emitted, and when does that matter?

level: seniorimportance: should knowfreq 45%

basics

~20 s

With several workers grabbing messages at once, a worker that gets an easy message can finish before another worker still stuck on a harder one — so messages can come out in a different order than they went in. That's fine for some jobs, but breaks ones where order matters, like applying account changes in sequence.

open as a page

When contention is severe enough that even reserved capacity can't keep the high-priority queue's oldest messages within their target latency, what design options exist to protect the high-priority tier further, and what do they cost?

level: seniorimportance: should knowfreq 45%

basics

~20 s

If just going first isn't enough anymore, you can also refuse or delay some of the low-priority work entirely — like temporarily rejecting or throttling non-urgent requests — so the important stuff always has enough room. It means some low-priority work gets sacrificed on purpose.

open as a page

In a saga, a compensating transaction for a shipping step (e.g., 'cancel shipment') times out and the orchestrator retries it. What could go wrong if the compensating transaction isn't idempotent, and what design measures make saga steps and their compensations safe to retry?

level: seniorimportance: should knowfreq 70%

basics

~10 s

If a retried step or compensation isn't idempotent, running it twice can double-charge, double-refund, or double-cancel something. Saga steps need unique operation IDs so duplicates can be detected and safely ignored.

open as a page

As a principal engineer setting integration standards across many teams, when would you actively tell a team NOT to build an Anti-Corruption Layer for a given integration, and how do you decide when an existing one should be retired?

level: principalimportance: should knowfreq 40%

basics

~20 s

Skip it when the outside system's model is already close to yours, the integration is small, short-lived, or has one consumer, or you're about to fully replace the outside system anyway. Retire an existing ACL once nothing is left using the system it protects against.

open as a page

You're designing an API for an operation that typically completes in under 300ms but occasionally, for large inputs, takes 10+ seconds. A colleague proposes using the Asynchronous Request-Reply pattern, HTTP 202 plus polling, for every call to keep the API uniform. What are the arguments against reflexively applying this pattern here, and what alternatives would you weigh instead?

level: principalimportance: should knowfreq 40%

basics

~20 s

For work that's usually fast, forcing every client through an accept-then-poll dance adds unnecessary round trips and complexity for the common case. Better options: keep it synchronous with a generous timeout for the fast path, or only switch to async for inputs that are actually large or slow.

open as a page

You're designing the autoscaling policy for a fleet of competing consumers that reads from a queue processing revenue-critical jobs (e.g., order fulfillment). What trade-offs do you weigh in deciding how aggressively to scale the consumer count up and down with queue depth?

level: principalimportance: should knowfreq 45%

basics

~20 s

You have to decide how fast to add or remove workers as the pile of waiting jobs grows or shrinks. Add workers too slowly and jobs pile up; add them too fast and you might overload the systems those workers depend on, or pay for capacity you don't need.

open as a page

A platform centralizes TLS termination, authentication, rate limiting, and caching into one API gateway layer in front of dozens of microservices. What does this concentration cost the system architecturally, and under what circumstances would you deliberately NOT centralize one of these concerns?

level: principalimportance: should knowfreq 40%

basics

~20 s

Putting all these shared jobs in one gateway makes every service simpler, but now that one gateway becomes critical: if it's slow, wrong, or down, everything behind it is affected, and one team owns something every other team depends on.

open as a page

When would you deliberately avoid the Pipes and Filters pattern for a data-processing workload, choosing a monolithic processor or a different architecture instead, and what specifically drives that call?

level: principalimportance: should knowfreq 35%

basics

~20 s

If the job is small, simple, needs to be lightning-fast, or a small team already struggles to run one service — splitting it into many pieces connected by queues adds cost and complexity that isn't worth it. Sometimes one well-written function does the job better.

open as a page

At what point does adding priority tiers to a messaging system stop being worth it, and what alternative approaches would you consider instead of (or alongside) the Priority Queue pattern at large scale?

level: principalimportance: should knowfreq 40%

basics

~20 s

If the different kinds of work are similar enough in urgency, or the system is small, adding separate priority queues can just be extra complexity for no real benefit. Sometimes it's simpler and cheaper to just give everything enough capacity, or to run truly urgent work on its own completely separate system instead.

open as a page

A platform has grown to support three very different client types, a mobile app, a web dashboard, and third-party partner integrations, each needing different composed views over the same dozen backend microservices. What architectural choices exist for where aggregation logic should live, and what are the trade-offs between a per-client Backend-For-Frontend, one shared generic aggregator, and a query-based alternative like GraphQL?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

You can build one aggregator per client type, each owned by that client's team, one shared aggregator for everyone, or replace hand-written aggregation with a query language like GraphQL where the client asks for exactly the fields it needs. Each trades ownership clarity and flexibility against operational overhead.

open as a page

At organizational scale, with many teams and multiple clusters or regions, how does gateway routing need to evolve beyond a single static routing table, and where does it start competing with service-mesh-based routing?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

One hand-edited routing table doesn't scale once many teams own routes and traffic spans regions. Routes need per-team ownership, dynamic publishing as services change, and sometimes a mesh of proxies per service instead of one central gateway.

open as a page

A saga has already completed step 1 (reserve inventory) and step 2 (charge payment) while step 3 (arrange shipping) is still pending. Another transaction reads the order in this intermediate state. What consistency problem does this create, and how does the 'semantic lock' technique help address it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A semantic lock is a flag added to a record (like 'pending') while a saga is still in progress, so other parts of the system know not to fully trust or freely modify that record until the saga finishes.

open as a page

showing 31–52 of 52