Messaging & Integration
Patterns for asynchronous integration and for the gateway in front of your services: competing consumers, priority queue, pipes and filters, gateway aggregation, offloading and routing, anti-corruption layer and saga.
part ofResilience & cloud-native patternsoverview, primer and where to startread it →on this pageshowhide
explore
- Competing Consumers6 questions
- Priority Queue6 questions
- Pipes and Filters6 questions
- Gateway Aggregation6 questions
- Gateway Offloading5 questions
- Gateway Routing6 questions
- Anti-Corruption Layer (bridge)6 questions
- Asynchronous Request-Reply5 questions
- Saga Pattern6 questions
- AI Engineerrole
- API Designskill
- Backend Developerrole
- Data Engineerrole
- DevOps / SRE Engineerrole
- Forward Deployed Engineerrole
- Full Stack Developerrole
- Game Developerrole
- Java Backend Developerrole
- Kotlin Backend Developerrole
- Server-Side Game Developerrole
- Software Architectrole
- Software Design & Architectureskill
- System Designskill
questions
page 2 of 2For filters connected by pipes in a message-processing pipeline to be truly reusable and composable — droppable into different pipelines without rewriting them — what do they need to agree on, and what breaks that reusability in practice?
basics
~20 sAll the workers need to speak the same language for the data passing between them — same fields, same format. If one worker expects data shaped differently than the last worker sends it, they can't be swapped or reused.
In a Priority Queue setup with separate high- and low-priority queues, strict 'always serve high-priority first' consumer logic can starve the low-priority queue during sustained high-priority load. Name two concrete techniques to prevent this, and explain how each works.
basics
~20 sInstead of always doing important work first no matter what, you can either give a small guaranteed share of time to the less-important queue (like 1 in 10 turns), or make old low-priority messages count as more important the longer they wait, so they eventually get processed.
In a cloud migration where a new set of microservices is being built to gradually replace an on-premises monolith, how does an Anti-Corruption Layer typically get deployed and operated as a cloud integration pattern, and how does that differ from its original description as a Domain-Driven Design concept?
basics
~20 sIn cloud migrations, the ACL is often its own small deployed service (or gateway) that talks to the old on-prem system and the new cloud services, translating between them while the old system is gradually replaced. Originally the same idea was described as code inside a team's own codebase for keeping domain models from mixing.
In a production system, an Anti-Corruption Layer sits in front of a flaky third-party API and has been running fine for months. What are the most likely ways this kind of layer fails or degrades in production, and how would you design against each one?
basics
~20 sThe translator can quietly mis-map data when the outside system changes without telling you, the extra layer can slow things down or become a single point of failure, and it can turn into a dumping ground for random logic if nobody owns it. Guard against these with tests, monitoring, and clear ownership.
An async request-reply API can notify completion either by having the client poll a status endpoint, or by having the client register a webhook URL that the server calls back with the result. What are the operational trade-offs between these two delivery mechanisms, and when would you choose one over the other?
basics
~20 sPolling means the client keeps checking back for updates; a callback (webhook) means the server calls the client back when the work is done. Polling is simpler to set up on the client side; callbacks are faster and less wasteful but need the client to run something reachable.
Why does scaling out a queue with more competing consumer instances typically break global message ordering, and how would you preserve ordering for related messages (e.g., all events for one customer) while still processing the queue in parallel?
basics
~20 sWhen many workers pull from one line, a later job can finish before an earlier one if it lands on a faster worker. To keep related jobs in order, you route them all to the same one worker instead of letting any worker grab them.
In a Competing Consumers setup, a single message keeps causing a processing exception no matter which consumer handles it, and it keeps getting redelivered — starving the queue for other work. How do you design around this 'poison message' problem?
basics
~20 sOne bad message that always fails can get retried forever and clog things up. The fix is to count how many times it's failed, and after a limit, move it out of the main queue into a separate 'failed' queue so a human can look at it later, instead of retrying forever.
An aggregation gateway fans out to five downstream services using a single shared thread pool. One of those services starts responding very slowly, though not erroring outright. What can happen to the gateway as a whole, and how would you isolate the damage?
basics
~20 sIf all five calls share the same limited pool of workers, the slow service can hog enough workers that the gateway can't process calls to the other four healthy services either, so one bad dependency drags down everything. Fix it by giving each downstream its own separate pool or limit and a circuit breaker, so one slow service can only exhaust its own slice.
A gateway is configured to cache GET responses and gzip-compress payloads on behalf of every backend service behind it. What has to be true of a response for gateway-level caching to be safe, and what commonly goes wrong when a team turns it on without checking?
basics
~20 sThe gateway can save a copy of a response and reuse it for later identical requests, saving backend work, but only if the response is truly the same for everyone who asks; if the response depends on who's asking (like personal data), caching it for all users is a serious bug.
In a production system, what are the specific ways a gateway's routing configuration goes wrong, and how does each one actually show up to an on-call engineer?
basics
~20 sRouting rules can overlap and steal each other's traffic, point at a service that's gone, or fall out of sync across gateway instances. To on-call, this looks like plain errors with no obvious cause, since the backends are healthy.
If you scale out a filter in a message pipeline to run multiple parallel instances for higher throughput, what happens to the order in which messages are processed and re-emitted, and when does that matter?
basics
~20 sWith several workers grabbing messages at once, a worker that gets an easy message can finish before another worker still stuck on a harder one — so messages can come out in a different order than they went in. That's fine for some jobs, but breaks ones where order matters, like applying account changes in sequence.
When contention is severe enough that even reserved capacity can't keep the high-priority queue's oldest messages within their target latency, what design options exist to protect the high-priority tier further, and what do they cost?
basics
~20 sIf just going first isn't enough anymore, you can also refuse or delay some of the low-priority work entirely — like temporarily rejecting or throttling non-urgent requests — so the important stuff always has enough room. It means some low-priority work gets sacrificed on purpose.
In a saga, a compensating transaction for a shipping step (e.g., 'cancel shipment') times out and the orchestrator retries it. What could go wrong if the compensating transaction isn't idempotent, and what design measures make saga steps and their compensations safe to retry?
basics
~10 sIf a retried step or compensation isn't idempotent, running it twice can double-charge, double-refund, or double-cancel something. Saga steps need unique operation IDs so duplicates can be detected and safely ignored.
As a principal engineer setting integration standards across many teams, when would you actively tell a team NOT to build an Anti-Corruption Layer for a given integration, and how do you decide when an existing one should be retired?
basics
~20 sSkip it when the outside system's model is already close to yours, the integration is small, short-lived, or has one consumer, or you're about to fully replace the outside system anyway. Retire an existing ACL once nothing is left using the system it protects against.
You're designing an API for an operation that typically completes in under 300ms but occasionally, for large inputs, takes 10+ seconds. A colleague proposes using the Asynchronous Request-Reply pattern, HTTP 202 plus polling, for every call to keep the API uniform. What are the arguments against reflexively applying this pattern here, and what alternatives would you weigh instead?
basics
~20 sFor work that's usually fast, forcing every client through an accept-then-poll dance adds unnecessary round trips and complexity for the common case. Better options: keep it synchronous with a generous timeout for the fast path, or only switch to async for inputs that are actually large or slow.
You're designing the autoscaling policy for a fleet of competing consumers that reads from a queue processing revenue-critical jobs (e.g., order fulfillment). What trade-offs do you weigh in deciding how aggressively to scale the consumer count up and down with queue depth?
basics
~20 sYou have to decide how fast to add or remove workers as the pile of waiting jobs grows or shrinks. Add workers too slowly and jobs pile up; add them too fast and you might overload the systems those workers depend on, or pay for capacity you don't need.
A platform centralizes TLS termination, authentication, rate limiting, and caching into one API gateway layer in front of dozens of microservices. What does this concentration cost the system architecturally, and under what circumstances would you deliberately NOT centralize one of these concerns?
basics
~20 sPutting all these shared jobs in one gateway makes every service simpler, but now that one gateway becomes critical: if it's slow, wrong, or down, everything behind it is affected, and one team owns something every other team depends on.
When would you deliberately avoid the Pipes and Filters pattern for a data-processing workload, choosing a monolithic processor or a different architecture instead, and what specifically drives that call?
basics
~20 sIf the job is small, simple, needs to be lightning-fast, or a small team already struggles to run one service — splitting it into many pieces connected by queues adds cost and complexity that isn't worth it. Sometimes one well-written function does the job better.
At what point does adding priority tiers to a messaging system stop being worth it, and what alternative approaches would you consider instead of (or alongside) the Priority Queue pattern at large scale?
basics
~20 sIf the different kinds of work are similar enough in urgency, or the system is small, adding separate priority queues can just be extra complexity for no real benefit. Sometimes it's simpler and cheaper to just give everything enough capacity, or to run truly urgent work on its own completely separate system instead.
A platform has grown to support three very different client types, a mobile app, a web dashboard, and third-party partner integrations, each needing different composed views over the same dozen backend microservices. What architectural choices exist for where aggregation logic should live, and what are the trade-offs between a per-client Backend-For-Frontend, one shared generic aggregator, and a query-based alternative like GraphQL?
basics
~20 sYou can build one aggregator per client type, each owned by that client's team, one shared aggregator for everyone, or replace hand-written aggregation with a query language like GraphQL where the client asks for exactly the fields it needs. Each trades ownership clarity and flexibility against operational overhead.
At organizational scale, with many teams and multiple clusters or regions, how does gateway routing need to evolve beyond a single static routing table, and where does it start competing with service-mesh-based routing?
basics
~20 sOne hand-edited routing table doesn't scale once many teams own routes and traffic spans regions. Routes need per-team ownership, dynamic publishing as services change, and sometimes a mesh of proxies per service instead of one central gateway.
A saga has already completed step 1 (reserve inventory) and step 2 (charge payment) while step 3 (arrange shipping) is still pending. Another transaction reads the order in this intermediate state. What consistency problem does this create, and how does the 'semantic lock' technique help address it?
basics
~20 sA semantic lock is a flag added to a record (like 'pending') while a saga is still in progress, so other parts of the system know not to fully trust or freely modify that record until the saga finishes.
showing 31–52 of 52