skip to content

Integration Design

Defining how the pieces communicate: synchronous versus asynchronous, API contracts, data formats, event-driven versus request-reply, and anti-corruption layers at the edges. Integration is usually where a solution design succeeds or falls apart.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

When integrating two systems, what is the core difference between a synchronous request-reply call and an asynchronous message-based call, and what does each cost you?

level: juniorimportance: must knowfreq 80%

answer

  1. phone call vs email
  2. temporal coupling
  3. at-least-once delivery
  4. thread pool exhaustion
  5. load leveling

basics

~20 s

Sync: caller waits for an instant answer, like a phone call. Async: caller sends a message and moves on, checking back later, like a letter. Sync is simple but ties your uptime to the other side; async is more resilient but harder to reason about.

solid answer

~40 s

Synchronous integration (REST/HTTP, gRPC unary) blocks the caller until the callee responds: simple mental model, immediate error visibility, but the caller's availability and latency become bounded by the callee's — a slow or down dependency can exhaust the caller's own thread/connection pool, so you need timeouts, circuit breakers, and retries to survive it. Asynchronous integration (a queue, an event, a callback) decouples the two in time: the caller doesn't block, the receiver processes when ready, giving load leveling and resilience to downstream outages. The cost is complexity — durable storage for in-flight messages, handling out-of-order or duplicate delivery, and the caller no longer knows synchronously whether the operation succeeded. Choose sync for latency-sensitive, user-blocking reads; async for write-heavy or cross-boundary work nobody is staring at a screen waiting for.

go deeper

for a junior

Can describe the basic difference (waiting vs not waiting) and give one example of each, without needing to discuss delivery guarantees or cascading failure.

for a middle

Knows to add timeouts/retries to sync calls and can name a queue or event as the async mechanism; starts to reason about which style fits a given feature.

for a senior

Designs the choice per-interaction (not per-system), discusses circuit breakers/bulkheads, idempotency for async consumers, and can trace how a synchronous chain cascades failure.

for a principal

Sets integration-style policy across a platform: defines when teams must use async for resilience/load reasons, owns org-wide retry/backoff and dead-letter conventions, and weighs the operational cost of async debugging against its resilience gains.

## What a synchronous call does **Synchronous integration** means the calling system opens a connection, sends a request, and blocks — holding a thread, a connection, and often a user-facing spinner — until the receiving system computes and returns a response over that same channel. The classic implementations are HTTP/REST calls, gRPC unary calls, or a direct network read. The caller's code proceeds linearly: 1. call 2. wait 3. get response 4. continue This is the model most engineers learn first because it maps directly onto function calls: you invoke something and get a return value back in the same execution flow. ## What an asynchronous call does instead **Asynchronous integration** breaks that direct temporal coupling. - Instead of waiting on the same channel, the caller fires a message into a durable intermediary (a queue, a topic, an outbox table) and moves on immediately, or accepts a callback/webhook the other side invokes later. - The receiving system consumes the message whenever it has capacity, processes it, and — if a reply is needed — sends it back through its own asynchronous channel rather than the original connection. The key shift: request and response become two independent messages instead of one blocking round trip. ## Why both exist Both patterns exist because they solve different problems. - **Synchronous calls** exist because some interactions are inherently read-now: a user clicking 'view profile' needs the data before the page can render — there's no meaningful way to decouple that. - **Asynchronous integration** exists because many interactions are write-now, process-eventually: placing an order doesn't require warehouse, billing, and shipping to all confirm instantly — it requires the order durably recorded, with downstream work happening independently and reliably even if one of those systems is temporarily down. ## The trade-offs The trade-offs cut both ways. | Style | What it buys | What it costs | |---|---|---| | **Synchronous integration** | Simple to trace and debug — a stack trace or distributed-trace span shows the exact causal chain, and errors surface immediately where you can retry or show the user a message. | Its cost is temporal coupling: the caller's availability is bounded by the callee's, latency compounds across chained calls, and a slow dependency can exhaust the caller's thread pool, taking it down too. Mitigations exist (timeouts, circuit breakers, bulkheads, retries with backoff) but they're bolted on, not free — each is another failure mode to test. | | **Asynchronous integration** | Buys resilience to outages (the message sits durably until the consumer is back) and load leveling (a burst becomes a queue drained at the consumer's own pace). | Its cost is that 'I sent it' is no longer 'it succeeded' — you need a separate mechanism to close that loop if the caller cares about the outcome. You also inherit distributed-systems problems: messages can arrive out of order or be delivered more than once (most brokers guarantee at-least-once, not exactly-once), so idempotent consumers become mandatory, not optional hardening. | ## Failure modes Failure modes differ sharply in production. - **A synchronous integration failing** looks like elevated latency, timeouts, and cascading errors — visible fast, but capable of taking multiple systems down together in a retry storm if every caller retries aggressively at once. - **An asynchronous integration failing** looks quieter: messages silently piling up in a queue, a consumer falling behind without anyone noticing until the backlog is hours old, or a poison message a consumer can never process, blocking everything behind it until someone builds a dead-letter queue and monitors depth/age. ## A concrete example A concrete example: an e-commerce checkout typically calls inventory-check and payment-authorization **synchronously** — the customer is staring at the screen and the flow cannot proceed without those answers, so a short timeout with a clear failure message is the right design. But order fulfillment, the confirmation email, analytics, and warehouse notification are dispatched **asynchronously** via an `OrderPlaced` event — none needs to happen before the customer sees 'Order confirmed,' and doing them synchronously would mean checkout hangs on the warehouse system's uptime. Choosing correctly per interaction, rather than picking one style for the whole system, is the actual skill being tested.

  • How do timeouts and circuit breakers help a synchronous integration survive a slow downstream dependency?
    A timeout caps how long the caller waits before giving up, preventing one slow call from exhausting the caller's thread/connection pool. A circuit breaker tracks failure rates and, once a threshold is crossed, stops calling the failing dependency entirely for a cooldown period, failing fast instead of piling up more blocked calls. Together they trade a guaranteed-instant answer for bounded, predictable degradation instead of cascading collapse.
  • If an asynchronous consumer can receive the same message twice, how do you keep processing correct?
    Make the consumer's handler idempotent — processing the same message twice produces the same end state as processing it once, typically by tracking a unique message key and short-circuiting on a duplicate. This is necessary because most brokers offer at-least-once delivery, not exactly-once, so duplicates are a normal occurrence, not an edge case.
  • What is a retry storm and why is it more dangerous in synchronous integrations?
    A retry storm is when many callers, hitting failures from an overloaded service, all retry near-simultaneously, multiplying load on the already-struggling service and prolonging the outage. It's sharper in synchronous integrations because callers are actively blocked and impatient, so naive fixed-interval retries synchronize; jittered exponential backoff spreads retries out to avoid this.

Sync is a phone call - both people are tied up until it ends. Async is dropping a letter in the mailbox - you walk away, and the reply arrives on its own schedule.

saying these in an interview costs you the question

  • Claims async integration is 'just sync but non-blocking' with no mention of durability or delivery guarantees
  • Assumes a message broker guarantees exactly-once delivery by default
  • Says synchronous calls never need timeouts because 'the network is reliable'
  • Treats 'fire and forget' as free of failure modes
  • Can't explain why chained synchronous calls compound latency

context

open as a page

When two teams integrate over an API, what makes a good API contract, and how do you evolve it without breaking existing consumers?

level: middleimportance: must knowfreq 75%

basics

~20 s

A contract is the agreed shape of requests and responses between two systems - like a form both sides fill out the same way. To change it safely, add new optional fields instead of removing or renaming old ones, so old callers keep working.

open as a page

A team is integrating a pricing service with three downstream consumers: a checkout UI that needs the current price before rendering, an analytics warehouse that aggregates prices nightly, and a recommendation engine that reacts to price drops. How should the choice between request-reply and event-driven integration differ across these three consumers, and what do you give up by picking event-driven for all three anyway?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Pick request-reply when a consumer needs the answer right now to keep working, like a checkout screen. Pick events when a consumer just needs to know 'something changed' and can react later, like analytics or recommendations. Using events everywhere adds delay and complexity where you didn't need it.

open as a page

When choosing a data exchange format for an integration - say, JSON versus a binary format like Protobuf or Avro - what are you actually trading off, and how does each handle schema evolution over time?

level: middleimportance: should knowfreq 55%

basics

~20 s

JSON is text you can read with your eyes and is easy to debug, but it's bigger and slower to parse. Formats like Protobuf or Avro are compact and fast but need a shared schema file and special tools to read.

open as a page

Why does an integration contract for a 'create order' operation typically need to define an idempotency key, and what breaks in production if it doesn't?

level: seniorimportance: should knowfreq 60%

basics

~20 s

An idempotency key lets a caller safely retry a request without accidentally doing it twice - like writing a unique order number on a form so resubmitting it doesn't create a second order. Without it, network retries can cause duplicate orders, charges, or emails.

open as a page

When integrating a modern service with a legacy system (or an external system your team doesn't control) whose domain model is a poor fit for yours, what does an anti-corruption layer do, and what does it cost to maintain?

level: principalimportance: should knowfreq 45%

basics

~20 s

An anti-corruption layer is a translation wall between your system and someone else's messy or outdated one - it converts their concepts into yours at the boundary, so their quirks don't leak into your code. It costs extra code to build and keep updated.

open as a page