skip to content

Communication & Data Management

How services actually talk and who owns the data: synchronous versus asynchronous styles, gateways and BFFs, contract versioning, and keeping data consistent when nobody shares a database.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

What is an API Gateway in a microservices architecture, and what problem does it solve for client applications that need data from multiple backend services?

level: juniorimportance: must knowfreq 85%

answer

  1. single entry point
  2. reverse proxy at the edge
  3. routing table
  4. cross-cutting concerns centralized
  5. SPOF risk

basics

~20 s

It's a single door that all client requests pass through before reaching the actual services behind it. Instead of a client knowing about and calling ten different services, it talks to one address, and the gateway figures out where to send each request.

solid answer

~30 s

An API Gateway is a reverse-proxy layer sitting between clients and a set of microservices. It gives clients one stable entry point (one host/port, one TLS cert) instead of requiring them to know every service's address. Concretely it does request routing (path/host-based dispatch to the right backend), and centralizes cross-cutting concerns—auth, rate limiting, logging, SSL termination—so each service doesn't reimplement them. This decouples client-facing API shape from internal service topology: services can be split, merged, or renamed without breaking clients, since only the gateway's routing table changes.

go deeper

for a junior

Can state that the gateway is a single entry point that routes requests to the right backend service and give one or two things it centralizes (auth, logging). Doesn't need deployment or failure-mode detail.

for a middle

Explains routing mechanics (path/host matching) and names 2-3 cross-cutting concerns handled at the gateway, and knows it must be deployed redundantly since it's a critical hop.

for a senior

Discusses routing-table drift, service-discovery integration, and how a slow downstream dependency called synchronously by the gateway can cascade into gateway-wide overload; can reason about latency budget added per hop.

for a principal

Weighs organizational cost of a shared gateway (release coordination, ownership bottleneck) against its benefits, and can articulate when NOT to route everything through one central gateway (e.g., internal east-west traffic via a mesh instead).

## What it is and how a request flows through it An **API Gateway** is a server that sits at the network edge of a microservices system, positioned between external clients (web apps, mobile apps, third-party integrators) and the internal fleet of backend services. Mechanically, it is a **reverse proxy**: 1. Every inbound request first hits the gateway, which **inspects the request** (typically the URL path, HTTP method, host header, or a version prefix). 2. It **looks up a routing table** that maps that pattern to a specific backend service, port, and protocol. For example, requests to `/orders/*` might route to the `order-service`, while `/users/*` routes to the `user-service`. 3. The gateway then **forwards the request** to the resolved backend, waits for the response, and relays it back to the client, often rewriting headers or paths along the way (stripping an internal-only prefix, adding a correlation ID for tracing). ## Why it exists **Why it exists:** before gateways became standard, clients called services directly, which meant every client needed to know the network address of every service, handle many different TLS certificates, and reimplement cross-cutting logic (authentication, throttling, request logging, CORS handling) once per service. That duplication is expensive and error-prone — a bug in the auth check on one service doesn't get fixed everywhere at once. A gateway collapses this to a single well-known entry point: - one **DNS name**; - one **TLS termination point**; - one place to enforce a **security policy**. It also decouples the client-visible API surface from internal topology. Internally, an order-processing capability might be one service today and split into `order-service` and `payment-service` tomorrow; as long as the gateway's routing table is updated, no client code changes. This is the same value reverse proxies like **Nginx** have always provided, extended with API-aware features: - **JSON-aware transformations**; - **service discovery** integration; - API-specific policies like **per-route rate limits**. ## Trade-offs, in both directions Trade-offs run in both directions. - **On the upside:** centralizing auth, rate limiting, and TLS termination means those concerns are implemented and audited once instead of many times, and operational visibility improves because all traffic funnels through one observable point (a single place to add access logs, metrics, and distributed tracing headers). - **On the downside**, the gateway becomes a mandatory hop added to every request's **latency budget** — even a well-tuned gateway adds a few milliseconds of proxying overhead, and a poorly tuned one can add tens of milliseconds. - It also becomes a **single point of failure** and a **scaling bottleneck**: if the gateway goes down, every service becomes unreachable from outside even though the services themselves are healthy, so gateways are typically deployed as a horizontally scaled, stateless fleet behind a load balancer, with no single instance holding critical state. - There is also an **organizational risk**: because the gateway is shared infrastructure, changes to its routing table or policies require coordination across teams, and it can become a bottleneck for release velocity if every team must go through a central gateway team to ship a new route — a common criticism when gateways evolve into what people call a 'distributed monolith' chokepoint. ## Failure modes in production Failure modes in production show up in recognizable patterns. 1. **First, routing table drift:** a service is renamed, redeployed on a new port, or decommissioned, but the gateway's route entry isn't updated, causing 502/504 errors that look like a backend outage but are actually a stale gateway config — this is why gateways are increasingly paired with service discovery (Kubernetes Services, Consul) rather than static route tables. 2. **Second, cascading overload:** if the gateway does synchronous work per request (token introspection calls to an auth server, for instance) and that auth server slows down, the gateway's own thread/connection pool can exhaust, taking down routing to every service, not just the one whose dependency is slow — this argues for timeouts, circuit breakers, and caching validated tokens at the gateway layer. 3. **Third,** misconfigured rate limits or CORS rules at the gateway can silently break legitimate traffic in ways that are hard to debug because the failure happens before the request ever reaches application code or logs. ## Where you have seen it A concrete real-world example: **Netflix's Zuul** edge-routing layer sits in front of the streaming service's hundreds of microservices, handling dynamic routing, load shedding, and authentication for all inbound traffic from client devices, so that individual services like the recommendation engine or playback service never have to implement their own edge security or throttling. Similarly, **AWS API Gateway** is commonly placed in front of a fleet of Lambda functions or ECS services precisely so that a mobile app talks to one HTTPS endpoint instead of directly invoking dozens of individually-addressed functions.

  • If the gateway becomes a single point of failure, how would you deploy it to avoid that?
    Run it as a stateless, horizontally scaled fleet of identical gateway instances behind a load balancer or DNS-based failover, with no instance holding session state that would make failover lossy. Health checks pull unhealthy instances out of rotation automatically, and routing config is pushed centrally so every instance stays in sync.
  • How does the gateway know where to route a request when services scale up/down or move?
    Instead of a static routing table, production gateways integrate with service discovery (Kubernetes DNS/Services, Consul, Eureka) so routes resolve to a live, changing set of healthy backend instances rather than hardcoded addresses. The gateway or its sidecar polls or subscribes to discovery updates and adjusts its routing/load-balancing pool accordingly.
  • What's the difference between routing at the gateway and load balancing?
    Routing decides which service a request goes to based on the request's content (path, host, headers); load balancing decides which specific instance of that already-chosen service handles it. A gateway typically does both: route to a service, then load-balance across that service's healthy instances.

Like a hotel concierge desk: guests don't wander the building looking for housekeeping, room service, or maintenance directly — they tell the concierge what they need, and the concierge routes the request to the right department, while also checking ID and handling common requests uniformly.

saying these in an interview costs you the question

  • Says the gateway 'is' a microservice rather than infrastructure/edge layer
  • Can't explain what happens if the gateway itself goes down
  • Thinks routing and load balancing are the same decision
  • Assumes the gateway always adds negligible latency regardless of what work it does per request
  • No mention of service discovery / static routing table not updating

context

open as a page

What problem does the Backend for Frontend (BFF) pattern solve, and how does it typically sit between a client app and the backend services?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A BFF is a small backend built just for one type of app (like a mobile app or a website) so that app gets exactly the data it needs in one call, instead of talking to lots of different services itself and stitching the results together.

open as a page

In a microservices system, what's the basic difference between a synchronous call between services (like REST over HTTP) and an asynchronous message (like publishing to a queue), and when would you pick each?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A synchronous call waits for an immediate answer before moving on, like a phone call. An asynchronous message is sent off while the sender keeps working, like leaving a voicemail; the other side handles it and replies whenever it's ready.

open as a page

In a microservices system, what does it mean for an API change to be 'backward compatible', and what is a concrete example of a change that breaks it?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A backward-compatible change lets old clients keep working without updates, like adding a new optional field. A breaking change forces every client to update at once, like renaming or removing a field they depend on.

open as a page

In a microservices architecture, why does each service typically own and manage its own database rather than multiple services sharing one database?

level: juniorimportance: must knowfreq 85%

basics

~10 s

Each service gets its own database so only that service can change its data directly; others must ask through its API, which stops hidden dependencies and lets services evolve independently.

open as a page

In a microservices order flow where Orders, Payments, and Inventory each own separate databases, why do teams typically reach for a saga instead of a two-phase commit (2PC) to keep the data consistent across them?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A saga splits one big transaction into small local ones, each service commits its own piece right away; if something later fails, earlier steps are undone with a follow-up 'undo' action instead of everyone waiting and locking together like 2PC.

open as a page

How does an API Gateway offload authentication and authorization so that individual backend microservices don't each have to implement their own login/token-checking logic, and what does the gateway typically pass downstream once a request is verified?

level: middleimportance: must knowfreq 75%

basics

~20 s

The gateway checks the caller's login token once, before the request goes anywhere else. If it's valid, the gateway lets the request through and tells the backend service who the caller is; if not, it rejects the request immediately.

open as a page

An API Gateway needs to stop any single client from making more than 100 requests per minute. Walk through how a token-bucket rate limiter enforces that, and why teams often layer both per-client and global limits at the gateway.

level: middleimportance: must knowfreq 70%

basics

~20 s

The gateway keeps a running count of how many requests each caller has made recently. Once a caller crosses the limit, the gateway rejects further requests for a while instead of forwarding them, protecting the backend services from being overwhelmed.

open as a page

Concretely, what does a BFF do differently when serving a mobile app's product page versus a desktop web app's product page, and how does it accomplish that?

level: middleimportance: must knowfreq 65%

basics

~20 s

The BFF talks to the same backend services either way, but it sends the mobile app a smaller, simpler bundle of data (fewer fields, smaller images) and sends the web app a fuller bundle, because phones have less screen space and often slower or costlier networks.

open as a page

What's the difference between the request/reply and publish/subscribe messaging patterns for service-to-service communication, and what coupling trade-off does each impose?

level: middleimportance: must knowfreq 80%

basics

~20 s

Request/reply is one service asking another a direct question and getting a direct answer, like calling a specific person. Publish/subscribe is one service announcing something happened, and any number of other services that care can listen in, like posting an announcement on a notice board.

open as a page

How does consumer-driven contract testing (e.g., with a tool like Pact) work, and what problem does it solve that end-to-end integration tests don't?

level: middleimportance: must knowfreq 70%

basics

~10 s

Each consumer writes down exactly what it expects from a provider's API. The provider runs those expectations as tests before deploying, catching breakage early without needing a slow, flaky full end-to-end environment.

open as a page

When Service A needs data that Service B owns (for example, a product's name and price displayed inside an order), what are the two main strategies for getting that data, and what does each trade off?

level: middleimportance: must knowfreq 75%

basics

~20 s

Either Service A asks Service B live every time it needs the data (simple but creates a runtime dependency), or Service A keeps its own local copy that's kept up to date via events (fast and independent, but the copy can be briefly out of date).

open as a page

What is a 'read model' built by replicating data from other services via events, and why might a Search service maintain its own copy of product data instead of querying the Catalog service directly for every search request?

level: middleimportance: must knowfreq 70%

basics

~20 s

A read model is a service's own local, search or query-optimized copy of data that another service owns, kept updated by listening to that service's change events, so it can answer queries fast and independently instead of calling the owner every time.

open as a page

When implementing a saga across Order, Payment, and Shipping services, what's the difference between a choreography-based saga and an orchestration-based saga, and what would push you to pick one over the other?

level: middleimportance: must knowfreq 85%

basics

~10 s

Choreography: each service listens for events and reacts on its own, no one's in charge. Orchestration: one central coordinator tells each service what to do next, step by step.

open as a page

A service updates its own database and then needs to publish an event so the next saga step can run. What is the 'dual write' problem this creates, and how does the transactional outbox pattern solve it?

level: middleimportance: must knowfreq 75%

basics

~20 s

If a service saves to its database and separately sends a message, one can succeed while the other fails, leaving things out of sync. The outbox pattern saves the message in the same database transaction as the data change, then a separate process reliably delivers it, so both always happen together or neither does.

open as a page

Describe the request-aggregation pattern at an API Gateway — where the gateway fans a single incoming client request out to multiple backend services and combines their responses into one — and the main risk this introduces compared to simple 1:1 routing.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Instead of the client making five separate calls to five services, it makes one call to the gateway, and the gateway calls all five services itself, combines the answers, and sends back one response. The risk is that the slowest of those five calls determines how long the client waits, and if any one fails, the whole combined answer can fail too.

open as a page

How does a Backend-for-Frontend pattern differ from a single shared API gateway serving all client types, and what's a concrete case where the shared-gateway approach breaks down?

level: seniorimportance: must knowfreq 60%

basics

~20 s

A shared gateway is one door for every app; a BFF is a separate, tailored door per app. The shared-gateway approach breaks down when different apps need very different data shapes and change at different speeds, because now every client's feature work has to queue behind the one team that owns the shared door.

open as a page

In an order-fulfillment flow spanning an Order, Payment, Inventory, and Shipping service, what's the difference between choreography (each service reacts to events from the others) and orchestration (a central coordinator directs each step), and what are the concrete trade-offs?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Choreography is like a dance where everyone reacts to what everyone else does, with no one in charge. Orchestration is like a conductor telling each musician exactly when to play. Choreography has no single controller; orchestration has one service driving the whole process.

open as a page

When a public-facing service needs to expose multiple versions of its API at once (e.g., /v1/orders vs /v2/orders, or via an Accept header), what are the main versioning strategies, and what's the operational cost of each?

level: seniorimportance: must knowfreq 75%

basics

~20 s

You can put the version in the URL path, in a request header, or in a query parameter. URL versioning is the simplest and most visible; header-based versioning keeps URLs stable but is less discoverable. All of them mean you're running and maintaining more than one version of your API at once.

open as a page

A team is choosing between REST with JSON over HTTP/1.1 and gRPC over HTTP/2 for synchronous calls between two internal services. What concrete technical trade-offs should drive that decision?

level: middleimportance: should knowfreq 65%

basics

~20 s

REST with JSON is simple, human-readable, and works everywhere, including browsers, but is slower to parse and less strict about structure. gRPC uses a compact binary format and a strict contract, making it faster and safer between services, but harder to read by hand and not natively usable from a browser.

open as a page

What is the tolerant reader pattern in service-to-service integration, and what specifically should a consumer's parsing code do (and not do) to implement it?

level: middleimportance: should knowfreq 55%

basics

~20 s

A tolerant reader only looks at the specific fields it actually needs from a response and ignores everything else, so it doesn't break when the provider adds new fields or extends the data it sends.

open as a page

A backend service behind the gateway starts responding slowly (not erroring, just slow) due to database contention. Without any protective mechanism at the gateway, explain how this can escalate into an outage for unrelated services, and what gateway-level mechanism prevents that.

level: seniorimportance: should knowfreq 50%

basics

~20 s

If the gateway just waits and waits for the slow service, it uses up all its capacity holding open those slow requests, leaving nothing free to handle requests to other, healthy services. A circuit breaker makes the gateway stop sending requests to the slow service for a while so it can recover and free up capacity for everything else.

open as a page

A company gives each frontend team - web, iOS, Android - its own BFF that they build and deploy themselves. What organizational benefit does this bring, and what failure mode commonly emerges a year later if it isn't actively managed?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Frontend teams can ship changes fast without waiting on a separate backend team, since they own their own small backend. But over time, each team's BFF often quietly reimplements the same business rules, and those copies drift apart until different apps disagree with each other.

open as a page

Service A calls Service B synchronously over REST, and B in turn calls Service C synchronously. C starts responding slowly under load. What failure mode does this create for A, and what are the standard mitigations?

level: seniorimportance: should knowfreq 70%

basics

~20 s

C being slow makes B slow, which makes A slow too, and if nothing stops it, all three can eventually run out of capacity and fail together, like one traffic jam backing up onto every road behind it. Fixes include setting time limits, giving up early instead of piling up more requests, and having a backup plan when the answer doesn't come.

open as a page

When a schema registry enforces 'BACKWARD' compatibility mode on a service's request/response schema, what exactly does it check before allowing a new schema version to be registered, and how does that differ from 'FORWARD' and 'FULL' compatibility modes?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A schema registry checks new versions of a data format before allowing them to be published, using rules like 'can old readers still read new data' (forward) or 'can new readers read old data' (backward), and 'full' means both must hold at once.

open as a page

When a service updates its own database and then needs to publish an event about that change so other services can update their read models, what can go wrong if it does these as two separate steps, and what's a common pattern to fix it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

If a service saves to its database and then separately sends a message, one of those two steps can fail after the other succeeds, for example the app crashes right between them, leaving the database and the message out of sync. A common fix is to write the event into the same database transaction as the data change, then have a separate process reliably forward it.

open as a page

When designing the compensating step for a saga stage — say, 'release reserved inventory' to undo an earlier 'reserve inventory' step — what makes some compensations straightforward and others genuinely hard or impossible to write correctly?

level: seniorimportance: should knowfreq 65%

basics

~20 s

A compensation is a new action that semantically cancels an earlier step, not a database undo. It's easy when the earlier step is fully reversible (like un-reserving stock); it's hard or impossible when the earlier step already had a real-world, external, or irreversible effect, like an email that was already sent or money that already left the system.

open as a page

A saga step's message handler receives the same 'PaymentCharged' event twice because the broker redelivered it after a timeout. What is the inbox pattern / idempotent consumer approach, and how does it prevent the handler from applying the event's effect twice?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The consumer keeps a record of which message IDs it already processed, in the same database transaction as the work it does. If the same message shows up again, it checks that record first and skips redoing the work, so a duplicate delivery has no extra effect.

open as a page

A platform team is choosing between Kong, AWS API Gateway, and a plain Nginx reverse proxy to front a set of microservices. What are the core architectural differences that should drive the choice, and under what circumstances would routing everything through one central gateway be the WRONG call?

level: principalimportance: should knowfreq 45%

basics

~20 s

Nginx is a fast general-purpose proxy you configure and run yourself; Kong adds API-specific features (plugins for auth, rate limiting) on top of that same self-hosted model; AWS API Gateway is a fully managed service with those features built in but tied to AWS. Sometimes it's better NOT to force all traffic through one shared gateway, especially service-to-service traffic inside the system.

open as a page

A startup has one web client, a small team, and a single small backend service. An architect proposes introducing a BFF layer in front of it 'for future scalability.' Why might this be premature, and what should trigger actually adopting the pattern later?

level: principalimportance: should knowfreq 40%

basics

~20 s

With only one app and one small backend, there's nothing to tailor a response for and no second team that needs independence - a BFF here just adds an extra network hop and another service to run, for no real benefit yet. Add it later, once there are genuinely different clients or teams pulling the API in different directions.

open as a page

showing 1–30 of 34