skip to content

questions

6

What is an API Gateway in a microservices architecture, and what problem does it solve for client applications that need data from multiple backend services?

level: juniorimportance: must knowfreq 85%

answer

  1. single entry point
  2. reverse proxy at the edge
  3. routing table
  4. cross-cutting concerns centralized
  5. SPOF risk

basics

~20 s

It's a single door that all client requests pass through before reaching the actual services behind it. Instead of a client knowing about and calling ten different services, it talks to one address, and the gateway figures out where to send each request.

solid answer

~30 s

An API Gateway is a reverse-proxy layer sitting between clients and a set of microservices. It gives clients one stable entry point (one host/port, one TLS cert) instead of requiring them to know every service's address. Concretely it does request routing (path/host-based dispatch to the right backend), and centralizes cross-cutting concerns—auth, rate limiting, logging, SSL termination—so each service doesn't reimplement them. This decouples client-facing API shape from internal service topology: services can be split, merged, or renamed without breaking clients, since only the gateway's routing table changes.

go deeper

for a junior

Can state that the gateway is a single entry point that routes requests to the right backend service and give one or two things it centralizes (auth, logging). Doesn't need deployment or failure-mode detail.

for a middle

Explains routing mechanics (path/host matching) and names 2-3 cross-cutting concerns handled at the gateway, and knows it must be deployed redundantly since it's a critical hop.

for a senior

Discusses routing-table drift, service-discovery integration, and how a slow downstream dependency called synchronously by the gateway can cascade into gateway-wide overload; can reason about latency budget added per hop.

for a principal

Weighs organizational cost of a shared gateway (release coordination, ownership bottleneck) against its benefits, and can articulate when NOT to route everything through one central gateway (e.g., internal east-west traffic via a mesh instead).

## What it is and how a request flows through it An **API Gateway** is a server that sits at the network edge of a microservices system, positioned between external clients (web apps, mobile apps, third-party integrators) and the internal fleet of backend services. Mechanically, it is a **reverse proxy**: 1. Every inbound request first hits the gateway, which **inspects the request** (typically the URL path, HTTP method, host header, or a version prefix). 2. It **looks up a routing table** that maps that pattern to a specific backend service, port, and protocol. For example, requests to `/orders/*` might route to the `order-service`, while `/users/*` routes to the `user-service`. 3. The gateway then **forwards the request** to the resolved backend, waits for the response, and relays it back to the client, often rewriting headers or paths along the way (stripping an internal-only prefix, adding a correlation ID for tracing). ## Why it exists **Why it exists:** before gateways became standard, clients called services directly, which meant every client needed to know the network address of every service, handle many different TLS certificates, and reimplement cross-cutting logic (authentication, throttling, request logging, CORS handling) once per service. That duplication is expensive and error-prone — a bug in the auth check on one service doesn't get fixed everywhere at once. A gateway collapses this to a single well-known entry point: - one **DNS name**; - one **TLS termination point**; - one place to enforce a **security policy**. It also decouples the client-visible API surface from internal topology. Internally, an order-processing capability might be one service today and split into `order-service` and `payment-service` tomorrow; as long as the gateway's routing table is updated, no client code changes. This is the same value reverse proxies like **Nginx** have always provided, extended with API-aware features: - **JSON-aware transformations**; - **service discovery** integration; - API-specific policies like **per-route rate limits**. ## Trade-offs, in both directions Trade-offs run in both directions. - **On the upside:** centralizing auth, rate limiting, and TLS termination means those concerns are implemented and audited once instead of many times, and operational visibility improves because all traffic funnels through one observable point (a single place to add access logs, metrics, and distributed tracing headers). - **On the downside**, the gateway becomes a mandatory hop added to every request's **latency budget** — even a well-tuned gateway adds a few milliseconds of proxying overhead, and a poorly tuned one can add tens of milliseconds. - It also becomes a **single point of failure** and a **scaling bottleneck**: if the gateway goes down, every service becomes unreachable from outside even though the services themselves are healthy, so gateways are typically deployed as a horizontally scaled, stateless fleet behind a load balancer, with no single instance holding critical state. - There is also an **organizational risk**: because the gateway is shared infrastructure, changes to its routing table or policies require coordination across teams, and it can become a bottleneck for release velocity if every team must go through a central gateway team to ship a new route — a common criticism when gateways evolve into what people call a 'distributed monolith' chokepoint. ## Failure modes in production Failure modes in production show up in recognizable patterns. 1. **First, routing table drift:** a service is renamed, redeployed on a new port, or decommissioned, but the gateway's route entry isn't updated, causing 502/504 errors that look like a backend outage but are actually a stale gateway config — this is why gateways are increasingly paired with service discovery (Kubernetes Services, Consul) rather than static route tables. 2. **Second, cascading overload:** if the gateway does synchronous work per request (token introspection calls to an auth server, for instance) and that auth server slows down, the gateway's own thread/connection pool can exhaust, taking down routing to every service, not just the one whose dependency is slow — this argues for timeouts, circuit breakers, and caching validated tokens at the gateway layer. 3. **Third,** misconfigured rate limits or CORS rules at the gateway can silently break legitimate traffic in ways that are hard to debug because the failure happens before the request ever reaches application code or logs. ## Where you have seen it A concrete real-world example: **Netflix's Zuul** edge-routing layer sits in front of the streaming service's hundreds of microservices, handling dynamic routing, load shedding, and authentication for all inbound traffic from client devices, so that individual services like the recommendation engine or playback service never have to implement their own edge security or throttling. Similarly, **AWS API Gateway** is commonly placed in front of a fleet of Lambda functions or ECS services precisely so that a mobile app talks to one HTTPS endpoint instead of directly invoking dozens of individually-addressed functions.

  • If the gateway becomes a single point of failure, how would you deploy it to avoid that?
    Run it as a stateless, horizontally scaled fleet of identical gateway instances behind a load balancer or DNS-based failover, with no instance holding session state that would make failover lossy. Health checks pull unhealthy instances out of rotation automatically, and routing config is pushed centrally so every instance stays in sync.
  • How does the gateway know where to route a request when services scale up/down or move?
    Instead of a static routing table, production gateways integrate with service discovery (Kubernetes DNS/Services, Consul, Eureka) so routes resolve to a live, changing set of healthy backend instances rather than hardcoded addresses. The gateway or its sidecar polls or subscribes to discovery updates and adjusts its routing/load-balancing pool accordingly.
  • What's the difference between routing at the gateway and load balancing?
    Routing decides which service a request goes to based on the request's content (path, host, headers); load balancing decides which specific instance of that already-chosen service handles it. A gateway typically does both: route to a service, then load-balance across that service's healthy instances.

Like a hotel concierge desk: guests don't wander the building looking for housekeeping, room service, or maintenance directly — they tell the concierge what they need, and the concierge routes the request to the right department, while also checking ID and handling common requests uniformly.

saying these in an interview costs you the question

  • Says the gateway 'is' a microservice rather than infrastructure/edge layer
  • Can't explain what happens if the gateway itself goes down
  • Thinks routing and load balancing are the same decision
  • Assumes the gateway always adds negligible latency regardless of what work it does per request
  • No mention of service discovery / static routing table not updating

context

open as a page

How does an API Gateway offload authentication and authorization so that individual backend microservices don't each have to implement their own login/token-checking logic, and what does the gateway typically pass downstream once a request is verified?

level: middleimportance: must knowfreq 75%

basics

~20 s

The gateway checks the caller's login token once, before the request goes anywhere else. If it's valid, the gateway lets the request through and tells the backend service who the caller is; if not, it rejects the request immediately.

open as a page

An API Gateway needs to stop any single client from making more than 100 requests per minute. Walk through how a token-bucket rate limiter enforces that, and why teams often layer both per-client and global limits at the gateway.

level: middleimportance: must knowfreq 70%

basics

~20 s

The gateway keeps a running count of how many requests each caller has made recently. Once a caller crosses the limit, the gateway rejects further requests for a while instead of forwarding them, protecting the backend services from being overwhelmed.

open as a page

Describe the request-aggregation pattern at an API Gateway — where the gateway fans a single incoming client request out to multiple backend services and combines their responses into one — and the main risk this introduces compared to simple 1:1 routing.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Instead of the client making five separate calls to five services, it makes one call to the gateway, and the gateway calls all five services itself, combines the answers, and sends back one response. The risk is that the slowest of those five calls determines how long the client waits, and if any one fails, the whole combined answer can fail too.

open as a page

A backend service behind the gateway starts responding slowly (not erroring, just slow) due to database contention. Without any protective mechanism at the gateway, explain how this can escalate into an outage for unrelated services, and what gateway-level mechanism prevents that.

level: seniorimportance: should knowfreq 50%

basics

~20 s

If the gateway just waits and waits for the slow service, it uses up all its capacity holding open those slow requests, leaving nothing free to handle requests to other, healthy services. A circuit breaker makes the gateway stop sending requests to the slow service for a while so it can recover and free up capacity for everything else.

open as a page

A platform team is choosing between Kong, AWS API Gateway, and a plain Nginx reverse proxy to front a set of microservices. What are the core architectural differences that should drive the choice, and under what circumstances would routing everything through one central gateway be the WRONG call?

level: principalimportance: should knowfreq 45%

basics

~20 s

Nginx is a fast general-purpose proxy you configure and run yourself; Kong adds API-specific features (plugins for auth, rate limiting) on top of that same self-hosted model; AWS API Gateway is a fully managed service with those features built in but tied to AWS. Sometimes it's better NOT to force all traffic through one shared gateway, especially service-to-service traffic inside the system.

open as a page