skip to content

A platform centralizes TLS termination, authentication, rate limiting, and caching into one API gateway layer in front of dozens of microservices. What does this concentration cost the system architecturally, and under what circumstances would you deliberately NOT centralize one of these concerns?

level: principalimportance: should knowfreq 40%

answer

  1. duplication traded for coupling
  2. gateway outage = total outage not partial
  3. shared ownership becomes a bottleneck
  4. extra hop latency on every request
  5. service mesh as decentralized alternative

basics

~20 s

Putting all these shared jobs in one gateway makes every service simpler, but now that one gateway becomes critical: if it's slow, wrong, or down, everything behind it is affected, and one team owns something every other team depends on.

solid answer

~60 s

Gateway offloading trades duplicated-but-independent concerns for centralized-but-coupled ones. The wins are consistency (one implementation of TLS/auth/rate limiting instead of N slightly different ones) and simpler backend services. The costs are: the gateway becomes a single point of failure whose outage or bug affects every service simultaneously, not just one; it becomes a shared-ownership bottleneck where every team's rollout can be blocked by gateway team capacity or a risk-averse change process; it adds a latency hop to every single request; and it can become a dumping ground for logic that doesn't cleanly belong there, blurring into gateway-aggregation or business logic if boundaries aren't disciplined. You'd deliberately avoid centralizing a concern when it genuinely needs per-service nuance the gateway can't express cheaply (e.g., wildly different rate-limit shapes per service that make a generic policy awkward), when the latency cost is unacceptable for a specific hot path, or in a topology where a service mesh's sidecar-per-service model is a better fit than a single chokepoint gateway, giving per-service policy without per-service reimplementation.

go deeper

for a junior

Should sense that putting everything in one place means that one place matters a lot if it breaks, even without detailed architectural vocabulary.

for a middle

Should articulate the single-point-of-failure concept clearly and name at least one concrete cost (latency, ownership bottleneck, or blast radius).

for a senior

Should discuss the duplication-vs-coupling trade-off explicitly and describe mitigations like self-service configuration or higher availability engineering for the gateway itself.

for a principal

Should reason across technical and organizational dimensions simultaneously, know when to reach for a service mesh instead of a monolithic gateway, and articulate concrete circumstances for deliberately not centralizing a given concern.

## Duplication traded for coupling Gateway offloading is fundamentally a trade of **duplication for coupling**. | Before centralization | After centralization | |---|---| | every backend service independently implements TLS, authentication, rate limiting, and caching, which means N slightly different implementations, N sets of bugs, N teams each needing security and HTTP expertise, but also N independent blast radii, a bug or outage in one service's auth code affects only that service | there's exactly one implementation of each concern, consistent everywhere, maintained by presumably one team of specialists, but now every service's behavior for that concern is only as correct, as available, and as fast as the shared gateway is | ## The first cost — a genuine single point of failure The most important cost is that the gateway becomes a genuine **single point of failure** for the entire system's public entry point. A bug in the gateway's rate-limiting logic, a bad certificate rollout, a memory leak under load, or a misconfigured deploy doesn't take down one service, it takes down every service simultaneously, because every request to every backend flows through this one component first. This is qualitatively different from a single backend service failing, which degrades one feature; a gateway failure is a full outage. Mitigating this requires the gateway itself to be built to a much higher availability bar than any individual backend service: - redundant instances across availability zones - careful canary/rolling deploys - circuit breakers and fallback behavior for its own dependencies (like the Redis store backing rate limiting) - aggressive monitoring, because its failure mode is total, not partial ## The second cost — ownership becomes a bottleneck The second cost is organizational, not just technical: centralizing these concerns usually means centralizing ownership too. One team, often called the platform or gateway team, now owns a component every other team's traffic depends on. This creates a **coordination bottleneck**: if a service team needs a new auth scope supported, a bespoke rate-limit shape, or a caching exception, they now depend on the gateway team's roadmap and risk tolerance, rather than just shipping a change in their own service. Mature organizations handle this by making the gateway highly configurable via self-service policy (declarative rate-limit configs per route, pluggable auth scopes) rather than requiring gateway-team code changes for every service's needs, but that configurability itself has to be built and maintained, which is real ongoing cost, not a one-time setup. ## The third cost — latency and design discipline The third cost is latency and request-path complexity: every request now takes an extra hop through the gateway before reaching its actual destination, and every additional cross-cutting check (TLS handshake already paid once at connection level, but auth lookup, rate-limit counter check, cache lookup) adds processing time on that hop. For most services this is negligible, single-digit milliseconds, but for genuinely latency-critical hot paths (real-time trading, some gaming backends), that overhead can matter enough to justify a different, more distributed approach. There's also a **design-discipline risk**: because the gateway is a convenient, central place to add logic, teams are tempted to push things there that don't belong, like fine-grained business authorization or multi-service response composition, drifting the gateway from cross-cutting concerns (offloading) toward domain logic or request orchestration (which is properly a separate concern, gateway aggregation, and arguably shouldn't live in the same component either, since it has very different scaling and failure characteristics). ## When to leave a concern where it is Deliberately not centralizing a concern makes sense in a few concrete circumstances. 1. **First**, when a concern genuinely needs deep per-service nuance that a shared gateway config can't express cheaply, for example wildly different rate-limiting shapes (one service needs a token-bucket keyed by a composite of user+resource, another needs a simple per-IP window), pushing that into a generic gateway policy layer can become more awkward than just implementing it locally, or via a library shared across services rather than a runtime gateway. 2. **Second**, when the latency cost of an extra hop is unacceptable for a specific hot path, some architectures route that path around the general-purpose gateway through a lighter, purpose-built entry point. 3. **Third**, and increasingly common at scale, a service mesh (Istio, Linkerd) offers a middle ground: instead of one central chokepoint gateway, each service gets its own sidecar proxy handling TLS, retries, and sometimes auth locally, giving the consistency benefits of centralization (one proxy implementation, uniformly deployed) without the single-chokepoint blast radius, since a sidecar crash only affects its own pod, not the whole fleet, at the cost of running many more proxy instances and a more complex control plane to configure them all consistently. Netflix's Zuul-then-service-mesh evolution and many large platforms' move from a monolithic API gateway toward mesh-plus-thin-edge-gateway architectures are real-world instances of organizations making exactly this trade-off as they outgrew a single chokepoint's blast radius and bottleneck.

  • How does a service mesh's sidecar model change the single-point-of-failure story compared to a single central gateway?
    With sidecars, each service instance has its own local proxy handling TLS, retries, and often mTLS-based auth, so a sidecar crash or bug only affects that one pod's traffic, not the entire fleet, unlike a central gateway whose failure is total. The trade-off is operational complexity: instead of managing one gateway, you're managing potentially hundreds of proxy instances that all need to be configured, upgraded, and monitored consistently via a control plane.
  • What organizational pattern helps avoid the gateway team becoming a bottleneck for every other team's needs?
    Making the gateway self-service and declaratively configurable, so service teams can define their own rate-limit shapes, auth scopes, or cache rules via configuration or a policy file rather than requiring the gateway team to write custom code for each request. This shifts the gateway team's role from implementing every change to maintaining the platform and guardrails that make self-service safe.
  • Why is it a design-discipline risk for a gateway to grow business-authorization or response-composition logic over time?
    Those concerns have fundamentally different characteristics from pure cross-cutting offloading (TLS, coarse auth, rate limiting): they require domain knowledge and touch multiple backend calls, which couples the gateway to business logic changes and gives it different scaling and failure behavior than a stateless traffic-shaping layer. Left unchecked, the gateway becomes a second monolith that's harder to reason about and test than the services it was meant to simplify.

Like routing every department's mail through one central mailroom instead of each department handling its own: far more consistent and efficient, until the mailroom itself jams, and suddenly no department gets any mail at all, not just the one that had a problem.

saying these in an interview costs you the question

  • treats centralizing everything into the gateway as a strictly free win with no cost
  • can't articulate why a gateway outage is worse than a single backend service outage
  • unaware of service mesh as an alternative decentralized approach to the same concerns
  • doesn't distinguish organizational/ownership cost from purely technical cost
  • would centralize every concern regardless of latency sensitivity or per-service nuance needs

context