skip to content

What does gateway routing cost a system in exchange for the topology decoupling it provides, and when would you deliberately avoid introducing it?

level: seniorimportance: must knowfreq 55%

answer

  1. extra hop = latency cost
  2. gateway becomes an SPOF unless made redundant
  3. routing table = config that can silently misroute
  4. skip it for tiny systems / latency-critical internal calls / early-stage churn

basics

~20 s

Every request now takes an extra hop through the gateway, adding latency, and the gateway itself must stay up and correctly configured or nothing behind it is reachable. Not worth it for a tiny system with one or two services.

solid answer

~40 s

Gateway routing trades client-side simplicity for a new piece of shared infrastructure on the critical path. Costs: an added hop and its latency on every request; the gateway becoming a single point of failure unless made redundant and horizontally scaled; a routing table that is now configuration which can drift, go stale, or misroute via overlapping rules; and an extra system to secure and operate. It pays off with multiple independently deployable services, version coexistence, blue-green/canary control, or hiding topology from many external clients. It's often not worth it for a small system with one or two services, for latency-critical internal calls where a service mesh sidecar is cheaper, or early on when topology is still churning too fast for stable rules.

go deeper

for a junior

Should recognize that adding a gateway adds a network hop and something new that must stay running.

for a middle

Should name the SPOF risk and the idea that routing config can be wrong, plus a basic sense of when it's overkill (very small system).

for a senior

Should give a concrete cost/benefit trade-off with named failure modes and articulate at least one specific scenario for deliberately not using a centralized gateway (e.g. internal mesh traffic).

for a principal

Should reason at the org/system-design level about where to scope gateway routing (edge vs. internal), how to keep the gateway from becoming a shared bottleneck across teams, and how to evolve away from an over-centralized routing layer as the system matures.

## The trade Gateway routing is not free; it buys **topology decoupling** and **centralized release control** at the cost of a new mandatory dependency sitting between every client and every backend, and a senior engineer should be able to name both sides of that trade concretely rather than treating the pattern as an unconditional best practice. ## The first cost: latency The first cost is latency. Every request that would otherwise go directly from client to backend now makes two hops: client to gateway, gateway to backend. For a well-operated gateway this is typically single-digit milliseconds of added round-trip time, but it is not zero, and it compounds when gateways are chained (an edge gateway in front of an internal gateway in front of a service mesh, which happens in larger organizations more often than architects intend). For most user-facing APIs this overhead is negligible against network and backend processing time, but for latency-sensitive internal service-to-service calls, especially high-fan-out calls where one request triggers dozens of downstream calls, that per-hop cost adds up and is a real reason many systems use gateway routing only at the edge (external-facing) while using direct service discovery or a sidecar mesh for internal east-west traffic. ## The second cost: availability risk concentration The second cost is availability risk concentration. Because every client request must pass through the gateway, the gateway's own uptime becomes a hard ceiling on the uptime of everything behind it: a backend can be perfectly healthy and still be completely unreachable if the gateway in front of it is down, overloaded, or mid-deployment. This is manageable, the standard answer is to run the gateway itself as a horizontally scaled, redundant fleet behind its own load balancer, with no single instance being load-bearing, but it means the gateway is not a free abstraction layer; it is a new tier of infrastructure that inherits the same availability requirements as the system it fronts, and someone has to own its capacity planning, deployment safety, and on-call. ## The third cost: configuration risk The third cost is configuration risk. The routing table is logic, even though it's declarative rather than code, and like any logic it can be wrong: - overlapping prefixes that shadow each other, - a rule left pointing at a decommissioned backend, - a rewrite rule that strips the wrong path segment. Unlike a bug in application code, a routing misconfiguration doesn't fail a unit test, it fails silently against live traffic, often only for a subset of paths, and the failure mode looks identical to a backend outage from the client's point of view (connection refused, 502, 404), which makes it slower to diagnose because the on-call engineer has to first rule out the gateway layer before looking at the backend that's actually fine. ## When it pays for itself Against these costs, the benefit is real and specific: gateway routing is worth introducing - when you have genuinely independent backend services whose internal topology needs to be free to change without coordinating every client, - when you need to run multiple versions of a service side by side (blue-green, canary, gradual API version migration), - or when many external, non-cooperative clients (mobile apps that can't be force-updated, third-party integrations) need a stable contract while the backend evolves underneath it. In those cases the alternative, every client independently discovering and calling backends directly, is strictly worse: it pushes topology-change coordination onto every caller and makes version-controlled rollouts impossible to do safely and centrally. ## When to skip it It is not worth it, or at least not worth introducing at the edge for every hop, in a few concrete situations. 1. **A small system with one or two services** and no near-term plan for splitting further. Adding a gateway is pure overhead: there's no topology to hide and no version-coexistence need, just an extra hop and an extra thing to operate. 2. **Latency-critical internal calls** between services that already trust each other and are deployed together. Many teams use a service mesh's sidecar proxies for routing and resilience instead of routing every internal call back out through a centralized gateway, because the sidecar sits in-process-adjacent (same host or pod) and avoids the extra network hop a centralized gateway implies. 3. **Early in a system's life**, when the service boundaries themselves are still being discovered and reshaped week to week, a routing table churns faster than it can be trusted. Teams often defer introducing a formal gateway until the service boundaries have stabilized enough that routing rules aren't rewritten in every sprint. ## Getting the scope wrong A concrete illustration of getting this wrong: a team that puts a single shared gateway on the critical path for high-volume internal service-to-service calls (not just external client traffic) can turn a routine gateway redeploy into a system-wide latency spike or outage, because internal calls that never needed the decoupling benefit are now coupled to the gateway's availability anyway. The fix in practice is usually to scope gateway routing to the boundary where the decoupling benefit is actually being consumed (external client traffic, or the specific paths under active migration) rather than routing every call, internal and external, through one shared choke point by default.

  • How would you make a gateway itself not a single point of failure?
    Run multiple gateway instances behind their own load balancer or DNS-based failover, with health checks removing unhealthy instances automatically, so no single gateway process is load-bearing. The routing configuration also needs to be distributed consistently and quickly to all instances, otherwise you trade a hard outage for a subtler inconsistent-routing problem during config propagation.
  • Why might a team route external traffic through a gateway but skip it for internal service-to-service calls?
    External traffic benefits most from topology decoupling and centralized version control since callers (mobile apps, partners) can't be coordinated the way internal teams can. Internal calls are typically latency-sensitive and high-volume, and a service mesh sidecar can provide similar routing/resilience benefits without adding a shared network hop, so many architectures reserve the centralized gateway for the edge.
  • What's a sign in production that a routing table has become too risky to change safely?
    If routing changes require a full gateway redeploy rather than a config push, if there's no way to validate a rule change against real traffic before it goes live, or if rule ordering isn't reviewed and overlapping prefixes have caused incidents before, that's a signal the routing configuration needs its own testing and review process, similar to application code.

Putting a security checkpoint at a building's single entrance makes it easy to control who goes where, but now everyone waits in that one line, and if the checkpoint itself jams, nobody gets in even if every office inside is fully staffed and ready.

saying these in an interview costs you the question

  • Presents gateway routing as strictly beneficial with no cost
  • Can't identify the gateway itself as a potential single point of failure
  • Doesn't recognize routing rules as configuration that can be wrong/stale
  • Would add a gateway to a two-service system with no articulated reason
  • No awareness that internal, latency-sensitive calls sometimes deliberately bypass a centralized gateway

context