skip to content

In a production system, what are the specific ways a gateway's routing configuration goes wrong, and how does each one actually show up to an on-call engineer?

level: seniorimportance: should knowfreq 45%

answer

  1. shadowing = broad rule steals specific rule's traffic
  2. stale rule = points at a service that's gone
  3. rewrite bug = right backend, wrong path
  4. propagation lag = different gateway instances, different rules, same moment

basics

~20 s

Routing rules can overlap and steal each other's traffic, point at a service that's gone, or fall out of sync across gateway instances. To on-call, this looks like plain errors with no obvious cause, since the backends are healthy.

solid answer

~40 s

The recurring failure modes are: rule shadowing, where a broad rule matches before a more specific one and silently steals its traffic; stale rules, where a rule still points at a decommissioned backend, producing connection errors that look like a backend outage even though the backend, if it existed, would be fine; rewrite bugs, where the gateway matches correctly but rewrites the path wrong before forwarding, so the backend 404s on what looks like a valid request; and config propagation lag, where a change is live on some gateway instances and not others, producing inconsistent behavior for the same request depending on which instance handled it. All present as generic HTTP errors identical to a backend problem, so the gateway layer often gets ruled out last rather than first.

go deeper

for a junior

Should recognize that a wrong routing rule can send a request to the wrong place and that this is different from a backend crashing.

for a middle

Should name at least two distinct failure modes (e.g. stale rule, rewrite bug) and how each looks different from a backend outage.

for a senior

Should walk through several failure modes with their concrete production symptoms and propose at least one mitigation (per-request routing logs, staged rollout of rule changes).

for a principal

Should design process and tooling around routing-table changes at organizational scale, review/validation gates, canarying routing changes themselves, and reducing propagation-lag windows across a large gateway fleet.

## Configuration that fails the way code fails A gateway's routing table is effectively logic expressed as configuration rather than code, and it fails the same broad ways code fails, **wrong precedence**, **stale references**, **incorrect transformations**, and **inconsistent rollout**, except that these failures happen against live production traffic with no compile step or test suite catching them first, and they present to on-call as generic, ambiguous errors rather than an obvious 'the routing table is wrong' signal. ## Rule shadowing The first and most common failure mode is rule shadowing. Routing tables are ordered or prioritized lists, and when two rules overlap (a broad prefix like `/api/**` alongside a specific one like `/api/orders/**`), the evaluation order or precedence policy determines which one actually fires. If the gateway uses first-match-wins semantics and the broad rule happens to be listed before the specific one, every request meant for the specific rule gets silently intercepted by the broad one instead, and unless the two rules point at meaningfully different backends, this can go unnoticed for a long time, since requests still get a response, just from the wrong service. It surfaces in production as: - requests that appear to succeed but return unexpected data or schema, - or one backend service mysteriously reporting far more request volume than its actual callers should produce, while the intended backend reports suspiciously low or zero traffic. ## Stale rules The second failure mode is stale rules: a rule that was correct when written but now points at a backend that has since been renamed, moved, or decommissioned. This typically happens after a service migration or decommissioning where the routing table update was missed or done incompletely. To on-call, this shows up as connection-refused errors, `502 Bad Gateway`, or DNS resolution failures on requests to a specific path, and it is easy to misdiagnose as 'the backend is down' when in fact the backend, if you tried to reach it directly, might not even exist anymore, or exists under a different address the gateway was never updated to point at. The **diagnostic tell** is that the backend's own health checks and monitoring show nothing wrong, because the backend genuinely isn't receiving the traffic at all; the failure is entirely upstream of it, at the gateway. ## The path-rewrite bug The third failure mode is a path-rewrite bug. Many gateways don't just forward a matched request as-is, they rewrite the path before forwarding (stripping a prefix, adding a version segment) so the backend sees a cleaner internal path than the one the client called. If the rewrite rule is subtly wrong, say it strips one segment too many or too few, the gateway correctly identifies the right backend and forwards the request, but the backend then returns a 404 or routing error of its own, because the path it received doesn't match any of its own internal routes. This is a particularly confusing failure because both the gateway's routing decision and the backend's own health are correct in isolation; the bug lives entirely in the transformation between them, and the fix requires someone to actually trace a single request's exact path through the gateway rewrite logic rather than checking gateway-up/backend-up status independently. ## Config propagation lag The fourth failure mode is specific to horizontally scaled gateway deployments: config propagation lag. When a routing change (a new rule, a canary weight bump, a rollback) is pushed, it typically needs to reach every gateway instance in the fleet, and depending on the distribution mechanism (a config file redeployed per-instance, a shared config store polled periodically, a push-based control plane), there can be a window, seconds to minutes, where some gateway instances are serving the old routing table and others the new one. During that window, identical requests from the same client can be routed differently depending purely on which gateway instance a load balancer happened to send them to, producing intermittent, seemingly random behavior that's very hard to reproduce on demand because it depends on backend instance selection that the person debugging doesn't control or observe directly. ## When shadowing hides a latent dependency A concrete real-world illustration combining a couple of these: a team decommissions an old backend service after a migration, updates the routing table to remove its rule, but forgets that a broader catch-all rule further down the list was relying on requests never reaching it because the specific rule intercepted them first (shadowing masking a latent dependency); once the specific rule is removed, the catch-all rule suddenly starts receiving traffic it was never designed to handle, and a service that had zero production traffic for months, and therefore zero recent operational attention, is abruptly serving live requests, often revealing bugs or capacity issues that had been dormant. ## The operational lesson The general operational lesson is that routing-table changes deserve the same rigor as code changes: - staged rollout, - the ability to diff old vs. new rules before applying, - canarying the routing change itself where possible, - and treating 'which backend actually served this request' as a first-class piece of observability (logged per request) rather than assuming it from the URL alone.

  • Why is rule shadowing particularly dangerous compared to a rule that fails loudly?
    Because the shadowed request still gets a response, just from the wrong backend, it can look like a successful call in dashboards and logs that only track HTTP status codes rather than which specific backend actually served the request. It can persist undetected for a long time until someone notices a data mismatch or an unexpected spike in one service's traffic relative to another.
  • How would you design routing-table changes to catch shadowing or stale-rule bugs before they hit production?
    Treat the routing table like code: run it through a linter or validator that flags overlapping prefixes and rules pointing at unknown/unregistered backends, and diff the proposed change against the current live table before applying it. Ideally, test the rule change against a shadow or staging traffic sample before promoting it, the same discipline used for canarying application code.
  • What single piece of observability would most help diagnose these failure modes quickly?
    Logging, per request, which specific routing rule matched and which backend target actually received the request, not just the client-facing status code. Without that, on-call has to infer the routing decision indirectly from symptoms, which is exactly what makes these failures slow to diagnose.

It's like a company directory with outdated forwarding instructions: some calls get silently redirected to the wrong department because an old blanket rule catches them first, some calls ring a desk that was cleared out months ago, and during an update to the directory itself, different receptionists are working from different versions of it at the same time.

saying these in an interview costs you the question

  • Thinks routing configuration can't have bugs the way code can
  • Doesn't distinguish a gateway-layer failure from a genuine backend outage
  • Has no answer for how rule ordering/precedence can cause misrouting
  • Assumes a routing-table update reaches every gateway instance instantly
  • No mention of logging which rule/backend actually served a request as a diagnostic tool

context