When an API gateway offloads authentication so backend services don't each implement login/token-checking logic, what typically gets validated at the gateway, and what does the gateway pass downstream so a service still knows who is calling it?
answer
- gateway checks token once
- backend trusts injected headers
- isolation makes trust valid
- coarse auth at gateway, fine-grained at service
- forged-header risk if bypassable
basics
~20 sThe gateway checks the caller's token or API key once, and if it's valid, forwards the request to the backend with a header saying who the caller is, so the backend trusts that header instead of re-checking credentials itself.
solid answer
~40 sOffloading authentication means the gateway, not each service, validates the incoming credential: it checks a JWT's signature and expiry, validates an API key against a store, or completes an OAuth2 token introspection call. If valid, the gateway injects trusted identity information, commonly a user ID, subject claim, or roles, into internal headers (like X-User-Id or X-Auth-Roles) and forwards the request. Backend services then trust those headers rather than re-parsing tokens themselves, which simplifies every service and centralizes credential logic, key rotation, and auth-provider integration in one place. The critical risk is that if the internal network is reachable directly, bypassing the gateway, a caller could forge those trusted headers; mitigations include network isolation so backends are unreachable except via the gateway, or having the gateway sign an internal token backends verify.
go deeper
Should know the gateway checks credentials once and that backends receive some form of identity info instead of raw tokens.
Should describe concretely what gets validated (signature, expiry, key lookup) and name the header-injection mechanism plus the network-isolation dependency.
Should articulate the coarse-vs-fine-grained authorization split, the forged-header risk when isolation fails, and mitigations like signed internal tokens.
Should reason about this as a trust-boundary design decision across the whole platform: when header-based trust is acceptable, when it needs a signed-token upgrade, and how it interacts with zero-trust network initiatives and multi-tenant isolation guarantees.
## What the gateway validates **Authentication offloading** moves the work of proving "who is calling" out of individual backend services and into the gateway that sits in front of all of them. Mechanically, a client sends a credential with its request, most commonly a bearer JWT, an OAuth2 access token, an API key, or a session cookie. The gateway intercepts the request before it reaches any backend and performs the validation step, which means: | Credential | What the gateway does with it | |---|---| | **JWT** | checking the cryptographic signature against a known public key or JWKS endpoint, confirming the token hasn't expired, and checking the issuer and audience claims match what's expected | | **Opaque OAuth2 token** | the gateway may call the authorization server's introspection endpoint to confirm the token is still active | | **API key** | it looks the key up against a store and checks it's valid and not revoked | Only after this check passes does the gateway let the request through; if it fails, the gateway rejects with a `401` before the backend ever sees the call. ## What it passes downstream Once validated, the gateway typically strips the original credential and replaces it with something the backend can trust cheaply: a set of internal headers carrying the caller's identity, such as a user ID, tenant ID, or a list of roles/scopes, or occasionally a re-signed internal token. The backend service reads those headers and treats them as ground truth, doing its own authorization logic (is this user allowed to do this specific action) without needing to know anything about JWT signature verification, OAuth2 flows, or session stores. This is the core value proposition: authentication is a **cross-cutting concern** that is identical in mechanism across every service in the system, so implementing it once, well, and centrally avoids dozens of slightly-different, independently-maintained, and independently-vulnerable implementations scattered across a codebase. It also gives a single place to rotate signing keys, swap identity providers, or add a new auth method (say, adding SSO) without touching every service. ## The trade-off — a shift in trust model The trade-off is a shift in trust model. Backend services stop verifying identity themselves and instead trust the network path: specifically, they trust that any request reaching them has already passed through the gateway's checks. That assumption only holds if backends are genuinely unreachable except via the gateway, enforced through: - network policies - security groups - a service mesh that only accepts traffic on the mesh's authenticated channel If that isolation has a gap, for example a backend service also exposed on a public load balancer for debugging, or a misconfigured Kubernetes `NetworkPolicy`, an attacker who can reach the backend directly can simply forge the trusted headers (`X-User-Id: admin`) and skip authentication entirely, since the backend was never taught to be suspicious of them. A more defense-in-depth approach has the gateway mint a short-lived signed internal token (for example, re-signing a minimal JWT) that backends independently verify, so a forged header alone isn't enough. ## Failure modes in production Failure modes in production tend to cluster around a few patterns. 1. **Clock skew or JWKS caching bugs** can cause the gateway to reject valid tokens (false negatives, seen as a spike in 401s) or, worse, accept tokens it shouldn't after a key rotation if the old public key is cached too long. 2. **Network isolation gaps**, as described above, are the most severe failure because they're silent: nothing errors, the system just becomes trivially bypassable, and this is usually only caught by a security audit or penetration test, not by normal monitoring. 3. **Authorization creeping into the gateway layer** beyond simple role-checking is another common issue; teams sometimes try to put fine-grained, resource-specific authorization logic ("can this user edit this specific document") into the gateway, which usually backfires because that decision needs domain data the gateway doesn't have, and it re-couples the gateway to business logic it was supposed to stay out of. The right split is: - the **gateway** handles authentication (who are you, is your token valid) and coarse authorization (role/scope-level gating) - **services** handle fine-grained, resource-level authorization using their own domain knowledge ## Where it shows up in practice A concrete real-world instance of this pattern is Kong or AWS API Gateway configured with a JWT or Cognito authorizer: the gateway validates the token against the identity provider's public keys before forwarding, and downstream Lambda functions or microservices receive claims already extracted into request context, never seeing the raw token or needing an AWS Cognito SDK dependency themselves.
- What happens to security if a backend service behind the gateway is also directly reachable from outside, bypassing the gateway entirely?An attacker can send requests straight to the backend and forge the trusted internal headers the gateway would normally inject, such as claiming an admin user ID, and the backend has no way to distinguish that from a legitimately gateway-validated request. This is why network isolation enforcing gateway-only ingress is a hard requirement for header-based trust to be safe, not an optional hardening step.
- Why shouldn't fine-grained, resource-level authorization (like "can this user edit this specific document") live in the gateway?The gateway generally doesn't have access to the domain data needed to make that call, such as document ownership or workflow state, so pushing that logic there either requires the gateway to query backend data (re-coupling it to business logic) or forces overly coarse rules. It's cleaner to keep the gateway responsible for coarse checks like role or scope, and let the owning service, which already has the domain data, make the fine-grained call.
- What's a safer alternative to plain trusted headers for passing identity from gateway to backend?Have the gateway mint a short-lived, signed internal token (a minimal re-signed JWT, for example) after successful validation, and have each backend independently verify that token's signature rather than blindly trusting an unauthenticated header. This adds defense-in-depth so a network isolation gap alone isn't enough to fully bypass authentication.
Like a nightclub bouncer who checks ID at the door and stamps your hand; staff inside don't re-check your ID, they just look for the stamp, which only works because there's no other way to get inside except past the bouncer.
saying these in an interview costs you the question
- thinks backends need no security controls at all once the gateway authenticates
- can't explain what happens if a backend is reachable outside the gateway
- puts fine-grained per-resource authorization logic entirely in the gateway
- doesn't distinguish authentication (who you are) from authorization (what you can do)
- assumes headers like X-User-Id are inherently trustworthy with no isolation argument