skip to content

How does an API Gateway offload authentication and authorization so that individual backend microservices don't each have to implement their own login/token-checking logic, and what does the gateway typically pass downstream once a request is verified?

level: middleimportance: must knowfreq 75%

answer

  1. JWT signature check at edge
  2. opaque token → introspection call
  3. inject trusted headers downstream
  4. coarse vs fine-grained authz split
  5. lock down backend network path

basics

~20 s

The gateway checks the caller's login token once, before the request goes anywhere else. If it's valid, the gateway lets the request through and tells the backend service who the caller is; if not, it rejects the request immediately.

solid answer

~40 s

The gateway terminates the auth concern at the edge: it validates an incoming credential — usually a JWT's signature/expiry, or a call out to an auth server for opaque-token introspection — before forwarding the request. If invalid, it returns 401/403 without the request ever reaching a backend service. If valid, it injects trusted, gateway-verified claims as headers (e.g., X-User-Id, X-Roles) so downstream services can trust identity without re-verifying signatures themselves. Coarse-grained authorization (route-level: 'is this caller allowed to hit this path at all') often happens at the gateway too, while fine-grained, resource-level authorization ('can this user edit this specific order') stays in the owning service, since only that service has the business context to decide.

go deeper

for a junior

Knows the gateway checks the login token before requests reach services and rejects bad requests early.

for a middle

Can explain JWT signature validation vs opaque-token introspection at a high level and knows the gateway passes identity downstream via headers.

for a senior

Reasons about the trust boundary this creates (backend services must be network-locked to the gateway), the coarse/fine-grained authorization split, and cache-staleness trade-offs for introspection.

for a principal

Weighs centralizing auth against the systemic blast radius of a gateway auth bug/outage, and can design the key-rotation and revocation-latency policies that keep the trade-off acceptable at scale.

## How the gateway proves who a caller is Authentication offloading means moving the work of proving who a caller is out of every individual microservice and into the single edge layer every request must pass through — the API Gateway. Mechanically this usually works one of two ways. - **With self-contained tokens (JWTs)**, the gateway holds the issuer's public key (or a JWKS endpoint it periodically fetches) and, on each request, verifies the token's cryptographic signature, checks the expiry (`exp`) claim, and checks the issuer/audience claims match what's expected — all without any network call, which keeps this fast (sub-millisecond, CPU-bound). - **With opaque tokens** (a random string that means nothing on its own, common with OAuth2 access tokens issued by systems like **Keycloak** or **Okta**), the gateway must call an introspection endpoint on the auth server to ask 'is this token still valid, and for whom' — this is a network round trip on the hot path, so it's often cached briefly (30-60 seconds) to avoid hammering the auth server on every request. Once verification succeeds, the gateway typically strips the raw `Authorization` header before forwarding and instead injects internal, trusted headers like `X-User-Id`, `X-Roles`, or `X-Tenant-Id` that the downstream service reads directly, skipping signature verification entirely because it trusts the gateway as the enforcement boundary. ## Why it exists This exists because re-verifying tokens in every service is both wasteful and risky. - **Wasteful** because ten services all parsing and validating the same JWT on the same request is pure duplicated CPU work and duplicated code to keep patched and correct. - **Risky** because security logic copy-pasted across a dozen codebases inevitably drifts — one service updates its JWT library to fix a vulnerability, another doesn't, and now there's an inconsistent security posture across the fleet. Centralizing at the gateway means there's exactly one place where the 'is this caller who they claim to be' logic lives, one place to patch, one place to audit. ## The trade-off: trust boundary and blast radius The trade-off is a shift in trust boundary and blast radius. - Every downstream service must now unconditionally trust the gateway's injected headers, which means the network path between the gateway and the backend services must be locked down (private network, mTLS, or at minimum non-spoofable — a service directly reachable from outside the gateway would let an attacker forge `X-User-Id` headers and impersonate any user). This is why gateway-fronted architectures usually also firewall backend services so they're unreachable except through the gateway. - There's also a **coupling cost**: if the gateway's auth logic has a bug or an outage, authentication for the entire system is down, even though every backend service is individually healthy — the classic single-point-of-failure trade against the benefit of centralization. - And **coarse vs. fine-grained authorization** has a real split: the gateway can cheaply enforce 'does this role have any access to `/admin/*`' but it usually cannot enforce 'can user 42 edit order 917' because that requires domain data (who owns order 917) that only the order service has — pushing that logic into the gateway would leak business knowledge into infrastructure and recreate a monolith at the edge. ## Failure modes in production Failure modes in production include: 1. **Clock skew** between the token issuer and the gateway causing valid tokens to be rejected as expired or not-yet-valid; 2. **A JWKS key rotation** that the gateway hasn't refreshed yet, causing a spike of false 401s right after the auth server rotates signing keys; 3. **The introspection-cache-staleness problem**, where a token that was just revoked (user logged out, or was suspended) is still accepted by the gateway for up to the cache TTL because it hasn't re-checked with the auth server yet — a real security/UX trade-off that teams tune deliberately. ## Where you have seen it A concrete real-world example: **Kong Gateway's** JWT and OAuth2 plugins do exactly this — they validate the incoming token at the plugin layer before the request is proxied to the upstream service, and inject consumer identity as headers/context that the upstream can read. **AWS API Gateway** offers the equivalent via Lambda authorizers or built-in Cognito/JWT authorizers in front of API routes, so a Lambda function serving business logic never has to parse a JWT itself — API Gateway has already validated it and passed the claims through in the event payload.

  • What stops an attacker who gets direct network access to a backend service from just forging the X-User-Id header the gateway normally injects?
    Nothing at the application layer if the network allows it — this is why backend services must be network-isolated so they're only reachable from the gateway (private subnet, firewall rules, or mTLS between gateway and services). The trust in the injected header is only as strong as the guarantee that no other path can reach the service.
  • Why would you cache the result of an opaque-token introspection call at the gateway instead of calling the auth server on every request?
    Introspection is a network round-trip, and doing it synchronously on every single request adds latency and load-tests the auth server at your full traffic volume. A short TTL cache amortizes that cost while bounding how long a revoked token can still be accepted, trading a small security window for much better throughput and resilience.
  • Why doesn't the gateway typically enforce resource-level authorization like 'can this user edit this specific order'?
    That decision needs business/domain data — who owns the order, what state it's in — that lives inside the owning service's database, not in the gateway's routing config. Pushing it into the gateway would force the gateway to know about every service's data model, recreating tight coupling and defeating the purpose of keeping the gateway generic.

Like a building's front-desk security checking ID badges: once you're past the front desk with a verified badge, the floors inside don't re-check your ID at every door — they trust the visitor sticker the front desk already gave you.

saying these in an interview costs you the question

  • Believes the gateway removes the need for any authorization logic in backend services
  • Doesn't mention that backend services must trust/verify the network path to the gateway
  • Confuses authentication (who you are) with authorization (what you can do)
  • Assumes opaque token validation is free/instant with no network cost
  • No awareness that revoked tokens can still be honored briefly if introspection results are cached

context