skip to content

Introspection adds a network call to every request. How do you design a Spring resource server to keep opaque-token validation performant, and what does that cost you?

level: seniorimportance: should knowfreq 34%

answer

  1. Per-request POST = bottleneck + SPOF
  2. Decorate introspector with Caffeine cache
  3. Key = hashed token; TTL <= exp
  4. Cost = revocation delayed to TTL
  5. Timeouts + fail-closed policy

basics

~20 s

Wrap the OpaqueTokenIntrospector in a cache keyed by the token so repeated requests skip the introspection call. Bound the cache by a short TTL and the token's exp. The cost is that a revoked token stays accepted until its cache entry expires.

solid answer

~50 s

By default SpringOpaqueTokenIntrospector calls the RFC 7662 endpoint on every request, making it the latency and throughput bottleneck. The standard fix is a custom @Bean OpaqueTokenIntrospector that decorates the real introspector with a cache (e.g. Caffeine) keyed on the token string, storing the resulting OAuth2AuthenticatedPrincipal. Bound entries by a short TTL and never beyond the token's exp so you don't serve expired tokens. This collapses N requests with the same token into one introspection call, cutting latency and endpoint load and reducing the availability coupling to the authorization server. The trade-off is revocation latency: a token invalidated centrally is still honored until its cache entry evicts, so the TTL is a tunable knob between immediacy and performance. You should also configure sensible connect/read timeouts and a failure policy so a slow AS doesn't stall request threads. Cache keys are sensitive (they are live tokens), so hash them and keep the cache in-memory/short-lived.

code

java · 13 lines
java
@Bean
OpaqueTokenIntrospector cachingIntrospector(OAuth2ResourceServerProperties props) {
    OAuth2ResourceServerProperties.Opaquetoken p = props.getOpaquetoken();
    OpaqueTokenIntrospector delegate = new SpringOpaqueTokenIntrospector(
            p.getIntrospectionUri(), p.getClientId(), p.getClientSecret());

    Cache<String, OAuth2AuthenticatedPrincipal> cache = Caffeine.newBuilder()
            .expireAfterWrite(Duration.ofSeconds(30))
            .maximumSize(10_000)
            .build();

    return token -> cache.get(token, delegate::introspect);
}

go deeper

for a junior

Know that introspection is a per-request network call and caching helps.

for a middle

Should describe wrapping the introspector in a cache and that it delays revocation.

for a senior

Should design the TTL<=exp bound, token hashing, timeouts, and fail-closed policy, and frame TTL as a revocation/perf dial.

for a principal

Weighs distributed vs local cache, circuit breaking, per-endpoint AuthenticationManagerResolver, and documents cache TTL as a security SLA.

## The problem `SpringOpaqueTokenIntrospector` performs an HTTP `POST` to the introspection endpoint **for every authenticated request**. That means: added p99 latency, the endpoint becomes a **throughput ceiling** and a **single point of failure**, and each resource-server instance hammers the AS. ## Caching decorator Wrap the delegate in your own bean and cache the returned principal keyed by the token: ```java @Bean OpaqueTokenIntrospector cachingIntrospector(OAuth2ResourceServerProperties props) { var p = props.getOpaquetoken(); OpaqueTokenIntrospector delegate = new SpringOpaqueTokenIntrospector( p.getIntrospectionUri(), p.getClientId(), p.getClientSecret()); Cache<String, OAuth2AuthenticatedPrincipal> cache = Caffeine.newBuilder() .expireAfterWrite(Duration.ofSeconds(30)) .maximumSize(10_000) .build(); return token -> cache.get(token, delegate::introspect); } ``` ### Design points - **Key = the token** (ideally a hash of it — the raw token is a live credential you don't want sitting in plaintext in cache/heap dumps). - **TTL bound**: `min(shortTTL, exp - now)`. Never cache past `exp`, or you'd accept an expired token. A caching layer that respects `exp` avoids serving structurally-expired tokens. - **Negative results**: decide whether to cache `active:false` briefly to blunt token-guessing storms, or never cache failures. - **Eviction size** caps memory. ## The cost: revocation latency Caching directly weakens opaque tokens' headline benefit — instant revocation. A token revoked at the AS remains accepted by the resource server **until its cache entry evicts**. So the TTL is an explicit dial: `TTL → 0` restores immediate revocation but full per-request cost; larger TTL improves performance but widens the revocation window. Choose based on how fast revocation must propagate for the domain. ## Resilience knobs - Set **connect/read timeouts** on the HTTP client the introspector uses; a hanging AS otherwise blocks a request thread (and, on the servlet stack, a container thread). - Decide a **fail-closed vs fail-open** policy on introspection errors. Fail-closed (401/503) is the secure default; fail-open is almost never acceptable for auth. - Consider a **circuit breaker** so a struggling AS sheds load rather than cascading. ## Alternatives / complements - **Move to JWTs** for the hot paths if revocation immediacy isn't required — no network at all. - **Distributed cache** (Redis) shared across instances lowers AS load further but adds its own latency/consistency and a new dependency; usually a per-instance in-memory cache is enough. - **AuthenticationManagerResolver** to introspect only for high-value endpoints while trusting short JWTs elsewhere. ## Gotchas - Declaring the bean **replaces auto-config** — rebuild the delegate from `OAuth2ResourceServerProperties`. - Don't let the cache outlive `exp`; don't log raw tokens as cache keys. - Caching makes revocation eventual — document the TTL as a security property, not just a perf tweak.

  • What security property do you weaken by caching introspection responses, and how do you bound it?
    Immediate revocation: a centrally-revoked token is still accepted until its cache entry evicts. You bound the exposure with a short TTL and by never caching past the token's exp, treating TTL as an explicit revocation-latency budget.
  • Why hash the token before using it as a cache key?
    The raw token is a live bearer credential; storing it in plaintext in a cache (and in heap/memory dumps) broadens its exposure. Hashing the key preserves lookup while not persisting the usable credential.
  • How do you prevent a slow authorization server from exhausting request threads?
    Set connect/read timeouts on the introspector's HTTP client, fail closed on timeout, and optionally add a circuit breaker so a degraded AS sheds load rather than blocking threads and cascading.

saying these in an interview costs you the question

  • Caching without bounding by exp (accepting expired tokens)
  • Storing raw tokens as plaintext cache keys / logging them
  • Fail-open on introspection errors (accepting requests when the AS is unreachable)
  • Claiming caching keeps revocation instantaneous

context