Introspection adds a network call to every request. How do you design a Spring resource server to keep opaque-token validation performant, and what does that cost you?
answer
- Per-request POST = bottleneck + SPOF
- Decorate introspector with Caffeine cache
- Key = hashed token; TTL <= exp
- Cost = revocation delayed to TTL
- Timeouts + fail-closed policy
basics
~20 sWrap the OpaqueTokenIntrospector in a cache keyed by the token so repeated requests skip the introspection call. Bound the cache by a short TTL and the token's exp. The cost is that a revoked token stays accepted until its cache entry expires.
solid answer
~50 sBy default SpringOpaqueTokenIntrospector calls the RFC 7662 endpoint on every request, making it the latency and throughput bottleneck. The standard fix is a custom @Bean OpaqueTokenIntrospector that decorates the real introspector with a cache (e.g. Caffeine) keyed on the token string, storing the resulting OAuth2AuthenticatedPrincipal. Bound entries by a short TTL and never beyond the token's exp so you don't serve expired tokens. This collapses N requests with the same token into one introspection call, cutting latency and endpoint load and reducing the availability coupling to the authorization server. The trade-off is revocation latency: a token invalidated centrally is still honored until its cache entry evicts, so the TTL is a tunable knob between immediacy and performance. You should also configure sensible connect/read timeouts and a failure policy so a slow AS doesn't stall request threads. Cache keys are sensitive (they are live tokens), so hash them and keep the cache in-memory/short-lived.
code
java · 13 lines@Bean
OpaqueTokenIntrospector cachingIntrospector(OAuth2ResourceServerProperties props) {
OAuth2ResourceServerProperties.Opaquetoken p = props.getOpaquetoken();
OpaqueTokenIntrospector delegate = new SpringOpaqueTokenIntrospector(
p.getIntrospectionUri(), p.getClientId(), p.getClientSecret());
Cache<String, OAuth2AuthenticatedPrincipal> cache = Caffeine.newBuilder()
.expireAfterWrite(Duration.ofSeconds(30))
.maximumSize(10_000)
.build();
return token -> cache.get(token, delegate::introspect);
}go deeper
Know that introspection is a per-request network call and caching helps.
Should describe wrapping the introspector in a cache and that it delays revocation.
Should design the TTL<=exp bound, token hashing, timeouts, and fail-closed policy, and frame TTL as a revocation/perf dial.
Weighs distributed vs local cache, circuit breaking, per-endpoint AuthenticationManagerResolver, and documents cache TTL as a security SLA.
## The problem `SpringOpaqueTokenIntrospector` performs an HTTP `POST` to the introspection endpoint **for every authenticated request**. That means: added p99 latency, the endpoint becomes a **throughput ceiling** and a **single point of failure**, and each resource-server instance hammers the AS. ## Caching decorator Wrap the delegate in your own bean and cache the returned principal keyed by the token: ```java @Bean OpaqueTokenIntrospector cachingIntrospector(OAuth2ResourceServerProperties props) { var p = props.getOpaquetoken(); OpaqueTokenIntrospector delegate = new SpringOpaqueTokenIntrospector( p.getIntrospectionUri(), p.getClientId(), p.getClientSecret()); Cache<String, OAuth2AuthenticatedPrincipal> cache = Caffeine.newBuilder() .expireAfterWrite(Duration.ofSeconds(30)) .maximumSize(10_000) .build(); return token -> cache.get(token, delegate::introspect); } ``` ### Design points - **Key = the token** (ideally a hash of it — the raw token is a live credential you don't want sitting in plaintext in cache/heap dumps). - **TTL bound**: `min(shortTTL, exp - now)`. Never cache past `exp`, or you'd accept an expired token. A caching layer that respects `exp` avoids serving structurally-expired tokens. - **Negative results**: decide whether to cache `active:false` briefly to blunt token-guessing storms, or never cache failures. - **Eviction size** caps memory. ## The cost: revocation latency Caching directly weakens opaque tokens' headline benefit — instant revocation. A token revoked at the AS remains accepted by the resource server **until its cache entry evicts**. So the TTL is an explicit dial: `TTL → 0` restores immediate revocation but full per-request cost; larger TTL improves performance but widens the revocation window. Choose based on how fast revocation must propagate for the domain. ## Resilience knobs - Set **connect/read timeouts** on the HTTP client the introspector uses; a hanging AS otherwise blocks a request thread (and, on the servlet stack, a container thread). - Decide a **fail-closed vs fail-open** policy on introspection errors. Fail-closed (401/503) is the secure default; fail-open is almost never acceptable for auth. - Consider a **circuit breaker** so a struggling AS sheds load rather than cascading. ## Alternatives / complements - **Move to JWTs** for the hot paths if revocation immediacy isn't required — no network at all. - **Distributed cache** (Redis) shared across instances lowers AS load further but adds its own latency/consistency and a new dependency; usually a per-instance in-memory cache is enough. - **AuthenticationManagerResolver** to introspect only for high-value endpoints while trusting short JWTs elsewhere. ## Gotchas - Declaring the bean **replaces auto-config** — rebuild the delegate from `OAuth2ResourceServerProperties`. - Don't let the cache outlive `exp`; don't log raw tokens as cache keys. - Caching makes revocation eventual — document the TTL as a security property, not just a perf tweak.
- What security property do you weaken by caching introspection responses, and how do you bound it?Immediate revocation: a centrally-revoked token is still accepted until its cache entry evicts. You bound the exposure with a short TTL and by never caching past the token's exp, treating TTL as an explicit revocation-latency budget.
- Why hash the token before using it as a cache key?The raw token is a live bearer credential; storing it in plaintext in a cache (and in heap/memory dumps) broadens its exposure. Hashing the key preserves lookup while not persisting the usable credential.
- How do you prevent a slow authorization server from exhausting request threads?Set connect/read timeouts on the introspector's HTTP client, fail closed on timeout, and optionally add a circuit breaker so a degraded AS sheds load rather than blocking threads and cascading.
saying these in an interview costs you the question
- Caching without bounding by exp (accepting expired tokens)
- Storing raw tokens as plaintext cache keys / logging them
- Fail-open on introspection errors (accepting requests when the AS is unreachable)
- Claiming caching keeps revocation instantaneous