Exchanging the caller's token once at the edge or again at every hop: what does each cost on a four-service donation chain?
answer
- one exchange, or one per hop
- an edge token is spendable across the zone
- each hop adds an issuer round trip
- cache key must include the subject
- narrow where the blast radius changes
basics
~20 sExchanging once at the edge costs one round trip but leaves the inward token spendable at every service behind it. Exchanging again at each hop narrows the token to its callee and adds a round trip and a failure point per hop.
solid answer
~50 sExchange once at the edge and the chain is cheap: one call to the authorization server per inbound donation, and the inward token carries the donor onward to the ledger, reconciliation and statement generation. What you give up is bounding — that token is spendable at all four services, so anything that captures it at the ledger can call statement generation directly. Exchange again at every hop and each service asks for a token whose `aud` is only its callee, so a captured token reaches one destination. You pay three more synchronous calls on the critical path, and the authorization server's exchange endpoint becomes a dependency in series with the whole chain. Caching helps less than people expect: the key must include the subject, and a token that lives seconds gets no reuse on a request that crosses each hop once. Most chains land on a hybrid — narrow at the hop that moves money, share one audience elsewhere.
code
pseudocode · 19 lines# run by a service before it calls the next one in the chain
function tokenForNextHop(inboundToken, callee, scopeNeeded):
subject = subjectOf(inboundToken)
key = (thisService, subject, callee, scopeNeeded)
cached = exchangeCache.get(key)
if cached != null and cached.expiresAt - now() > SAFETY_MARGIN:
return cached.token
issued = authorizationServer.exchange(
presenting = inboundToken,
targetAudience = callee,
scope = scopeNeeded)
ttl = min(issued.expires_in - SAFETY_MARGIN, HARD_CAP)
if ttl > 0:
exchangeCache.put(key, issued.token, ttl)
return issued.tokengo deeper
Recall that the token the donor presented is not automatically the token the fourth service should see, and that something has to carry the donor's identity along the chain deliberately.
Explain the two shapes and their mechanics: one exchange at the boundary with a zone-wide audience, against an exchange per hop with an audience naming one callee, and what each means if a token is captured mid-chain.
Demonstrate the production judgment: quantify the added round trips, name the cache key including the subject, and describe what an exchange-endpoint outage does to work already admitted versus work still at the door.
Take the position: which hops genuinely differ in blast radius, what you will pay to bound those and not the others, and how you keep the narrowing honest as services are added to the chain.
## The two topologies One donation crosses intake, the ledger, reconciliation and statement generation before anything is written back, and the fourth service still has to know on whose behalf it is acting. The specification for token exchange (RFC 8693) tells you how to ask for a token in someone's name; it deliberately does not tell you *how often* to ask. That is the design decision, and there are two honest shapes with a spectrum between them. **Exchange once at the edge.** The gateway obtains one inward token per inbound request, and every hop forwards it unchanged. Cost: one exchange per request, one extra dependency, one place to reason about. **Exchange again at every hop.** Each service presents the token it received and asks for a new one whose audience is the service it is about to call. Cost: an exchange per edge in the call graph, all of them on the critical path. ## What edge-once leaves open An edge-minted token has to be acceptable to every service behind the edge, which means its `aud` names the whole zone rather than one callee. The consequence is that identity propagation becomes **transitive**: anything that can read the token at reconciliation can present it to statement generation, or back to the ledger, and be served. The chain's reach is the union of everything in the zone, not the path the request actually took. That is not automatically wrong. If the four services share a blast radius already — same deploy, same data, same on-call — a zone audience describes the truth. It becomes wrong the moment one hop is materially more dangerous than the others. ## What per-hop costs, stated concretely | | exchange once at the edge | exchange at every hop | |---|---|---| | exchanges per donation | 1 | 1 per edge in the call graph | | added latency | one round trip, before any work | one round trip per hop, inside the request | | token reach if captured mid-chain | every service in the zone | the one callee it names | | issuer load | request rate | request rate multiplied by chain depth | | issuer outage | new requests cannot enter; in-flight ones finish | every internal call stops, including admitted work | The last row is the one that decides most arguments. Under edge-once, a degraded exchange endpoint fails requests at the door and the system stays consistent. Under per-hop, work that has already been admitted and has already written a provisional ledger entry cannot reach reconciliation, and you are now handling partial donations rather than refused ones. ## Caching the exchanged token Caching is the first reflex and it is worth doing carefully, because the cache key is a security boundary: - The key must include **the subject** — which donor this token speaks for — as well as the calling service, the callee and the scope asked for. A cache keyed only by callee and scope will serve one donor's request a token minted for another donor. That is not a performance bug, it is an identity bug, and it will look like a data-mixing incident long before anyone suspects the cache. - The entry must never outlive the token: take the smaller of the token's remaining lifetime minus a safety margin, and a hard cap. An expired token served out of a cache turns into a mid-chain refusal. - Expect a poor hit rate. Cardinality is per donor per hop, and the lifetime is seconds. A request that crosses each hop exactly once never reuses anything. The cache pays only where one service calls the same callee several times inside a single request, or where the same donor is active continuously. ## What most chains actually do The defensible answer is neither extreme. Exchange at the edge so the donor's own credential stops there, then **re-narrow only at boundaries where the blast radius genuinely changes** — the hop into the ledger that can move money, the hop that reaches donor contact details — and let the rest of the zone share one inward audience. That buys the bounding where it matters and pays for it three times instead of ten. Two things make the hybrid work. First, each callee must actually refuse a token whose `aud` is not itself: the gateway or the caller obtains the narrowed token, but only the callee can enforce it, and an audience nobody checks is decoration. Second, the narrowing has to be recorded somewhere a reader can see the call graph — otherwise the next service added to the chain quietly inherits the zone audience and the bound you paid for stops describing the system.
- What must the exchanged-token cache be keyed by, and what goes wrong if the subject is left out?Caller, subject, callee and requested scope. Leave the subject out and the key stops distinguishing donors: a hit serves a token minted for whoever populated the entry, so one donor's request acts as another. It surfaces as mixed-up ledger entries, not as a cache miss, which is why it is usually found late.
- Why does caching an exchanged token pay off less than people expect?Its cardinality is one entry per caller, subject, callee and scope, and its lifetime is seconds. A donation that crosses each hop exactly once never gets a second look at the same entry. The cache earns its place only on fan-out — one service calling the same callee repeatedly inside a single request — or for a caller that is continuously busy for the same subject.
- What does an outage of the exchange endpoint do to each topology?Under edge-once, requests are refused at the boundary and work already admitted runs to completion — a clean failure. Under per-hop, calls fail wherever the chain has got to, so donations stall mid-flight with provisional state written and nothing to finish it. The second failure is much harder to clean up than the first.
- A fifth service is added behind the edge. What changed about the edge-once topology?Its reach. The inward token's audience covers the zone, so the new service is now a destination for every token already in flight, without anyone editing a caller. That silent widening is the cost of a zone audience, and it is why the money-moving hops are usually the ones pulled out into their own audience first.
A visitor badge that opens every door on the floor, against a fresh badge issued at each door for the next room only. The second is safer and there is a queue at every door.
saying these in an interview costs you the question
- Says forwarding the inbound token downstream costs nothing
- Exchanges at every hop without measuring the added tail latency
- Caches the exchanged token keyed only by the callee and scope
- Writes a cache entry that outlives the token it holds
- Assumes the authorization server absorbs chain depth for free
- Calls an edge-minted token bounded when its audience covers the whole zone