skip to content

Exchanging identity at every hop puts the authorization server on the critical path of every internal call — at what chain depth does that stop being worth it?

level: principalimportance: should knowfreq 27%

answer

  1. count edges, not depth
  2. availability lands in series
  3. narrow where risk differs, not everywhere
  4. an audience nobody checks is decoration
  5. coarse first, narrow the money hop

basics

~20 s

There is no fixed depth. It stops being worth it when the added exchanges and the availability you put in series exceed the reach you actually remove, so narrow at boundaries whose blast radius genuinely differs rather than at every hop.

solid answer

~50 s

Depth is the wrong unit — count **edges in the call graph**, because one hop that fans out to three callees costs three exchanges, not one. Per-hop exchange makes exchange load the inbound request rate multiplied by that edge count, and puts one service's availability in series with every internal call: a chain of five hops is less available than any hop in it, and its p99 carries a round trip per hop. Against that, narrowing only buys something where the callee's authority genuinely differs from its caller's. So the answer is not a depth, it is a map: exchange at the edge, re-narrow at the hops that move money or touch donor records, and let the rest share one inward audience. Two caveats keep it honest — an `aud` no callee checks is decoration, and narrowing bounds what a *stolen* token reaches, not what a *compromised* service can legitimately ask for.

code

pseudocode · 14 lines
pseudocode
# the audience map, held as data rather than scattered through callers
audienceFor = {
    ("intake",         "ledger"):        "ledger",          # money moves: narrow
    ("ledger",         "reconciliation"): "reconciliation", # ledger mutation: narrow
    ("reconciliation", "statements"):     "internal-zone",  # formatting only: shared
    ("reconciliation", "fx-rates"):       "internal-zone",  # formatting only: shared
}

function audienceForEdge(caller, callee):
    target = audienceFor[(caller, callee)]
    if target == null:
        # a new edge nobody classified must not silently inherit the zone
        raise ConfigurationError(caller + " -> " + callee + " has no audience")
    return target

go deeper

for a junior

Recall that making each internal call fetch its own token is not free: something has to issue every one of those tokens, and it has to be up for the call to proceed.

for a middle

Explain the mechanics of the cost: one round trip per call graph edge, added inside the request, and an issuer whose load scales with how many internal calls a single inbound request produces.

for a senior

Show that you have operated this: describe what a slow exchange endpoint does to the chain's tail, where the timeouts surface, and what an outage leaves half-finished under each topology.

for a principal

Take a position with numbers behind it: which hops you will narrow and why, what availability and latency you are buying that with, and why starting coarse and narrowing later is the reversible direction.

## The arithmetic you are signing up for Per-hop exchange is usually argued as a security property and paid for as a capacity decision, so do the capacity sum first. - **Load is edges, not depth.** A donation that crosses intake, the ledger, reconciliation and statement generation is three edges. If reconciliation also calls a currency service and a fraud service, that same request is five. Exchange load is the inbound request rate multiplied by the number of edges in the call graph, and fan-out grows that faster than depth does. - **Availability moves into series.** Under per-hop exchange, every internal call needs the exchange endpoint to answer. Five dependent calls each needing one more service mean the chain is strictly less available than the least available component in it, and the exchange endpoint is now in the denominator of every internal call rather than only of admissions. - **Latency is per hop and lands on the tail.** One round trip per hop, inside the request, with the chain's p99 accumulating each hop's own tail. A chain that was comfortably inside its budget at three hops can miss it at five for no reason a trace will attribute to business logic. - **Degradation hurts before failure does.** An exchange endpoint that is slow rather than down multiplies its added latency by the edge count. Watch exchange latency per hop, not just its availability, or the first symptom will be timeouts attributed to whichever service happens to be last. ## What narrowing buys, and what it does not Narrowing each exchanged token to one callee makes propagation **bounded instead of transitive**. A token captured at reconciliation reaches statement generation and nothing else, rather than reaching every service in the zone. What it does **not** do is contain a compromised service. If reconciliation is owned by an attacker, it can still ask for exactly the token it is entitled to and call statement generation with it, legitimately. Narrowing shortens the reach of a *stolen credential*; it does not shorten the reach of a *service acting within its rights*. Teams that expect the second from it over-buy: they pay per-hop exchange everywhere and are still surprised that a compromised middle service can do what that middle service is allowed to do. ## The failure mode that usually decides it | | exchange once at the edge | exchange at every hop | |---|---|---| | exchange endpoint down | requests refused at the door; admitted work completes | calls fail mid-chain; admitted work stalls | | exchange endpoint slow | one round trip of added admission latency | added latency multiplied by the edge count | | new service added behind the edge | silently becomes a valid destination for tokens in flight | must be given an audience before anything can call it | The first row is what a donation platform actually cares about. A refused donation is a retry. A donation whose provisional ledger entry exists but which can never reach reconciliation is a reconciliation exception with a human on the end of it, and there will be one per in-flight request for the duration of the incident. ## A graded answer rather than a depth Rank the hops by what an attacker who could call them freely would get: 1. **Hops that move money or mutate the ledger** — narrow these individually, always. The exchange cost is trivial next to the loss. 2. **Hops that read donor records** — narrow these, because the reach is a disclosure rather than a loss and is still worth bounding. 3. **Hops that compute or format** — statement rendering, currency lookup, a template service. A shared zone audience is usually the honest description: nothing behind them is worth reaching. That gives a chain with two or three narrowed edges instead of ten, and the rule generalises: **narrow where the blast radius changes, not where the call graph has an edge.** It is also the reversible direction. Starting coarse and narrowing a specific hop later is an incremental change to two components. Starting with a per-service audience everywhere and discovering the exchange endpoint cannot carry it is an unwind across the whole chain. ## How to tell whether the narrowing is real Two checks, both cheap, and a design that fails either is paying for nothing: - **Does the callee refuse a token whose `aud` is not itself?** The caller obtains the narrowed token; only the callee can enforce the narrowing. If a callee accepts any well-signed token from the issuer, every exchange above it is decoration and the chain is transitive again. - **Can you name the audience for every edge in the call graph?** If a service was added and nobody chose, it inherited the zone audience, and the bound you believe you have stopped describing the system some deploys ago.

  • What does narrowing each exchanged token to one callee not prevent?
    A compromised service doing what it is allowed to do. Reconciliation owned by an attacker can still obtain the token it is entitled to and call statement generation with it. Narrowing bounds the reach of a stolen token, not the authority of a service acting within its rights — that is a different control, in a different place.
  • One service in the chain fans out to three callees. What does that do to the exchange count?
    It triples that level. Exchange load is the number of edges in the call graph, not the depth of the deepest path, so a five-deep chain with one three-way fan-out is seven exchanges per request rather than four. Size the exchange endpoint from the call graph, and re-size it when a fan-out is added.
  • What breaks first when the exchange endpoint degrades rather than fails outright?
    The tail, everywhere at once. Each hop's added round trip grows, the chain's p99 accumulates all of them, and the timeouts surface at whichever service happens to be last — which is never the one at fault. Alert on exchange latency per hop; availability alone will stay green through the whole incident.

Issuing a fresh key at every internal door is safer than one master key, right up to the point where the key desk is the reason nobody can get anywhere.

saying these in an interview costs you the question

  • Says exchanging at every hop is always the safer design
  • Ignores that the exchange endpoint is now in series with every internal call
  • Assumes narrowing the audience stops a compromised service calling its callee
  • Sizes exchange load from chain depth and forgets fan-out
  • Treats an audience no callee verifies as a control
  • Picks a per-service audience before knowing which hops differ in risk