skip to content

How should credentials flow from a user through an agent to a third-party tool server?

level: principalimportance: should knowfreq 34%

answer

  1. Ask: who was this token issued for?
  2. Deputies act with borrowed authority
  3. No passthrough between hops
  4. The audience claim must name you
  5. One narrowly scoped token per downstream

basics

~20 s

Each hop gets its own token. A tool server must never replay a token minted for something else at a downstream API, and every service must reject tokens whose audience does not name it — otherwise the server becomes a confused deputy acting with authority it was never granted.

solid answer

~50 s

Model each hop as its own trust relationship. The agent platform holds a token for the tool server, and the tool server obtains a separate, narrowly scoped token for each downstream API it calls — by exchanging its own credential with delegation context, never by forwarding whatever token arrived. Every service validates that the `aud` claim in the token it received actually names it, and rejects otherwise; this is the standard OAuth 2.1 audience-validation rule and it is what stops replay across services. Token passthrough — the tool server taking the caller's token and reusing it downstream — collapses those boundaries: the downstream API cannot tell whether the agent or the user initiated the call, scopes cannot be narrowed per hop, and a compromised tool server inherits the caller's full authority. Add per-user consent recorded at the client rather than a blanket admin grant, short TTLs, and no refresh tokens parked on tool servers.

code

python · 9 lines
python
def check_audience(claims: dict, this_service: str) -> None:
    aud = claims.get("aud")
    allowed = aud if isinstance(aud, list) else [aud]
    if this_service not in allowed:
        raise PermissionError("token was not issued for this service")


check_audience({"aud": ["logs-api"]}, "logs-api")  # ok
check_audience({"aud": ["logs-api"]}, "billing-api")  # rejected

go deeper

for a junior

Know that an access token is issued for one specific service and should not be reused at another, and that a tool server should not simply forward the token it received.

for a middle

Explain audience validation and what passthrough breaks: no scope narrowing, lost attribution, and a compromised server holding live tokens for every caller.

for a senior

Design the chain — per-hop tokens obtained by exchange with delegation context, per-client consent, short TTLs, no refresh tokens on tool servers — and be able to say what each control prevents.

for a principal

Own where the line sits between hops that require their own token and hops that may share a service identity, and the delegated-versus-agent-owned authority policy for long-running work. Defend the extra complexity against delivery pressure with the confused-deputy argument.

## The shape of the problem An agent architecture almost always has three or more parties: the user, the agent platform, a tool server, and one or more downstream APIs that actually hold the data. Credentials have to move through that chain, and the naive implementation — pass the user's token along at every hop so everything "just works" — recreates one of the oldest failures in authorization, the confused deputy. A confused deputy is a component that holds authority and can be tricked into exercising it for someone else. A proxying tool server is a natural deputy: it sits between an agent driven by untrusted content and an API holding real data. If it forwards tokens, or holds a broad pre-consented grant it will use for any caller, then persuading the agent is equivalent to persuading the API. ## Why token passthrough is prohibited in practice Token passthrough means a service accepts a token minted for someone else and replays it downstream. It is attractive because it is one line of code, and it is wrong for several independent reasons: - **Audience is the point of a token.** An access token names the service it was issued for. Replaying it at a different service means either that service does not check the audience — in which case any token from anywhere works there — or the check fails and you have built something fragile. - **No scope narrowing.** The forwarded token carries the caller's full scope. The tool server cannot reduce it to the operation it actually needs, so every downstream call runs at maximum authority. - **Attribution is destroyed.** The downstream API's logs show the user, not the tool server acting for the user, so nobody can reconstruct which component initiated a call. - **Blast radius on compromise.** A compromised or poisoned tool server holds live tokens for every caller who ever used it. The corresponding positive rule is audience validation: a resource server accepts a token only if the token's audience claim names that resource server. It is a few lines of code and it is the check that makes every other boundary meaningful. ## What to do instead **Per-hop tokens with delegation context.** The tool server authenticates as itself, and where it must act for a user it obtains a downstream token through a token-exchange flow that carries the delegation — the resulting token names the downstream API as its audience, carries a scope matching just that operation, and records both the acting service and the user. Standard token exchange exists precisely for this. **Per-client consent, not blanket grants.** A proxying server that registers once with a broad grant and serves many clients has consented on behalf of users who never saw a prompt. Consent should be recorded for the specific client and the specific scopes, so that an attacker who can reach the proxy does not automatically inherit an existing approval. **Short lifetimes, minimal storage.** Access tokens should be short-lived, and refresh tokens should not sit on a tool server if the architecture can avoid it — a long-lived refresh token on a third-party component is a standing key to a user's account. **The agent's own identity where delegation is not required.** Plenty of agent work is not on behalf of a specific person. Those calls should use the agent's own principal with its own narrow role, which is simpler to audit and to revoke than a delegated chain. ## The tradeoffs a lead actually owns This design is more moving parts. Per-hop exchange needs an authorization server that supports it, every service needs correct audience validation, and debugging a failed call now means reading three tokens instead of one. Teams under delivery pressure reach for passthrough for exactly this reason. The judgment call is where to spend the complexity. A reasonable position: any hop that crosses a trust boundary — into a third-party tool server, into a system holding another team's or another tenant's data — gets its own token, non-negotiably. Hops entirely inside one trust domain may share a service identity, with the user recorded as delegation context rather than as the token subject. Draw that line explicitly and write it down, because the default in every codebase is to forward whatever arrived. The second judgment call is delegated user authority versus agent-owned authority. Delegation gives you least privilege naturally — the agent can never exceed the requesting user — but it multiplies tokens and makes long-running or scheduled agent work awkward, since the user's session may be gone. Agent-owned identities are simpler and auditable but require you to decide the agent's standing authority up front. Most mature platforms use both and are explicit about which flows use which. ## Answering this well Open with the confused deputy, state the no-passthrough rule and audience validation as the two hard requirements, then describe per-hop tokens with delegation and per-client consent. Finish with the honest tradeoff — complexity versus containment — and where you would draw the line. An interviewer at this level is testing whether you can defend the boundary under delivery pressure, not whether you can recite the flow.

  • What exactly goes wrong if a downstream API skips audience validation?
    It accepts any token the issuer signed, regardless of which service that token was minted for. Then a token obtained for a low-value service becomes a valid credential at a high-value one, and every service in the estate shares a single effective trust level. Audience validation is what keeps separately issued tokens from being interchangeable, so skipping it silently removes the boundary between services.
  • When is it acceptable for an agent to act under its own identity rather than the user's?
    When the work is not attributable to a specific person's authority — scheduled maintenance, fleet-wide reporting, background reconciliation — or when the user's session will not outlive the task. Use the agent's own principal with a narrow standing role, and record which human or system requested the work as separate context. What you must not do is use the agent's broad identity to perform work the requesting user would not be allowed to do.
  • Why is a proxying tool server with one blanket grant a confused deputy?
    Because it holds pre-consented authority and will exercise it for whoever reaches it. Users never saw a consent prompt for the specific client, so any caller that can talk to the proxy inherits an approval granted for someone else's purpose. Per-client consent and per-hop tokens remove the standing authority that makes the deputy exploitable.

saying these in an interview costs you the question

  • Forwarding the caller's token unchanged to downstream APIs
  • Skipping audience validation because the issuer is trusted
  • One blanket admin grant covering every user of a proxy
  • Storing long-lived refresh tokens on a third-party tool server
  • Assuming a token is safe to reuse anywhere in the same estate

context