skip to content

Scope and Authorisation

Permission for an AI test comes from the app owner and from the model provider's terms, and staging is a different agreement from production. Interviewers ask who signed for the target you hit.

on this pageshow

explore

questions

5

Your client owns a customer-facing chat application built on a third-party hosted model API. Before you red-team that application, whose permission do you need besides the client's, and why?

level: juniorimportance: must knowfreq 72%

answer

  1. two signatures: app owner and provider
  2. traffic runs on provider infrastructure
  3. acceptable-use terms bind the tester
  4. which account gets suspended
  5. local stand-in for prohibited-output work

basics

~20 s

You also need the model provider's. The client owns the application, but the traffic you generate lands on the provider's infrastructure under their acceptable-use terms, which often restrict deliberate attempts to defeat safety measures. Check the provider's testing policy, and get the client's written sign-off naming the accounts and endpoints you may hit.

solid answer

~50 s

There are two permissions, and only one of them is the client's. The client authorises you against **their** application: the domains, the accounts, the data, the hours. That authorisation covers everything the client owns. It does not cover the hosted model. Every adversarial prompt you send is executed on the provider's infrastructure, billed to a provider account, and evaluated against the provider's acceptable-use and safety terms. Those terms typically prohibit deliberately eliciting prohibited content and may require registration or a separate agreement for safety testing. A client signature cannot waive a contract between the client and the provider, and certainly cannot waive one you are not party to. So the practical answer is: read the provider's usage and testing policy, confirm which account the traffic will be billed to, and if the policy requires notification or an authorised-testing path, use it. Where the provider will not clear the traffic, negotiate a self-hosted or open-weights stand-in for the parts of the test that require producing genuinely prohibited output.

go deeper

for a junior

Names both parties: the application owner and the model provider, and knows the provider has published usage terms that testing must respect.

for a middle

Explains that the traffic executes on the provider's infrastructure under an agreement the client cannot waive on the provider's behalf, and identifies which account carries the risk.

for a senior

Writes the scope so the provider account, key, volume ceiling and enforcement escalation path are named, and moves prohibited-content work onto a locally-run stand-in.

for a principal

Decides organisational policy: which classes of AI engagement the firm accepts, what warranty it requires from clients, and how it structures testing so a provider's enforcement never takes a client's production down.

## The estate is split, and authorisation follows ownership A *hosted model API* means the client's application does not run inference. It assembles a request — system prompt, conversation history, retrieved documents, tool definitions — sends it over HTTPS to an endpoint the provider operates, and receives generated tokens back. The provider meters those tokens against an API key, and that key belongs to a provider *account* governed by a contract the client signed. So the deployment has **two owners**: - **The client** owns the prompt template, the retrieval corpus, the tool wiring, the output handling, the session store and the domain name. - **The provider** owns the model weights, the serving fleet, the server-side safety stack and the terms of use that govern all of it. Authorisation is granted by an owner, for that owner's estate. One signature therefore cannot cover both halves, no matter how comprehensively it is worded. ## What the provider's terms actually bear on Three clauses decide scope. 1. First, *prohibited use*: providers generally forbid deliberately eliciting disallowed content or circumventing safety mitigations, and several carve out a narrow authorised path — a registered safety-testing or researcher programme — for exactly the work a red team wants to do. 2. Second, *volume and abuse*: sustained machine-rate attempts to break a safety boundary are the precise signature automated enforcement is built to catch. 3. Third, *attribution*: terms routinely forbid obscuring who is sending traffic, which rules out proxying around the problem and would in any case destroy the audit trail an authorised engagement depends on. Note the legal shape, because candidates get it backwards. - The client's contract with the provider **binds the client**. - The client can indemnify you; it cannot grant you a right it does not itself hold. - Acting as the client's agent inherits the client's obligations rather than dissolving them. - And a bug-bounty safe harbour published by the client covers the client's own systems and nobody else's platform. ## What it costs Reading the current usage and testing policy costs an hour. Getting onto a provider's **authorised-testing path**, where one exists, costs days to weeks of lead time and sometimes returns no answer at all — which means it belongs in the project schedule at kickoff, not in the week you planned to run. The cost of skipping it is **asymmetric**, because enforcement is account-level and automatic. If the account carrying your traffic is the one behind the live product, a suspension is a total outage of that product until a support ticket is resolved, and standard-tier support response is measured in days. Provisioning a separate project and key, by contrast, costs minutes. The direct bill matters too: adversarial prompts are long, many published techniques are multi-turn, and any judge model scoring the transcripts bills again — so a campaign's token spend is a multiple of a naive attempts-times-unit-price figure, not that figure. ## Where the reasoning misleads The famous error is treating the client's letter as **global cover**. The quieter and more damaging one runs the other way. A team reads the terms, concludes the model layer is off limits, tests only the application, and files a report containing no model-layer findings — and the reader takes zero findings there as evidence of a robust model. It is nothing of the sort. It is an **absence manufactured by a contract**. Any coverage figure you quote is a fraction whose denominator is *what you were permitted to run*, not the attack surface that exists. If part of that surface was closed by terms rather than by testing, the deliverable has to say so in the same breath as the number, or the number is a lie told by omission. ## What goes in the scope document - The application and its environments; - the exact endpoints; - the provider account, project and key the test traffic bills to; - the calendar window; - a request-rate and spend ceiling; - a **warranty** from the client that it holds the right to authorise testing of its own deployment; - a clause covering the harmful content you will deliberately generate, where it is stored and who may read it; - and an **escalation contact**, reachable within a stated time, for the case where enforcement fires mid-test. ## What you check before day one - The provider's **currently published** usage and testing policy, not the copy you read last year. - Which key the tooling is actually configured with — read the environment variable the runner exports, do not trust the runbook. - The spend and rate ceilings on that project. - And whether an **open-weights model** you host yourself can absorb the parts of the campaign that would breach the terms, so that the traffic which does reach the provider looks ordinary.

  • The client insists their contract with the provider is their problem and tells you to proceed. What do you do?
    Put the risk in writing: the exposure is account enforcement and a possible production outage, and it falls on them. Ask them to confirm in writing that they have the right to authorise it, and keep the high-volume prohibited-content work on a locally-run model where the provider's terms do not apply.
  • Which parts of the test can you keep entirely inside the client's estate?
    Anything that exercises the application layer rather than model safety: prompt-template extraction attempts, retrieval-corpus poisoning, tool-invocation abuse, output handling, and authorisation checks between users. These consume inference, but they do not require eliciting prohibited content from the provider.

saying these in an interview costs you the question

  • Says the client's authorisation letter covers everything because the client pays the provider's bill.
  • Has never read a model provider's acceptable-use or testing policy and assumes any testing is fine.
  • Plans to run the whole campaign on the production API key the live product uses.
  • Treats a public bug-bounty safe harbour as covering an unrelated third party's platform.
  • Cannot say what would happen if the provider's abuse enforcement fired mid-engagement.

context

open as a page

You are asked to red-team a staging copy of a client's LLM application instead of production. What differences between the copy and production would stop your results transferring, and what do you record about the environment you tested?

level: middleimportance: must knowfreq 62%

basics

~20 s

Staging results transfer only as far as the configuration matches. Check the system prompt, the model identifier and routing, the retrieval corpus, tool permissions, and whether input and output filters are enabled the same way. Record all of it with the test dates, and say in the report that findings apply to that configuration.

open as a page

A client asks you to red-team an assistant whose tools can send email, file tickets and call a partner company's API on a user's behalf. How do you scope what the assistant may actually do during the test, and what goes in writing?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The assistant's tools reach systems your client does not own, so a successful injection could send real mail or write to a partner. Agree in writing which tools stay live, which are stubbed, which recipients and accounts are test-only, and who you call if an action escapes into a real third-party system.

open as a page

You plan to run an automated adversarial-prompt campaign of tens of thousands of requests against a client's production LLM endpoint, billed to a metered model-provider account. Beyond permission to test, what do you agree before you start?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Agree whose provider account and key pay for it, a spend cap, a request-rate ceiling and a stop condition. Warn that sustained adversarial traffic can trip the provider's abuse detection and throttle or suspend the account the live product uses, so ask for a separate key and a named contact who can unblock it.

open as a page

The model behind a client's application is a hosted third-party service the client cannot pin or freeze. How do you write scope and findings validity so the report still means something three months later?

level: principalimportance: should knowfreq 38%

basics

~20 s

Record the exact model identifier, endpoint and dates tested, and state that findings describe that snapshot. Write a validity window and named re-test triggers into the scope: a provider model update, a system-prompt change, a new tool or corpus. Push the client toward durable application-side controls rather than one-off prompt patches.

open as a page