How do you authenticate to the GitHub Models inference API from code and from CI?
answer
- bearer header, GitHub identity
- fine-grained token, one permission
- CI needs no stored secret
- permissions block on the job
- authorised is not the same as unthrottled
basics
~20 sSend a GitHub token as an HTTP bearer credential. Locally that is a fine-grained personal access token carrying the Models read permission; inside GitHub Actions it is the workflow's built-in GITHUB_TOKEN, after the job declares the models: read permission.
solid answer
~50 sAuthentication is a plain `Authorization: Bearer <token>` header — no vendor key, no OAuth dance. Outside CI you mint a fine-grained personal access token and give it the **Models** read permission; a token without it gets a 401/403 rather than an inference response. Inside GitHub Actions you normally do not create a secret at all: the automatically-provisioned `GITHUB_TOKEN` can reach the endpoint once the workflow or job block declares `permissions: models: read`, and you pass it to your step as an environment variable. Because both SDK styles map the token onto the same header, the code is identical either way — the Azure AI Inference client wraps it in `AzureKeyCredential`, the OpenAI client accepts it as `api_key`. Access is also governable above the individual: organisation and enterprise administrators control whether GitHub Models is available to members at all, so a valid token can still be refused.
code
bash · 4 linescurl -sS https://models.github.ai/inference/chat/completions \
-H "Authorization: Bearer $GITHUB_TOKEN" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"ping"}]}'go deeper
Know that the credential is a GitHub token sent as a bearer header, and that a personal access token needs the Models read permission explicitly granted.
Explain the two paths — fine-grained PAT locally, built-in workflow token in Actions with permissions: models: read — and why the permissions block is exhaustive per job.
Show you separate the failure modes: 401/403 for credential or policy, 404 for a bad model identifier, 429 for quota, and that you never respond to throttling by rotating credentials.
Own the policy layer: whether the org enables GitHub Models at all, what token lifetimes and permission sets you mandate, and how credentials are held in configuration so a provider swap is a values change.
## The credential GitHub Models does not issue its own API keys. Every call carries a GitHub token in the standard header: `Authorization: Bearer <token>` Which token depends on where the code runs. ### Local and server code: a fine-grained personal access token Create a fine-grained PAT and grant it the **Models** permission at read level. Fine-grained tokens are the modern shape: they carry an explicit expiry, they are scoped to an account or organisation, and each capability is granted individually rather than by a broad scope. A token missing the Models permission is the single most common first-call failure — the request is syntactically fine and the endpoint rejects it as unauthorised, which people misread as a bad URL. Treat the token like any other secret: environment variable or secret manager, never a literal in source, and rotate on the expiry you set. It is worth noting that this token is *your GitHub identity*, not a narrow inference key — one more reason to keep the permission set minimal and the lifetime short. ### GitHub Actions: the built-in workflow token A workflow gets a short-lived `GITHUB_TOKEN` automatically, valid only for the duration of the run. Its capabilities are declared in the workflow's `permissions` block, and model access is one of them: ``` permissions: contents: read models: read ``` Declare it at workflow level or on the individual job, then expose the token to the step that needs it (commonly as `GH_TOKEN` for the CLI, or your own variable name for a script). This is the preferred CI path because there is no long-lived secret to store, share or leak — the credential dies with the run. Remember that a `permissions` block is exhaustive for the job it appears on: listing `models: read` alone silently drops the other permissions the job relied on, which is why the example above also keeps `contents: read` for the checkout step. ## How the SDKs consume it Both supported client styles reduce to the same header: - The Azure AI Inference client takes it as a key credential object alongside the endpoint URL. - The OpenAI client accepts it in its `api_key` argument with the base URL overridden. That symmetry matters when you later migrate: the credential slot in your configuration does not change shape, only its value and the endpoint beside it. ## Layers above the token A valid, correctly-permissioned token is necessary but not sufficient: - **Organisation and enterprise policy.** Admins decide whether members may use GitHub Models. If it is disabled for the org, member tokens are refused regardless of permission bits. - **Rate limiting.** Authorisation and quota are different gates. A token that authenticates perfectly still receives 429 responses once the account's allowance is exhausted. - **Model availability.** A model can be present in the catalog but not enabled for your account or region, which surfaces as an error naming the model rather than the credential. ## Debugging the failure modes Work through them in order: 1. **401 / 403** — the token is missing, expired, or lacks the Models read permission; or the org has not enabled Models. Regenerate with the permission explicitly checked and retry. 2. **403 in Actions specifically** — nearly always a missing or overwritten `permissions: models: read` on the job. 3. **404 on the model** — usually a model identifier problem rather than auth; identifiers on this endpoint are publisher-qualified. 4. **429** — you are authenticated and over quota; back off rather than re-minting tokens. ## Practices worth stating in an interview Use the workflow token in CI rather than a stored PAT; keep PAT lifetimes short and permissions minimal; never log the header; and hold the credential in configuration next to the base URL so that pointing the same code at a paid provider endpoint later is a values change, not a code change.
- Why prefer the workflow token over a stored personal access token in CI?The workflow token is minted per run and expires with it, so there is no long-lived secret to store, rotate, or leak through a fork or a log. Its capabilities are declared in the workflow file, which makes model access reviewable in the diff. A stored PAT is a standing credential tied to a person, and it keeps working after that person leaves the team.
- Your call worked yesterday and now returns 403 in CI only. Where do you look first?At the job's permissions block. A job-level permissions declaration replaces the workflow-level one wholesale, so adding an unrelated permission to that job can drop models: read. Check whether the block was edited, whether the job is running from a fork with reduced token capabilities, and whether an org admin disabled GitHub Models for members.
- Does a correctly-permissioned token guarantee your request succeeds?No. Authorisation and quota are separate gates: a valid token still gets 429 once the account's per-minute or per-day allowance is exhausted, and the org or enterprise policy can disable GitHub Models for members entirely. A model that exists in the catalog but is not enabled for your account fails on the model identifier, not the credential.
saying these in an interview costs you the question
- Thinks GitHub Models issues its own separate API keys
- Stores a long-lived PAT as a secret instead of using the workflow token
- Forgets that a job-level permissions block replaces the workflow-level one
- Treats a 429 as an authentication problem and re-mints tokens
- Assumes a valid token bypasses organisation-level policy