skip to content

GitHub Models

GitHub's catalog and playground let you try and then call models using the GitHub token you already have. It is the lowest-friction way to prototype before committing to a paid provider account.

on this pageshow

questions

6

What is GitHub Models, and when would you use it instead of a paid provider account?

level: juniorimportance: must knowfreq 62%

answer

  1. catalog, playground, one endpoint
  2. credential you already have
  3. free tier is for prototyping
  4. account-scoped limits, not per app
  5. swap base URL to graduate

basics

~20 s

GitHub Models is a hosted catalog and playground of models you can call with your GitHub token. It gives free, rate-limited inference for prototyping, so you can compare models from several publishers before opening any paid provider account.

solid answer

~40 s

GitHub Models bundles three things: a **catalog** of hosted models from several publishers (OpenAI, Meta, Mistral, Microsoft, DeepSeek, xAI and others), a **playground** for trying them in the browser, and an **inference endpoint** you call from code at `https://models.github.ai/inference`. The selling point is friction removal: you authenticate with a GitHub token you already have instead of creating a separate account, billing profile and key with each vendor. That makes it the natural place to prototype a feature, A/B a prompt across model families, and run prompt evaluations in CI. The free allowance is rate-limited and shaped for experimentation rather than production traffic, so real workloads either enable paid usage or move to a dedicated provider endpoint. Because the API is chat-completions shaped, that move is mostly an endpoint and credential swap.

go deeper

for a junior

Be able to say what the three pieces are — catalog, playground, inference endpoint — and that you authenticate with a GitHub token instead of a vendor key.

for a middle

Explain why the free tier suits prototyping and evaluation but not production: account-scoped rate limits, no capacity guarantee, and a paid-usage path that still is not a provider SLA.

for a senior

Show that you keep base URL, credential and model id in configuration from the first commit, so the prototype-to-production move is a deploy rather than a refactor, and that you re-run evaluations after switching endpoints.

for a principal

Frame it as procurement leverage: a place to compare vendors on real prompts before committing spend or signing a contract, with an explicit exit criterion for when a workload must move off it.

## What it is GitHub Models is GitHub's front door to hosted language models. Three surfaces get conflated, so separate them: 1. **The catalog** — a browsable list of models from multiple publishers, each with a model card and sample code. The line-up spans OpenAI, Meta Llama, Mistral, Microsoft Phi, DeepSeek, Cohere and xAI models, and it changes over time as models are added and retired. 2. **The playground** — an in-browser chat surface where you send a prompt, tune parameters, and compare two models side by side on the same input. 3. **The inference API** — an HTTP endpoint (`https://models.github.ai/inference`) that serves every catalog model behind one base URL, one credential and one request shape. ## Why the token matters The defining property is the credential. Calling five vendors directly means five accounts, five billing relationships, five key-rotation stories and five SDKs. GitHub Models collapses that to a GitHub token with a Models read permission. Inside GitHub Actions it goes further: the automatically-provisioned workflow token can be granted model access, so a CI job can call a model with no secret to manage at all. ## Where it fits in a project's life - **Model selection.** You have a summarisation feature and four plausible models. Run the same prompt against all four in the playground or via the CLI, look at output quality, latency and token counts, then commit to one. - **Prototyping.** Wire the endpoint into a spike branch, ship a demo, decide whether the feature is worth building before anyone signs a vendor contract. - **CI evaluation.** Store prompts as files in the repository and run them on pull requests so a prompt edit is reviewed like code. - **Learning and internal tools.** Low-volume, non-critical usage that never justifies procurement. ## Where it does not fit The free tier is rate-limited on requests per minute, requests per day, tokens per request and concurrency, and those limits are scoped to your account rather than to one application. A user-facing product sharing an account-level quota with every other experiment in the org is a capacity incident waiting to happen. Paid usage exists and is billed per token once enabled on the account, but even then you should ask whether you want production traffic depending on a developer-experience surface rather than a provider or cloud endpoint with its own SLA and regional controls. ## What it is not - **Not a local runtime.** Weights run on GitHub's hosted infrastructure; nothing is downloaded to your machine. Tools that run open weights locally are a different category. - **Not GitHub Copilot.** Copilot is a coding assistant product; GitHub Models is raw model access for your own code. They are billed and configured differently, though your Copilot plan can influence the model rate limits you get. - **Not a single vendor's API.** The endpoint is multi-publisher, and a model's identifier there carries the publisher, so the same underlying model is named differently here than in the vendor's own API. ## The migration story The reason teams tolerate the limits is that leaving is cheap. Code written against the endpoint uses a chat-completions request body, so moving to a vendor endpoint or a cloud-hosted deployment means changing a base URL, a credential and a model identifier. Keep those three values in configuration from day one, and the prototype-to-production step stays a config change rather than a rewrite. What does **not** travel automatically is anything vendor-specific you leaned on, and any evaluation baseline you recorded — re-run your evals after the switch, because the same model family served from a different endpoint can differ in defaults and available parameters.

  • What actually changes in your code when you graduate from GitHub Models to a production provider endpoint?
    Three values: the base URL, the credential, and the model identifier — GitHub Models qualifies models with their publisher, while a vendor's own API uses its bare name. Keep all three in configuration and the switch is a deploy, not a rewrite. Re-run your prompt evaluations afterwards, since defaults and supported parameters can differ between endpoints.
  • Would you put user-facing production traffic on GitHub Models?
    Not by default. The free allowance is rate-limited and scoped to the account, so unrelated experiments in the same org compete with your product for quota, and the surface is positioned for developer experimentation rather than backed by a production SLA. Prototype there, then move to a provider or cloud endpoint with capacity and regional guarantees you control.
  • How does GitHub Models differ from running open weights locally?
    GitHub Models is hosted inference: you send requests over HTTPS and never hold the weights. Local runtimes download open-weight models and run them on your own GPU or CPU, which trades network dependency and per-token cost for hardware cost, VRAM sizing and operational work. The two also differ in model availability — closed models are only reachable through hosted APIs.

saying these in an interview costs you the question

  • Thinks GitHub Models downloads and runs weights locally
  • Assumes it only serves OpenAI models
  • Confuses it with GitHub Copilot's coding assistant
  • Believes the free tier is sized for production traffic
  • Expects a separate vendor API key per model

context

open as a page

How do you authenticate to the GitHub Models inference API from code and from CI?

level: middleimportance: must knowfreq 55%

basics

~20 s

Send a GitHub token as an HTTP bearer credential. Locally that is a fine-grained personal access token carrying the Models read permission; inside GitHub Actions it is the workflow's built-in GITHUB_TOKEN, after the job declares the models: read permission.

open as a page

How do you run GitHub Models prompt evaluations from a GitHub Actions workflow?

level: middleimportance: should knowfreq 35%

basics

~20 s

Store prompts as .prompt.yml files in the repository, grant the job models: read so the workflow token can call inference, install the gh-models CLI extension, and run gh models eval on those files so prompt changes are reviewed and scored like code.

open as a page

Which SDKs call GitHub Models, and what changes in OpenAI SDK code to target it?

level: middleimportance: should knowfreq 50%

basics

~20 s

Two client styles work: the Azure AI Inference SDK, and any OpenAI-compatible client pointed at the GitHub Models base URL. With the OpenAI SDK you override base_url, pass the GitHub token as api_key, and use publisher-qualified model ids such as openai/gpt-4o-mini.

open as a page

Your GitHub Models prototype starts returning HTTP 429 under load — how do you respond?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Treat 429 as quota, not a bug: GitHub Models throttles per account on requests per minute and per day, tokens per request and concurrency, with tighter limits on larger models. Back off using the wait the error states, and move sustained load to a paid endpoint.

open as a page

When should an organisation standardise on GitHub Models, and when must a workload move off it?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Standardise on it for evaluation and prototyping, where one credential across publishers speeds model selection. Move a workload off once it needs guaranteed capacity, regional or contractual data controls, or dedicated quota, since limits are account-scoped and admins gate availability org-wide.

open as a page