Which SDKs call GitHub Models, and what changes in OpenAI SDK code to target it?
answer
- one wire format, two clients
- three edits: url, key, model
- publisher-qualified identifiers
- compatible on inference routes only
- capabilities vary per model card
basics
~20 sTwo client styles work: the Azure AI Inference SDK, and any OpenAI-compatible client pointed at the GitHub Models base URL. With the OpenAI SDK you override base_url, pass the GitHub token as api_key, and use publisher-qualified model ids such as openai/gpt-4o-mini.
solid answer
~40 sThe endpoint speaks a chat-completions-shaped protocol, so you have two idiomatic options. The **Azure AI Inference SDK** (`azure-ai-inference`) builds a `ChatCompletionsClient` from the endpoint plus an `AzureKeyCredential` wrapping your GitHub token, and calls `complete(...)`. The **OpenAI SDK** works if you construct the client with `base_url="https://models.github.ai/inference"` and `api_key=<GitHub token>`; from there `chat.completions.create(...)` behaves as usual, including streaming and tool calls where the chosen model supports them. Two things bite. First, model identifiers are **publisher-qualified** — `openai/gpt-4o-mini`, not `gpt-4o-mini` — so a copy-pasted vendor snippet fails on the model name. Second, OpenAI-compatible means the inference routes, not the account-scoped ones: stateful and org-level APIs such as assistants, file storage, batch jobs and fine-tuning are not part of this surface, and per-model parameter support varies by publisher.
code
python · 13 linesimport os
from openai import OpenAI
client = OpenAI(
base_url="https://models.github.ai/inference",
api_key=os.environ["GITHUB_TOKEN"],
)
response = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Say hello in one word."}],
)
print(response.choices[0].message.content)go deeper
Remember the three edits to OpenAI SDK code: base_url to the GitHub Models endpoint, api_key set to a GitHub token, and a publisher-qualified model id.
Explain that compatibility covers the inference request shape only — stateful and account-scoped APIs are absent, and tool calling or JSON output support varies per model and publisher.
Demonstrate that you feature-test the exact model id you ship, centralise retry and error handling at the HTTP layer since both clients share the wire format, and keep endpoint, credential and model in configuration.
Own the client choice as an architectural bet: which SDK minimises rewrite cost across the vendors you might land on, and what abstraction you allow teams so a model swap never reaches business logic.
## Two clients, one wire format GitHub Models exposes an inference API whose request and response bodies follow the widely-copied chat-completions shape: a `messages` array of role/content objects, parameters like `temperature` and `max_tokens`, and a response carrying `choices[0].message.content` plus a usage block. Anything that speaks that shape can talk to it. **Azure AI Inference SDK.** Construct `ChatCompletionsClient(endpoint=..., credential=AzureKeyCredential(token))` and call `complete(model=..., messages=[...])`, building messages from the `SystemMessage`/`UserMessage`/`AssistantMessage` types. There is a matching `EmbeddingsClient` with an `embed(...)` call for embedding models in the catalog. This is the client the catalog's own samples lean on, and it is publisher-neutral by design. **OpenAI SDK.** Instantiate the client with `base_url` set to the GitHub Models inference URL and `api_key` set to your GitHub token, then use `chat.completions.create(...)` normally. This is popular because existing code, wrappers and framework integrations already speak it. ## What actually changes in OpenAI SDK code Exactly three things: 1. **`base_url`** — point it at `https://models.github.ai/inference` instead of the vendor default. 2. **`api_key`** — supply the GitHub token; the SDK just puts it in the bearer header. 3. **`model`** — use the publisher-qualified identifier. This is the step people miss. The same model is `gpt-4o-mini` in one vendor's own API and `openai/gpt-4o-mini` here, because a multi-publisher catalog needs a namespace. Meta, Mistral, Microsoft and DeepSeek models are likewise prefixed with their publisher. Everything else — message construction, streaming iteration, reading `choices` — is unchanged, which is precisely why the prototype-to-production move later is cheap. ## Where the compatibility stops "OpenAI-compatible" is a statement about *request shape on the inference routes*, never a promise of feature parity. Expect these gaps: - **Account-scoped and stateful APIs are absent.** Server-side conversation state, hosted file storage, vector stores, batch job submission and fine-tuning belong to a vendor's own platform account. GitHub Models is inference against a catalog; there is no account there to hold that state. - **Per-model parameter support varies.** The catalog spans publishers whose models genuinely differ: some support tool calling, some structured JSON output, some vision inputs, some none of these. A parameter accepted by one model can be rejected or quietly ignored by another under the same base URL. The model card in the catalog is the authority, and your integration tests should cover the specific model you ship. - **Reasoning-style models behave differently.** Models that emit internal reasoning treat sampling parameters and token accounting on their own terms; do not assume settings carry across families. - **Vendor-specific extensions do not exist here.** Anything a provider added outside the common chat-completions surface is unlikely to be plumbed through a multi-publisher gateway. ## Practical consequences - **Do not hardcode the model id.** Put base URL, credential and model in configuration. Then the same binary runs against GitHub Models in development and a paid endpoint in production, with the publisher prefix stripped or replaced by config rather than by an edit. - **Feature-detect rather than assume.** If your code path depends on tool calling or JSON-schema output, verify it against the exact model id you deploy, not against the family name. - **Write the error handling once.** Because the wire format is shared, a 429 or a malformed-request error looks the same from either client; centralise retry and logging around the HTTP layer instead of the SDK. - **Pick a client for the right reason.** The OpenAI SDK maximises reuse of existing code and ecosystem integrations. The Azure AI Inference client is the neutral option and maps naturally onto a later move to a cloud-hosted deployment of the same models.
- You copy a vendor snippet, change only the base URL and key, and get an error about the model. Why?Because GitHub Models namespaces model identifiers by publisher. A snippet written against a vendor's own API uses the bare name, while the catalog needs the qualified form such as openai/gpt-4o-mini or a meta/… identifier for Llama models. The fix is the model argument, not the URL or credential — and it is the reason to keep the model id in configuration rather than inline.
- Which OpenAI SDK features should you expect not to work against this endpoint?The account-scoped and stateful ones: server-side conversation state, hosted file storage and vector stores, batch job submission, and fine-tuning. Those depend on a vendor platform account that does not exist behind a multi-publisher catalog. Inference routes — chat completions, streaming, and embeddings for embedding models — are what the compatibility covers.
- Two catalog models from different publishers behave differently for the same request. Is that a bug?No — a shared request shape does not mean shared capabilities. Tool calling, structured JSON output and vision inputs are per-model properties, so a parameter honoured by one model may be rejected or ignored by another under the same base URL. Check the model card and pin integration tests to the exact model identifier you deploy.
saying these in an interview costs you the question
- Uses bare vendor model names instead of publisher-qualified ids
- Assumes every OpenAI SDK method works against the endpoint
- Believes OpenAI-compatible implies identical model capabilities
- Hardcodes the model id so the production swap becomes a code change
- Expects fine-tuning or hosted file storage from the catalog endpoint